When localization means more than just EN, DE, IT, ES, FR et al.
Posted on by Roger Tuan , Ronak Ganatra and Silvano Stralla
TLDR
A customer's S'gaw Karen content was showing weird dotted circles in Firefox. Instead of fixing it for one language, we went down a rabbit hole that eventually fixed it for 43 scripts.
The CMS now falls back to Google's Noto fonts for scripts your OS can't draw, and only loads them when a page actually needs them.
Turns out when it comes to localization, locale codes, fallbacks, text direction, and fonts still very much matter too.
Our fix covers the CMS, not your frontend. There's a checklist further down to help you go font hunting.
Wait, what's S'gaw Karen?
უკეთესი ការបង្ហាញ ለ ရှားပါး ⵜⵉⴼⵉⵏⴰⵖ
That's literally "better display for rare scripts" in Georgian, Khmer, Amharic, Burmese and Tifinagh.
Until last week, whether the CMS showed all of them properly depended on which browser you opened it in. 🫠
We recently got a support request from a customer to "Fix S'gaw Karen text rendering in Firefox". This led to some googling.
To save you the trouble: S'gaw Karen is spoken mostly in Myanmar and Thailand, and one of our customers manages their content in it. Chrome was fine. Firefox, though, was rendering those little dotted circles around characters where they definitely shouldn't be.
We fixed it (more on that in a bit). But the rabbit hole we fell into kinda really lit up things about "localization" I don't think of as often as the Holy Roman Empire Unicode. And if you run a multilingual site on a headless CMS, a good chunk of what we found applies to your frontend too.
So half of this is "here's what we did", and the other half is "here's what might be relevant on your end".
Localization is a stack
Be honest. When you say a site is "localized", you probably mean it's been translated into English, German, Italian, Spanish and French.
That's five languages that all use the same alphabet - the Latin one. Which is... not much of a stretch, tbh.
OK fine let's say you went fancier - you added language support for Mandarin too.
Localization. ✅.
Or ❓
REAL localization tho, is really a stack of decisions, and translation is just that one layer of it:
| Layer | The question it answers | What DatoCMS gives you |
|---|---|---|
| Model | Which content gets translated? | Localization per field, and translations that can be optional per model |
| Locale | Which language, for which region? | Language and region codes like en-GB and pt-BR (not just language, because German in Switzerland is COMPLETELY different from German in Austria) |
| Workflow | Who translates, who publishes, who QAs, and when does it go live? | Locale-based permissions, publishing and scheduling per locale, the AI Translations plugin, translation tools |
| Delivery | What happens when a translation is missing? | fallbackLocales, _locales and _allFieldLocales in the Content Delivery API |
| Direction | Which way does the text run? | Left to right (most), right to left (Arabic, for example), top to bottom (Kanji, for example) |
| Rendering | Can the screen actually draw it? | Noto fallback fonts for 43 scripts in the CMS (i.e. can the CMS render what you're typing, not whether the customer can see it on your website) |
If you want to dig into any of these, the docs have you covered: localization in DatoCMS, region-specific language codes, translating content with AI and localization in the Content Delivery API cover a lot on localization with Dato.
Most of the conversation happens at the top of that stack (which fields, which workflow, which AI model). The bottom is where text in less common scripts breaks, and its kinda hard to know unless you're super confident that your font choices CAN render it on the frontend for your customers.
Language ≠ script ≠ locale
This is the bit that made me go 🐰🕳️ hunting.
TLDR?
A language is WHAT people speak.
A script is HOW it's written down.
A locale is a language AND where it's used.
And they don't always line up in a very straightfoward way.
S'gaw Karen, for example, is written in the Myanmar script, same as Burmese, Mon, and Shan. So fixing Karen meant fixing Myanmar, which fixed all four. Nice bonus.
It works the other way round too. Serbian is written in both Cyrillic and Latin, and Punjabi uses Gurmukhi in India but Shahmukhi in Pakistan.
Even the language codes get weird. Karen doesn't have a two-letter ISO code. kar covers all the Karen languages as a group, and S'gaw Karen on its own is ksw. TIL.
Then there's what these scripts actually do on screen, which Latin-only reading does NOT prepare you for:
Myanmar and Balinese stack consonants on top of each other
In Myanmar and Khmer, some vowel signs get typed after a consonant but get drawn before it
Myanmar, Khmer, and Lao don't put spaces between words
Six of the 43 new scripts we now cover run right to left: Adlam, Hanifi Rohingya, Mandaic, N'Ko, Syriac and Thaana
Here's a few of them, each one writing its own name, and it should render just fine 🤞🏾🤞🏾🤞🏾🤞🏾🤞🏾🤞🏾:
| Script | Sample | Literally reads/sounds as | Used for |
|---|---|---|---|
| Myanmar | မြန်မာ | Myanmar | Burmese, S'gaw Karen, Mon, Shan |
| Georgian | ქართული | Kartuli (Georgian) | Georgian, Mingrelian, Svan |
| Ethiopic | ግዕዝ | Ge'ez | Amharic, Tigrinya, Tigre, Ge'ez |
| Khmer | ខ្មែរ | Khmer | Khmer |
| Tibetan | བོད་ | Böd (Tibet) | Tibetan, Dzongkha, Ladakhi |
| Cherokee | ᏣᎳᎩ | Tsalagi (Cherokee) | Cherokee |
| Tifinagh | ⵜⵉⴼⵉⵏⴰⵖ | Tifinagh | Tashelhit, Central Atlas Tamazight |
Tofu and dotted circles
When a script doesn't render, it usually breaks in one of two ways.
Tofu (▯). The browser checked every font it could find, none of them had that character, so it gives up and draws an empty box. It's called tofu because it looks like a block of tofu.
Dotted circles (◌). Sneakier. The browser did find a font, but the text shaping engine couldn't stitch the pieces together. Scripts like Myanmar build each syllable out of a base letter plus marks above, below or around it. When the font can't combine them, the leftover mark gets drawn on a dotted circle, which is basically the browser saying "this mark lost its letter".
That's what our customer was seeing.
So why Chrome and not Firefox? Because what you see depends on your browser, your OS, and whatever fonts happen to be installed. Every combo can grab a different fallback font for the same character, and some handle Karen while others don't.
Which means you can't just reproduce it once, fix it, and move on. There are a LOT of combinations, and you won't see most of them from your own laptop.
Fixing the category instead of the bug
The quick fix would have been to load one new font for Karen, LGTM it to Firefox, and ciao.
Pretty sure I even saw "look into it without making an obsession of it" somewhere in that ticket.
So instead of patching one language, we went a level up:
Fix it by script, not by language: Noto already ships one font per script, so one rule covers every language written in it.
Cover every script in modern use: Unicode keeps an official list of these (the Recommended and Limited Use scripts in UAX #31), so that became the scope.
Skip the ones operating systems already handle: Latin, Cyrillic, Greek, Arabic, CJK, Hangul, Devanagari, Thai and a few others render fine pretty much everywhere.
In the end (it does even matter), 43 scripts and dozens of languages now render just fine, covered by a handful of CSS rules that shouldn't need touching every time someone shows up with a new language.
If you're debugging something similar on your end, check whether your "language bug" is actually a script bug. If it is, fixing the whole script is about the same amount of work usually.
Why Noto?
During review, Silvano asked this too, for a reason other than "Google says so"?
The answer is literally in the name. Noto = "no tofu".
DID YOU KNOW THAT? LITERALLLLY TIL. Anyways.
Google built it specifically to be the fallback font that covers everything (NPR has a really nice backstory). On top of that:
It's open source under the SIL Open Font License, so we can bundle it and serve it ourselves
It's one design family across hundreds of scripts, so mixed-language text like some of my nonsense doesn't end up looking like alphabet soup.
It ships as one font per script, which is exactly how we wanted to load it
Android already uses it as its fallback, and Apple ships Noto fonts for scripts it doesn't have its own fonts for
Fun fact: We did also look at GNU Unifont, which covers a massive range of characters. But it's a pixel font, and it can't handle the combining and stacking that Myanmar needs. So it would've broken Karen before we went down the other 🐰🕳️🕳️🕳️.
No thanks.
Also, we're Italian. And Noto is a town in Sicily. Great granita.
How it works
It just does.
Each Noto font gets registered with @font-face and a unicode-range that only covers its script. The browser only downloads a font when the page actually contains a character in that range. And the fonts sit in the CMS font stack after our interface font and before the generic sans-serif, so Latin text looks exactly the same as before.
The whole thing is about 1 KB of CSS once it's minified and gzipped.
So if your projects are in Latin, Cyrillic, CJK or Arabic, nothing extra will load. If one of your editors opens a record in Tifinagh, their browser grabs one font, once. That's it.
Remember tho, this is all inside the CMS. The editor, the record forms, the previews. Your API responses haven't changed and neither has your website. That part's on 🫵
The changelog entry is here, and everything on localization in DatoCMS lives in the docs.
And in case you were curious, here's a list of all the new scripts that DatoCMS supports👇
Adlam (Fula, Pular, Fulfulde)
Armenian
Balinese
Bamum
Batak (Toba, Karo, Mandailing)
Canadian Syllabics (Cree, Inuktitut, Ojibwe)
Chakma
Cham
Cherokee
Ethiopic (Amharic, Tigrinya, Tigre, Ge'ez)
Georgian (Georgian, Mingrelian, Svan)
Hanifi Rohingya (Rohingya)
Javanese
Kayah Li (Red Karen)
Khmer
Lao
Lepcha
Limbu
Lisu
Mandaic
Meetei Mayek (Manipuri)
Miao (A-Hmao, other Miao languages)
Myanmar (Burmese, S'gaw Karen, Mon, Shan)
N'Ko (Bambara, Mandinka, Maninka)
New Tai Lue (Tai Lü)
Newa (Newar)
Nyiakeng Puachue Hmong (White Hmong, Green Hmong)
Ol Chiki (Santali)
Osage
Saurashtra
Sinhala (Sinhala, Pali)
Sundanese
Syloti Nagri (Sylheti)
Syriac (Classical Syriac, Assyrian, Turoyo)
Tai Le (Tai Nüa)
Tai Tham (Northern Thai, Tai Khün, Tai Lü)
Tai Viet (Tai Dam, Tai Dón)
Thaana (Dhivehi)
Tibetan (Tibetan, Dzongkha, Ladakhi)
Tifinagh (Tashelhit, Central Atlas Tamazight)
Vai
Wancho
Yi (Nuosu)