Technology . Souk Weekly
Arabic-First AI and the Politics of the Language Model
Building a model that thinks in Arabic first is technical, cultural, and political all at once
Updated

The world's most capable language models once learned Arabic like a clever tourist might: well enough to charm but missing the subtle nuances that make a language truly native. These models translated rather than thought, defaulting to English logic wrapped in Arabic words. Now, a new wave of regional effort aims to change this, building systems that reason in Arabic from the ground up.
A Language That Is Many Languages
Arabic is not uniform; it's a spectrum. There’s Modern Standard Arabic (MSA), the formal register used in news and scripture, which everyone reads but few speak at home. Around MSA orbit various dialects, Levantine, Gulf, Egyptian, Maghrebi, each with its own vocabulary and rhythm, often incomprehensible to speakers of other dialects. A language model trained primarily on formal text speaks like a newsreader at a wedding: it misses the warmth and nuance of everyday conversation.
To truly feel native, a model must encompass this entire spectrum. But capturing that complexity is far more challenging than compiling a neat dataset. It requires understanding the subtle shades of meaning that exist between dialects and registers, which is why many models fall short when trying to capture Arabic's rich tapestry.
Whose Arabic Gets to Be the Default?
Every decision in training a language model carries cultural weight. If a model learns predominantly from one country’s media, it absorbs that nation’s idioms, assumptions, and even its silences. The dialect it finds easiest becomes, by default, the machine age's preferred dialect. Communities whose speech is underrepresented online risk being seen as errors to correct. Thus, choosing which Arabic a model speaks is also about deciding who gets heard and who might be quietly erased.
The Thin and Contested Corpus
Compared with English, high-quality Arabic text on the open internet is sparse, much of it translated rather than originally composed. This scarcity tempts builders to rely heavily on machine-translated data, which embeds English sentence structures into what should be a purely Arabic mind. The harder path involves commissioning, digitizing, and cleaning original material, much of it sitting in archives, libraries, or the memories of those never asked for their stories. Data here is not discovered; it's crafted.
Sovereignty in a Server Rack
There’s also a strategic dimension to this effort. A region that relies entirely on foreign models for its administration, education, and media outsources part of its cognition to companies beyond its governance. Building capable local models asserts that a society should understand itself through its own language without intermediaries deciding what can be said. This is why these projects attract not just engineers but also government ministries.
The Risk of the Official Voice
But sovereignty comes with risks. A model shaped under close state interest might become fluent and obedient, eloquent on safe topics but vague when it comes to contentious ones. The danger isn’t that Arabic-first AI fails to sound native; it’s that it sounds too native while quietly narrowing what a native voice can say.
Building a model that dreams in Arabic is a worthy and overdue project, deserving respect for the engineering alone. Yet language transcends mere technology, it carries memory, hierarchy, and dispute. Whoever builds this machine inherits all of these elements, and the most honest builders will admit they are not just modeling Arabic; they’re making an argument about what it means to be Arabic.
Language is more than a tool; it’s a mirror reflecting society's complexities. As we strive to create AI that truly understands Arabic, let us remember that this endeavor is also a reflection of our own values and aspirations.
The Weekly
One email a week.
The good stuff, the strange stuff, the souk stuff.