Falcon-Emirati-7B Can Actually Speak Emirati Arabic, Not Just Translate It
The Technology Innovation Institute has released a 7-billion-parameter model trained specifically on the Emirati dialect, scoring 84.83% on a purpose-built benchmark where models far larger fall short.

Key points
- Falcon-Emirati-7B scored 84.83% on Alyah, a 1,173-question benchmark built by native Emirati speakers, beating every other Arabic and multilingual model tested including ones many times its size.
- The model was built on top of Falcon-H1-Arabic, an existing Arabic model family, not created from scratch.
- Training used three data sources: real Emirati web content, formal Arabic material about Emirati culture, and synthetic text constrained by Emirati vocabulary glossaries.
- Emirati Arabic is primarily spoken with little written text online, making it one of the hardest Arabic varieties for AI to learn.
- This release follows a pattern of AI labs targeting dialect gaps that standard multilingual models leave wide open.
Arabic has a problem most people outside the Arab world don't know about. The formal version taught in schools and used in news broadcasts, called Modern Standard Arabic, is nobody's mother tongue. Real conversations happen in regional dialects: Egyptian, Levantine, and in the UAE, a Gulf variety called Emirati Arabic, with its own grammar, humour and poetry. A model that knows the formal version fluently can still read an Emirati sentence, translate every word correctly, and completely miss the point.
That's the gap the Technology Innovation Institute set out to close with Falcon-Emirati-7B, released this week on Hugging Face.
What makes this harder than it sounds?
Emirati Arabic is mostly spoken, not written. Very little of it appears as text online, which is exactly what AI models learn from. Nabati poetry, a classical oral tradition, and everyday proverbs carry meaning tied to shared cultural knowledge that a literal reading can't recover.
The team didn't build from scratch. Falcon-Emirati-7B sits on top of Falcon-H1-Arabic, a family of Arabic language models available in three sizes: 3B, 7B and 34B parameters, where a parameter is roughly a unit of learned knowledge. The 7B tier was chosen deliberately. It's large enough to hold cultural nuance but small enough to keep deployment practical. The 34B variant would likely push quality further, but the training and serving cost doesn't make sense for a dialect-focused chat model.
Training data came from three places. First, crawled Emirati web content written natively in the dialect. Second, formal-Arabic articles about Emirati customs, history and social norms, which teach the model what it's discussing even when they're not written in dialect. Third, synthetic text produced under strict vocabulary rules and dictionaries built specifically for Emirati grammar, covering topics the real-world crawl couldn't reach.
How does it compare to other models?
The benchmark is Alyah (الياه, meaning "North Star"), a multiple-choice test of 1,173 questions collected manually from native Emirati speakers. It covers everyday greetings, figurative language, heritage knowledge and poetry: the categories where generic Arabic models tend to perform worst.
| Model | Alyah score |
|---|---|
| Falcon-Emirati-7B | 84.83% |
| Leading multilingual competitors | Below 84.83% |
The institute didn't publish individual competitor scores in the summary available, but stated the 7B model outperformed every Arabic and multilingual model in the comparison, including several many times its size.
Scoring combined automatic metrics with review by native Emirati speakers, who judged naturalness, tone and cultural fit alongside factual correctness. That human check matters. A benchmark number alone can't tell you whether a sentence sounds like something a real Emirati would actually say.
Our 18 August story on non-English AI reasoning found that multilingual models are closing the gap on English in formal tasks, but dialect comprehension is a different problem entirely, one that aggregate benchmarks tend to obscure.
For Emirati users, the practical difference is whether an AI assistant understands a joke, gets the weight of a proverb, or catches a poetry reference without flattening it into formal Arabic. That's not a small thing.
What matters most here is that TII is betting dialect specificity beats scale. An 84.83% score from 7B parameters is a credible result. What to watch: whether independent native-speaker testing holds up outside the institute's own evaluation, and whether the training recipe scales to other under-resourced Arabic varieties.



