Bulbul بلبل
A Dataset for Dialectal Arabic Speech Recognition
Countries
Dialects
Speakers
Hours of Speech
ASR Models Tested
Arabic Isn't One Language
Most Arabic speech recognition benchmarks treat "Arabic" as a single language. It isn't. A speaker from Morocco and a speaker from Yemen can sound almost unintelligible to each other, yet ASR systems are still mostly evaluated as if dialect doesn't exist.
Bulbul is our attempt to close that gap: a benchmark built from real people, reading real sentences, the way they actually talk at home.
Same meaning, different Arabic
"My card still hasn't arrived."
البطاقة بتاعتي لسه موصلتش لحد دلوقتي
البطاقة ديالي ما وصلتش لحد دابا
Same idea, barely any shared words. Now picture a model trained mostly on news broadcasts trying to transcribe both.
11 Countries, 16 Dialects
Every country had its own team that collected text in its dialect, recorded it, and had each clip checked by a native speaker. Saudi Arabia and Yemen go one level deeper, into sub-dialects. Pick a country to see what it sounds like on paper.

فك الباب أبغى أكلمك شوي
Meaning: Open the door, I want to talk to you for a bit.
Sub-dialects
- Hijazi10.73h
- Eastern (Qatifi + Hassawi)1.01h
- Najdi0.85h
- Southern0.56h
- Hijazi-Bedouin0.49h
How a Dataset Gets Built
16 months, 11 country teams, a Slack group, weekly check-ins, and a lot of people reading sentences into microphones.
- 1
Text Collection
My partDialect text gathered per country from social media, donated WhatsApp and Telegram chats, and public corpora like SADA, MADAR, and Casablanca. Classical and MSA sentences came from the Tarjamat dataset.
- 2
Review & Preprocessing
Every sentence was reviewed by hand. Offensive, politically sensitive, or off-dialect text was removed, then normalized before upload.
- 3
Recording
My partParticipants recorded in a custom Gradio app, with sentences organized by country and sub-dialect, plus a training video and written guidelines.
- 4
Two-Level Verification
My partSpeakers could replay and check their own clip before submitting. Then a native speaker from the same country reviewed it for clarity, text match, and dialect authenticity.
- 5
Domain Classification
GPT-5-mini tagged each sentence with one of 11 topics, from Social to Sports. Anything below 0.70 confidence was reviewed by a human.
- 6
Benchmarking
My partSpeaker-disjoint dev and test splits, then zero-shot evaluation of 12 ASR models on both dialectal and accented formal speech.
Saudi speakers joined through the National Volunteer Platform. Every 10 minutes of recorded audio earned them 1 hour of official volunteer work.
My Role
Bulbul was a team of 35 researchers. I worked on three pieces of it, from the first sentence collected to the final leaderboard.
Saudi Dialect Data
Helped build the Saudi subset, one of the largest in Bulbul: collecting dialect text across Hijazi, Najdi, Southern, Hijazi-Bedouin, and Eastern varieties, onboarding volunteers, and verifying recordings.
The Recording Tool
Worked on the Gradio app every participant recorded in: sentences grouped by country and sub-dialect, speaker metadata, and a self-check step before each clip is submitted.
Whisper Benchmarking
Ran the zero-shot evaluation of OpenAI's Whisper models (Large-V3 and Large-V3-Turbo) on Bulbul's dialectal and accented test sets, scoring WER and CER through the shared Arabic normalization pipeline so they compare fairly against every other model.
The Leaderboard
12 models, tested zero-shot: no fine-tuning, default decoding, straight out of the box. The number is word error rate (WER), so lower is better.
Even the best model, OmniASR-LLM-7B, makes about 3x more word errors on dialects than on formal Arabic spoken by the same communities.
What We Learned
Some dialects are much harder
Palestinian and Saudi speech got the lowest error rates. Algerian and Sudanese were the hardest, thanks to scarce public data and a big distance from MSA.
Bigger isn't always better
LLM-decoder models improve steadily as they grow. CTC models don't: OmniASR-CTC-1B beats its own 3B and 7B siblings.
Egyptians talk fastest
Egyptian speakers averaged 1.94 words per second, Yemeni speakers 1.09. Tempo alone changes how hard alignment is for a model.
Where models trip
Where ق becomes a glottal stop, models flip between قديش and أديش. Emphatic ص gets confused with plain س, and the ب- prefix shows up as both بضل and بيضل.
Read the Paper
Built with 35 researchers from 12 institutions, led by Dr. Hamzah Luqman (PI) at KFUPM, with Ahmed Ashraf and Aisha Alansari. And none of it exists without the 275 volunteers who lent us their voices.