Brewing ideas, coding intelligence

0%
Accepted · EMNLP 2026 Main Conference

Bulbul بلبل

A Dataset for Dialectal Arabic Speech Recognition

0

Countries

0

Dialects

0

Speakers

0

Hours of Speech

0

ASR Models Tested

PythonPyTorchHugging FaceGradioSpeech RecognitionArabic NLP

Arabic Isn't One Language

Most Arabic speech recognition benchmarks treat "Arabic" as a single language. It isn't. A speaker from Morocco and a speaker from Yemen can sound almost unintelligible to each other, yet ASR systems are still mostly evaluated as if dialect doesn't exist.

Bulbul is our attempt to close that gap: a benchmark built from real people, reading real sentences, the way they actually talk at home.

Same meaning, different Arabic

"My card still hasn't arrived."

Egyptian

البطاقة بتاعتي لسه موصلتش لحد دلوقتي

Moroccan

البطاقة ديالي ما وصلتش لحد دابا

Same idea, barely any shared words. Now picture a model trained mostly on news broadcasts trying to transcribe both.

11 Countries, 16 Dialects

Every country had its own team that collected text in its dialect, recorded it, and had each clip checked by a native speaker. Saudi Arabia and Yemen go one level deeper, into sub-dialects. Pick a country to see what it sounds like on paper.

Map of the Arab world with a sample sentence from each of the 11 Bulbul countries

فك الباب أبغى أكلمك شوي

Meaning: Open the door, I want to talk to you for a bit.

11.89Hours of speech
1.75Words per second
19.6%Best model WER

Sub-dialects

  • Hijazi
    10.73h
  • Eastern (Qatifi + Hassawi)
    1.01h
  • Najdi
    0.85h
  • Southern
    0.56h
  • Hijazi-Bedouin
    0.49h

How a Dataset Gets Built

16 months, 11 country teams, a Slack group, weekly check-ins, and a lot of people reading sentences into microphones.

  1. 1

    Text Collection

    My part

    Dialect text gathered per country from social media, donated WhatsApp and Telegram chats, and public corpora like SADA, MADAR, and Casablanca. Classical and MSA sentences came from the Tarjamat dataset.

  2. 2

    Review & Preprocessing

    Every sentence was reviewed by hand. Offensive, politically sensitive, or off-dialect text was removed, then normalized before upload.

  3. 3

    Recording

    My part

    Participants recorded in a custom Gradio app, with sentences organized by country and sub-dialect, plus a training video and written guidelines.

  4. 4

    Two-Level Verification

    My part

    Speakers could replay and check their own clip before submitting. Then a native speaker from the same country reviewed it for clarity, text match, and dialect authenticity.

  5. 5

    Domain Classification

    GPT-5-mini tagged each sentence with one of 11 topics, from Social to Sports. Anything below 0.70 confidence was reviewed by a human.

  6. 6

    Benchmarking

    My part

    Speaker-disjoint dev and test splits, then zero-shot evaluation of 12 ASR models on both dialectal and accented formal speech.

Fun fact

Saudi speakers joined through the National Volunteer Platform. Every 10 minutes of recorded audio earned them 1 hour of official volunteer work.

My Role

Bulbul was a team of 35 researchers. I worked on three pieces of it, from the first sentence collected to the final leaderboard.

Saudi Dialect Data

Helped build the Saudi subset, one of the largest in Bulbul: collecting dialect text across Hijazi, Najdi, Southern, Hijazi-Bedouin, and Eastern varieties, onboarding volunteers, and verifying recordings.

The Recording Tool

Worked on the Gradio app every participant recorded in: sentences grouped by country and sub-dialect, speaker metadata, and a self-check step before each clip is submitted.

Whisper Benchmarking

Ran the zero-shot evaluation of OpenAI's Whisper models (Large-V3 and Large-V3-Turbo) on Bulbul's dialectal and accented test sets, scoring WER and CER through the shared Arabic normalization pipeline so they compare fairly against every other model.

The Leaderboard

12 models, tested zero-shot: no fine-tuning, default decoding, straight out of the box. The number is word error rate (WER), so lower is better.

41.7%WER on dialects
13.3%WER on formal Arabic

Even the best model, OmniASR-LLM-7B, makes about 3x more word errors on dialects than on formal Arabic spoken by the same communities.

OmniASR-LLM-7B
41.7%
OmniASR-LLM-3B
43.6%
OmniASR-LLM-1B
45.1%
OmniASR-CTC-1B
49.5%
SeamlessM4T
49.8%
OmniASR-LLM-300M
50.8%
OmniASR-CTC-7B
51.8%
OmniASR-CTC-3B
53.3%
WhisperV3
53.8%
WhisperV3-Turbo
63.2%
OmniASR-CTC-300M
65.7%
MMS-1B
77.1%

What We Learned

01

Some dialects are much harder

Palestinian and Saudi speech got the lowest error rates. Algerian and Sudanese were the hardest, thanks to scarce public data and a big distance from MSA.

02

Bigger isn't always better

LLM-decoder models improve steadily as they grow. CTC models don't: OmniASR-CTC-1B beats its own 3B and 7B siblings.

03

Egyptians talk fastest

Egyptian speakers averaged 1.94 words per second, Yemeni speakers 1.09. Tempo alone changes how hard alignment is for a model.

04

Where models trip

Where ق becomes a glottal stop, models flip between قديش and أديش. Emphatic ص gets confused with plain س, and the ب- prefix shows up as both بضل and بيضل.

Read the Paper

Built with 35 researchers from 12 institutions, led by Dr. Hamzah Luqman (PI) at KFUPM, with Ahmed Ashraf and Aisha Alansari. And none of it exists without the 275 volunteers who lent us their voices.