data source
Common Voice
Mozilla's crowdsourced speech dataset; counts here are recorded and validated speech hours per language, across its scripted-speech and spontaneous-speech corpora.
speechcrowdsourced
Also on IDLI
Common Voice has a dedicated observatory on IDLI — the fuller view of this source.
Languages
333
Hours (scripted)
42,376
recorded, across every language
Hours (spontaneous)
515
Updated
Aug 5, 2026
Languages in Common Voice
9 of 333 shown
| Detail | |||||
|---|---|---|---|---|---|
Swahili (macrolanguage) swa | 1,065 | 1,523 | 392 | — | detail → |
Oriya (macrolanguage) ori | 37 | 164 | 6 | — | detail → |
Batanga bnm | 16 | 21 | 15 | — | detail → |
Yangben yav | 10 | 10 | 9 | — | detail → |
Fang (Equatorial Guinea) fan | 10 | 43 | 9 | — | detail → |
Mundang mua | 10 | 17 | 10 | — | detail → |
Malay (macrolanguage) msa | 4 | 29 | 0 | 6 | detail → |
Nepali (macrolanguage) nep | 2 | 63 | 1 | — | detail → |
Lango (Uganda) laj | 0 | 2 | — | — | detail → |
Community portal
Opportunities, events, and experts across the digital language inclusion ecosystem.