data source
Common Voice
Mozilla's crowdsourced speech dataset; counts here are recorded and validated speech hours per language, across its scripted-speech and spontaneous-speech corpora.
speechcrowdsourced
Also on IDLI
Common Voice has a dedicated observatory on IDLI — the fuller view of this source.
Languages
333
Hours (scripted)
42,376
recorded, across every language
Hours (spontaneous)
515
Updated
Aug 5, 2026
Languages in Common Voice
7 of 333 shown
| Detail | |||||
|---|---|---|---|---|---|
Luo (Kenya and Tanzania) luo | 112 | 47 | 28 | — | detail → |
Indus Kohistani mvy | 25 | 57 | 22 | — | detail → |
Kohistani Shina plk | 17 | 10 | 13 | — | detail → |
Batanga bnm | 16 | 21 | 15 | — | detail → |
Occitan (post 1500) oci | 13 | 149 | 3 | — | detail → |
Marwari (Pakistan) mve | 10 | 20 | 10 | — | detail → |
Standard Moroccan Tamazight zgh | 2 | 51 | 1 | — | detail → |
Community portal
Opportunities, events, and experts across the digital language inclusion ecosystem.