data source
Common Voice
Mozilla's crowdsourced speech dataset; counts here are recorded and validated speech hours per language, across its scripted-speech and spontaneous-speech corpora.
speechcrowdsourced
Also on IDLI
Common Voice has a dedicated observatory on IDLI — the fuller view of this source.
Languages
333
Hours (scripted)
42,376
recorded, across every language
Hours (spontaneous)
515
Updated
Aug 5, 2026
Languages in Common Voice
10 of 333 shown
| Detail | |||||
|---|---|---|---|---|---|
Eastern Mari mhr | 301 | 500 | 281 | — | detail → |
Hungarian hun | 181 | 1,679 | 137 | — | detail → |
Western Mari mrj | 37 | 60 | 34 | — | detail → |
Bulgarian bul | 22 | 152 | 17 | — | detail → |
Pahari-Potwari phr | 14 | 63 | 14 | — | detail → |
Parkari Koli kvx | 11 | 22 | 11 | — | detail → |
Gujari gju | 11 | 7 | 10 | — | detail → |
Marwari (Pakistan) mve | 10 | 20 | 10 | — | detail → |
Goaria gig | 10 | 20 | 10 | — | detail → |
Amharic amh | 3 | 51 | 2 | — | detail → |
Community portal
Opportunities, events, and experts across the digital language inclusion ecosystem.