data source
Common Voice
Mozilla's crowdsourced speech dataset; counts here are recorded and validated speech hours per language, across its scripted-speech and spontaneous-speech corpora.
speechcrowdsourced
Also on IDLI
Common Voice has a dedicated observatory on IDLI — the fuller view of this source.
Languages
333
Hours (scripted)
42,376
recorded, across every language
Hours (spontaneous)
515
Updated
Aug 5, 2026
Languages in Common Voice
5 of 333 shown
| Detail | |||||
|---|---|---|---|---|---|
Belarusian bel | 1,893 | 8,621 | 1,818 | — | detail → |
Persian fas | 431 | 4,660 | 373 | — | detail → |
Russian rus | 292 | 3,744 | 254 | 3 | detail → |
Western Frisian fry | 215 | 2,105 | 70 | 0 | detail → |
Indonesian ind | 67 | 671 | 34 | — | detail → |
Community portal
Opportunities, events, and experts across the digital language inclusion ecosystem.