data source
Common Voice
Mozilla's crowdsourced speech dataset; counts here are recorded and validated speech hours per language, across its scripted-speech and spontaneous-speech corpora.
speechcrowdsourced
Also on IDLI
Common Voice has a dedicated observatory on IDLI — the fuller view of this source.
Languages
333
Hours (scripted)
42,376
recorded, across every language
Hours (spontaneous)
515
Updated
Aug 5, 2026
Languages in Common Voice
7 of 333 shown
| Detail | |||||
|---|---|---|---|---|---|
Sakizaya szy | 14 | 26 | 14 | — | detail → |
Atayal tay | 12 | 19 | 11 | — | detail → |
Malayalam mal | 11 | 157 | 4 | — | detail → |
Sabah Bisaya bsy | 11 | 25 | — | — | detail → |
Northwest Gbaya gya | 10 | 31 | 10 | — | detail → |
South Ucayali Ashéninka cpy | 10 | 15 | 10 | — | detail → |
Ayacucho Quechua quy | 2 | 8 | 0 | — | detail → |
Community portal
Opportunities, events, and experts across the digital language inclusion ecosystem.