data source
Common Voice
Mozilla's crowdsourced speech dataset; counts here are recorded and validated speech hours per language, across its scripted-speech and spontaneous-speech corpora.
speechcrowdsourced
Also on IDLI
Common Voice has a dedicated observatory on IDLI — the fuller view of this source.
Languages
333
Hours (scripted)
42,376
recorded, across every language
Hours (spontaneous)
515
Updated
Aug 5, 2026
Languages in Common Voice
13 of 333 shown
| Detail | |||||
|---|---|---|---|---|---|
Ukrainian ukr | 118 | 1,177 | 103 | — | detail → |
Luo (Kenya and Tanzania) luo | 112 | 47 | 28 | — | detail → |
Estonian est | 67 | 1,069 | 52 | — | detail → |
Armenian hye | 62 | 1,041 | 39 | — | detail → |
Macedonian mkd | 55 | 468 | 26 | — | detail → |
Romanian ron | 52 | 464 | 25 | — | detail → |
Lithuanian lit | 36 | 333 | 30 | — | detail → |
Slovenian slv | 22 | 1,003 | 17 | — | detail → |
Gheg Albanian aln | 11 | 14 | — | — | detail → |
Albanian sqi | 9 | 155 | 9 | — | detail → |
Sardinian srd | 4 | 50 | 3 | — | detail → |
Arvanitika Albanian aat | 2 | 5 | — | — | detail → |
Macedo-Romanian rup | 0 | 9 | 0 | — | detail → |
Community portal
Opportunities, events, and experts across the digital language inclusion ecosystem.