← All data sources← All Language Resources observatories on IDLI

data source

Common Voice

Mozilla's crowdsourced speech dataset; counts here are recorded and validated speech hours per language, across its scripted-speech and spontaneous-speech corpora.

speechcrowdsourced

Also on IDLI

Common Voice has a dedicated observatory on IDLI — the fuller view of this source.

Open the Common Voice observatory
Languages
333
Hours (scripted)
42,376
recorded, across every language
Hours (spontaneous)
515
Updated
Aug 5, 2026

Languages in Common Voice

10 of 333 shown

Detail
Eastern Mari
mhr
301500281detail →
Hungarian
hun
1811,679137detail →
Western Mari
mrj
376034detail →
Bulgarian
bul
2215217detail →
Pahari-Potwari
phr
146314detail →
Parkari Koli
kvx
112211detail →
Gujari
gju
11710detail →
Marwari (Pakistan)
mve
102010detail →
Goaria
gig
102010detail →
Amharic
amh
3512detail →

Community portal

Opportunities, events, and experts across the digital language inclusion ecosystem.