← All data sources← All Language Resources observatories on IDLI

data source

Common Voice

Mozilla's crowdsourced speech dataset; counts here are recorded and validated speech hours per language, across its scripted-speech and spontaneous-speech corpora.

speechcrowdsourced

Also on IDLI

Common Voice has a dedicated observatory on IDLI — the fuller view of this source.

Open the Common Voice observatory
Languages
333
Hours (scripted)
42,376
recorded, across every language
Hours (spontaneous)
515
Updated
Aug 5, 2026

Languages in Common Voice

13 of 333 shown

Detail
Chinese
zho
1,2069,911319detail →
Japanese
jpn
7267,834372detail →
Yue Chinese
yue
4514,296319detail →
Portuguese
por
2293,843188detail →
Min Nan Chinese
nan
2430022detail →
Vietnamese
vie
234178detail →
Aragonese
arg
1853172detail →
Maltese
mlt
182259detail →
Pu-Xian Chinese
cpx
1129detail →
Min Dong Chinese
cdo
1032detail →
Assamese
asm
8513detail →
Javanese
jav
13detail →
Sundanese
sun
12detail →

Community portal

Opportunities, events, and experts across the digital language inclusion ecosystem.