5 data sources · updated Aug 5, 21:09 UTC
Language Resources
Lexical, textual, and archival data on the world's languages, side by side across every source that tracks it.
Start from a language to see how every data source in this category covers it, side by side.
6,392 languages have data in at least one data source.
PanLex
6,241 langCrowdsourced lexical translation database aggregating thousands of dictionaries; counts here are meaning–expression entries per language, from the "meanings" export.
- Languages
- 6,241
- Records
- 79,088,196
PARADISEC
1,388 langPacific and Regional Archive for Digital Sources in Endangered Cultures; counts here are catalogued items and collections tagged with each language, from PARADISEC's own OLAC/OAI-PMH feed.
- Languages
- 1,388
- Items
- 51,204
ELAR
55 langEndangered Languages Archive collections catalogue; counts here are collections tagged with each language, from ELAR's own "imdi.language" facet.
- Languages
- 55
- Collections tagged
- 99
Tatoeba
423 langCrowdsourced sentence and translation database; counts here are sentence, translation-link, and audio-recording totals per language, from Tatoeba's own exports.
- Languages
- 423
- Sentences
- 12,829,057
Common Voice
333 langMozilla's crowdsourced speech dataset; counts here are recorded and validated speech hours per language, across its scripted-speech and spontaneous-speech corpora.
- Languages
- 333
- Hours (scripted)
- 42,376