13 data sources · updated Aug 5, 16:09 UTC
Language Processing
How well a language is served by the open-source NLP ecosystem — tokenizers, morphological analyzers, language identifiers, and transliteration tools — one card per source, per language.
Start from a language to see how every data source in this category covers it, side by side.
8,314 languages are searchable; 2,129 have data in at least one data source so far.
Unicode CLDR (segmentation)
10 langCanonical language-specific word/line/sentence break rules -- the reference point for tokenization support.
- Languages with an override
- 10
- Max break types
- 4
spaCy
78 lang85 language directories; presence of tokenizer_exceptions and a lemmatizer distinguishes a real pipeline from a stub.
- Languages with a directory
- 78
- Max pipeline modules
- 6
Stanza
88 langPer-language pipeline resource listing (tokenize/mwt/pos/lemma/depparse/ner/sentiment/constituency/coref) -- the cleanest pipeline-capability source.
- Languages with a pipeline
- 89
- Max processors
- 9
UniMorph
185 lang197 public repos, one per ISO 639-3 code, of inflectional paradigms.
- Languages with a repo
- 185
- Raw repos scanned
- 197
Universal Dependencies
174 lang477 repos named UD_<Language>-<Treebank>; token counts and annotation layers.
- Languages with a treebank
- 174
- Total treebanks
- 463
GiellaLT
150 lang471 repos of FST morphology, spellers, and hyphenators for Sami, Uralic, and Indigenous North American languages -- best coverage of languages absent from mainstream NLP.
- Languages with tooling
- 150
- Raw repos scanned
- 471
GlotLID
1,977 lang~2,000 language labels; the widest language-identification coverage available.
- Languages identifiable
- 2,150
- Raw tags scanned
- 2,164
fastText LID (NLLB)
211 lang217 FLORES-200 labels from Meta's NLLB language-identification model.
- Languages identifiable
- 211
- Total script labels
- 218
CLD3
102 lang~107 languages; the language identifier embedded in Chrome.
- Distinct languages
- 102
- Raw entries scanned
- 109
franc
385 lang~400 languages, JavaScript; useful as a breadth datapoint rather than a quality signal.
- Languages identifiable (franc-all)
- 385
- franc-min tier
- 61
Unicode CLDR (transforms)
32 lang381 transform files; the canonical transliteration registry, shipped downstream as ICU transliterator IDs.
- Languages with a transform
- 32
- Canonical-form files
- 185
Aksharamukha
97 lang~120 scripts; strongest coverage for Brahmic and Southeast Asian script conversion.
- Languages covered
- 97
- Scripts scanned
- 102
uroman
0 langScript-agnostic romanization covering any Unicode script -- a script-level, not language-level, capability.
- Coverage
- Universal
- Languages with a dedicated row
- 0