← IDLI

Unevenly Digitized Languages

Filter languages by low or high coverage in any combination of sources.

Combine None, Any, Low, or High across any of the sources below; results match all selected filters. None and Any test whether a source has an entry; Low and High test quantity, defined per source.

Filter thresholds are currently in beta and are intended as preliminary reference points rather than definitive high/low classifications. Feedback is welcome.

CLDR

Low: Unlabeled or basic tier

High: Modern tier

Wikipedia

Low: Fewer than 5,000 articles (or Incubator-only, no full edition)

High: 25,000 or more articles

Keyman (keyboards)

Low: Fewer than 10,000 downloads

High: 50,000 or more downloads

OPUS (NLP datasets)

Low: 1 or fewer corpora

High: 4 or more corpora

Common Voice (speech)

Low: Fewer than 10 hours recorded

High: 30 or more hours recorded

Google Fonts

Low: 4 or fewer font families

High: 6 or more font families

SIL Fonts

Low: 1 recommended font family

High: 3 or more recommended font families

Wikimedia jQuery.IME

Low: Fewer than 3 input methods

High: 5 or more input methods

LibreOffice Localization

Low: Less than 35% translated

High: 70% or more translated

Glottolog (AES)

Low: Moribund, nearly extinct, or extinct

High: Not endangered

WALS

Low: 10 or fewer features documented

High: 36 or more features documented

PanLex

Low: Fewer than 40 meanings

High: 200 or more meanings

PARADISEC

Low: Fewer than 3 items

High: 8 or more items

Tatoeba

Low: Fewer than 50 sentences

High: 1,000 or more sentences

Tesseract (OCR)
ICANN IDN
LanguageTool
HarfBuzz (text shaping)
eSpeak NG dictsource
NVDA locale
AOSP LatinIME
translatewiki.net
spaCy
Stanza
UniMorph
Universal Dependencies
GiellaLT
GlotLID
fastText LID (NLLB)
CLD3
franc
ELAR
Piper TTS
Machine Translate Foundation
Apertium
Hunspell

No filters set yet. Choose None, Any, Low, or High on at least one source, or a region or country, above.