Observatories
7 observatories measure NLP.
OPUS
Open index →spaCy
Open index →78 languages
85 language directories; presence of tokenizer_exceptions and a lemmatizer distinguishes a real pipeline from a stub.
Stanza
Open index →88 languages
Per-language pipeline resource listing (tokenize/mwt/pos/lemma/depparse/ner/sentiment/constituency/coref) -- the cleanest pipeline-capability source.
Tatoeba
Open index →423 languages
Crowdsourced sentence and translation database; counts here are sentence, translation-link, and audio-recording totals per language, from Tatoeba's own exports.
NLLB-200
Open index →196 languages
Meta's No Language Left Behind translation models covering ~200 languages (FLORES-aligned).
FLORES-200
Open index →194 languages
Evaluation benchmark for ~200 languages — the de facto baseline for whether a language is evaluable at all.
Belebele
Open index →96 languages
Reading comprehension across 100+ language variants. Presence means evaluable, not merely trainable.