NLP for Malay (macrolanguage)

Measured on OPUS, spaCy, Stanza, Tatoeba, NLLB-200, FLORES-200, Belebele.

Observatories covering Malay (macrolanguage)

7 observatories measure NLP. Cards linking straight to this selection have data for it.

735,865,784 sentence pairs

21 corpora

5 pipeline modules

of 6 possible

No data for this selection

Per-language pipeline resource listing (tokenize/mwt/pos/lemma/depparse/ner/sentiment/constituency/coref) -- the cleanest pipeline-capability source.

No data for this selection

Crowdsourced sentence and translation database; counts here are sentence, translation-link, and audio-recording totals per language, from Tatoeba's own exports.

No data for this selection

Meta's No Language Left Behind translation models covering ~200 languages (FLORES-aligned).

FLORES-200

Open index →

No data for this selection

Evaluation benchmark for ~200 languages — the de facto baseline for whether a language is evaluable at all.

Yes in belebele