← IDLIHow to read this data
← Back to all languages

How to read this data

What "sentence pairs" means

OPUS is a collection of parallel corpora: pairs of documents that are the same text in two different languages, aligned sentence by sentence. A "sentence pair" is one aligned pair of sentences between a language and a partner language in one of those documents. A language's total sentence-pair count is the sum of its aligned pairs across every corpus and every partner language it appears with in OPUS.

Raw counts reflect tagging, not real-world translation supply

These numbers describe what OPUS has indexed and how its source corpora tag their language codes — not how much genuine translation exists between two languages in the world, and not a judgment about corpus quality, licensing, or fitness for any particular task.

A few things worth knowing before reading too much into a count:

Languages without an ISO badge

Most rows show an ISO 639-3 code (e.g. "ISO cmn"). Some don't — these are languages this observatory couldn't confidently resolve to a specific ISO 639-3 code, usually for one of two reasons:

Rather than guess, these are left without an ISO code so the uncertainty is visible instead of hidden.

Snapshots and history

Each time the data pipeline runs, it writes a new dated snapshot rather than overwriting the last one. OPUS corpora don't update often enough for week-to-week change figures to be meaningful, so this site doesn't compute or display them. What each language's page does show is a history chart of its sentence-pair count across every snapshot collected — straight lines between the real data points, with each point clearly marked, so it reads honestly as sparse periodic measurements rather than implying continuous change. A language with only one snapshot on record won't have a chart yet; that's expected, not missing data.

Source

All underlying data comes from OPUS (opus.nlpl.eu), the open parallel corpus collection maintained by the University of Helsinki. OPUS is the authoritative source for the corpora, resource counts, and language codes this observatory aggregates — this site adds no data of its own beyond that aggregation.