← All data sources← All Language Processing observatories on IDLI

data source

Unicode CLDR (segmentation)

Open Unicode CLDR (segmentation)

Canonical language-specific word/line/sentence break rules -- the reference point for tokenization support.

tokenization

Also on IDLI

Unicode CLDR (segmentation) has a dedicated observatory on IDLI — the fuller view of this source.

Open the Unicode CLDR (segmentation) observatory
Languages with an override
10
Max break types
4

Languages in Unicode CLDR (segmentation)

10 of 10 shown

Community portal

Opportunities, events, and experts across the digital language inclusion ecosystem.