data source
Unicode CLDR (segmentation)
Open Unicode CLDR (segmentation)Canonical language-specific word/line/sentence break rules -- the reference point for tokenization support.
tokenization
Also on IDLI
Unicode CLDR (segmentation) has a dedicated observatory on IDLI — the fuller view of this source.
Languages with an override
10
Max break types
4
Languages in Unicode CLDR (segmentation)
10 of 10 shown
Community portal
Opportunities, events, and experts across the digital language inclusion ecosystem.