← All data sources← All Language Processing observatories on IDLI

data source

85 language directories; presence of tokenizer_exceptions and a lemmatizer distinguishes a real pipeline from a stub.

tokenizationmorphological-analysis
Languages with a directory
78
Max pipeline modules
6
Unmatched codes
1
not in the crosswalk

Languages in spaCy

1 of 78 shown

Detail
Arabic
ara
4detail →

Community portal

Opportunities, events, and experts across the digital language inclusion ecosystem.