Overview of Featured Indic NLP Resources
masterThe catalog highlights several key trends and resources in Indian language NLP:
- Universal Language Contribution API (ULCA): A standard API and open scalable data platform (part of the Bhasini mission) for discovering and uploading Indian language datasets and models.
- Large-scale Datasets: Examples include IndicCorp (9B tokens), Samanantar (50M parallel sentence pairs), Naamapadam (5.7M NER sentences), HiNER (100k NER sentences), and Aksharantar (26M transliteration pairs).
- High Language Coverage: Resources like Aksharantar (21 languages), FLORES-200 (27 languages), and IndoWordNet (18 languages) are expanding coverage across most constitutional Indian languages.
- Low-resource Support: Increasing focus on languages like Bodo, Kangri, and Khasi.
- Key Contributors: Major groups include AI4Bharat, BUET CSE NLP, KMI, L3Cube, iNLTK, and IIT Patna.