CoolFace
Datasetpublic

yash-ingle/ILID_Indian_Language_Identification_Dataset

ILID: Native Script Language Identification for Indian Languages Paper | Code | Project Page πŸ—£ ILID: Indian Language Identification Dataset (23 Languages)Authors: Yash Ingle, Dr. Pruthwik MishraInstitute: Sardar Vallabhbhai National Institute of Technology (SVNIT), Surat, India πŸ“„ Dataset Description The ILID (Indian Language Identification Dataset) benchmark contains 250,000 sentences from English and 22 official Indian languages, designed for training and… See the full description on the dataset page: https://huggingface.co/datasets/yash-ingle/ILID_Indian_Language_Identification_Dataset.

sourceHugging Facemitupdated 9mo agoView on Hugging Face
0likes98downloads
Dataset Card

ILID: Native Script Language Identification for Indian Languages

**Paper** | **Code** | **Project Page**

πŸ—£ ILID: Indian Language Identification Dataset (23 Languages) Authors: Yash Ingle, Dr. Pruthwik Mishra Institute: Sardar Vallabhbhai National Institute of Technology (SVNIT), Surat, India


πŸ“„ Dataset Description

The ILID (Indian Language Identification Dataset) benchmark contains 250,000 sentences from English and 22 official Indian languages, designed for training and evaluating language identification models. The dataset supports the task of distinguishing between Indian languages, many of which share scripts, vocabulary, and structure.


πŸ“Š Dataset Statistics

LanguageCodeTrainDevTestTotal
Assameseasm80001000100010000
Bengaliben80001000100010000
Bodobrx80001000100010000
Dogridoi80001000100010000
Konkanigom80001000100010000
Gujaratiguj80001000100010000
Hindihin80001000100010000
Kannadakan80001000100010000
Kashmirikas80001000100010000
Maithilimai80001000100010000
Malayalammal80001000100010000
Marathimar80001000100010000
Manipuri (Bengali)mni_Beng80001000100010000
Manipuri (Meitei)mni_Mtei80001000100010000
Nepalinpi80001000100010000
Odiaory80001000100010000
Punjabipan80001000100010000
Sanskritsan80001000100010000
Santalisat80001000100010000
Sindhi (Arabic)snd_Arab80001000100010000
Sindhi (Devanagari)snd_Deva80001000100010000
Tamiltam80001000100010000
Telugutel80001000100010000
Urduurd80001000100010000
Englisheng80001000100010000
Totalβ€”2000002500025000250000

πŸ“ Files Provided

  • β€”shuffled_train_sentences: Training sentences (80% split – 200,000 samples)
  • β€”shuffled_train_labels: Corresponding labels for training sentences
  • β€”shuffled_dev_sentences: Validation (dev) sentences (10% split – 25,000 samples)
  • β€”shuffled_dev_labels: Corresponding labels for dev sentences
  • β€”shuffled_test_sentences: Test sentences (10% split – 25,000 samples)
  • β€”shuffled_test_labels: Corresponding labels for test sentences

πŸ“Œ Tasks

  • β€”Language Identification (LID)
  • β€”Multilingual Text Classification
  • β€”Benchmarking ML & DL Models on Indian Languages

🧹 Data Collection & Cleaning

  • β€”13 languages collected using web scraping from Wikipedia, news portals, and blogs.
  • β€”10 languages sampled from large monolingual corpora (Bhashaverse).
  • β€”Each sentence underwent cleaning, normalization, and language filtering via FastText.

🧠 Models & Results

Baseline models include:

  • β€”TF-IDF + Machine Learning: SVM, Logistic Regression, Random Forest, etc.
  • β€”FastText Classifier
  • β€”Fine-tuned MuRIL (BERT for Indian languages)

Best ensemble models achieve F1-scores of up to 0.99 on test/dev sets.


πŸ“š Citation

If you use this dataset, please cite:

bibtex
@misc{ingle2025ilidnativescriptlanguage,
      title={ILID: Native Script Language Identification for Indian Languages}, 
      author={Yash Ingle and Pruthwik Mishra},
      year={2025},
      eprint={2507.11832},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2507.11832}, 
}