cfilt/IITB-IndicMonoDoc
IITB Document level Monolingual Corpora for Indian languages. 22 scheduled languages of India + English (1) Assamese, (2) Bengali, (3) Gujarati, (4) Hindi, (5) Kannada, (6) Kashmiri, (7) Konkani, (8) Malayalam, (9) Manipuri, (10) Marathi, (11) Nepali, (12) Oriya, (13) Punjabi, (14) Sanskrit, (15) Sindhi, (16) Tamil, (17) Telugu, (18) Urdu (19) Bodo, (20) Santhali, (21) Maithili and (22) Dogri. Language Total (#Mil Tokens) bn 5258.47 en 11986.53 gu 887.18 hi 11268.33 kn… See the full description on the dataset page: https://huggingface.co/datasets/cfilt/IITB-IndicMonoDoc.
IITB Document level Monolingual Corpora for Indian languages.
22 scheduled languages of India + English
(1) Assamese, (2) Bengali, (3) Gujarati, (4) Hindi, (5) Kannada, (6) Kashmiri, (7) Konkani, (8) Malayalam, (9) Manipuri, (10) Marathi, (11) Nepali, (12) Oriya, (13) Punjabi, (14) Sanskrit, (15) Sindhi, (16) Tamil, (17) Telugu, (18) Urdu (19) Bodo, (20) Santhali, (21) Maithili and (22) Dogri.
To cite this dataset:
@inproceedings{doshi-etal-2024-pretraining,
title = "Pretraining Language Models Using Translationese",
author = "Doshi, Meet and
Dabre, Raj and
Bhattacharyya, Pushpak",
editor = "Al-Onaizan, Yaser and
Bansal, Mohit and
Chen, Yun-Nung",
booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
month = nov,
year = "2024",
address = "Miami, Florida, USA",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2024.emnlp-main.334/",
doi = "10.18653/v1/2024.emnlp-main.334",
pages = "5843--5862",
}