CoolFace
Datasetpublic

cfilt/IITB-IndicMonoDoc

IITB Document level Monolingual Corpora for Indian languages. 22 scheduled languages of India + English (1) Assamese, (2) Bengali, (3) Gujarati, (4) Hindi, (5) Kannada, (6) Kashmiri, (7) Konkani, (8) Malayalam, (9) Manipuri, (10) Marathi, (11) Nepali, (12) Oriya, (13) Punjabi, (14) Sanskrit, (15) Sindhi, (16) Tamil, (17) Telugu, (18) Urdu (19) Bodo, (20) Santhali, (21) Maithili and (22) Dogri. Language Total (#Mil Tokens) bn 5258.47 en 11986.53 gu 887.18 hi 11268.33 kn… See the full description on the dataset page: https://huggingface.co/datasets/cfilt/IITB-IndicMonoDoc.

sourceHugging Facecc-by-4.0updated 2y agoView on Hugging Face
11likes25kdownloads
Dataset Card

IITB Document level Monolingual Corpora for Indian languages.

22 scheduled languages of India + English

(1) Assamese, (2) Bengali, (3) Gujarati, (4) Hindi, (5) Kannada, (6) Kashmiri, (7) Konkani, (8) Malayalam, (9) Manipuri, (10) Marathi, (11) Nepali, (12) Oriya, (13) Punjabi, (14) Sanskrit, (15) Sindhi, (16) Tamil, (17) Telugu, (18) Urdu (19) Bodo, (20) Santhali, (21) Maithili and (22) Dogri.

LanguageTotal (#Mil Tokens)
bn5258.47
en11986.53
gu887.18
hi11268.33
kn567.16
ml845.32
mr1066.76
ne1542.39
pa449.61
ta2171.92
te767.18
ur2391.79
as57.64
brx2.25
doi0.37
gom2.91
kas1.27
mai1.51
mni0.99
or81.96
sa80.09
sat3.05
sd83.81
Total=39518.51

To cite this dataset:

@inproceedings{doshi-etal-2024-pretraining,
    title = "Pretraining Language Models Using Translationese",
    author = "Doshi, Meet  and
      Dabre, Raj  and
      Bhattacharyya, Pushpak",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.334/",
    doi = "10.18653/v1/2024.emnlp-main.334",
    pages = "5843--5862",
}