datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
raw_maithili_audio_data_newmaithili_syspin_female_tts_22050
Maithili TTS Dataset (IISc SYSPIN Female)
This is a Maithili female TTS dataset from the IISc SYSPIN project.
It has been converted to 22050 Hz (mono) for seamless use in TTS fine-tuning, following the same schema as Firoj112/nepali_openslr43_tts_22050.
Dataset Summary
Language: Maithili (mai)
Speaker: Spk0001 (Female)
Total Duration: ~59 hours 40 mins
Total Utterances: 34,412
Sampling Rate: 22050 Hz (Resampled from 48kHz)
Format: Mono channel, float32 PCM… See the full description on the dataset page: https://huggingface.co/datasets/Firoj112/maithili_syspin_female_tts_22050.Synthetic-Multispeaker-Maithili-SantaliMaithili_Sentiment_8K
Maithili 64K Dataset
🌾 Maithili Multi-Dimensional Sentiment Corpus
📌 Executive Summary
Standard sentiment analysis in Indian vernaculars relies on flat, one-dimensional classification. The Maithili Multi-Dimensional Sentiment Corpus (64,215 rows) introduces a high-resolution, socio-linguistically grounded architecture. It utilizes a dual-axis classification system—predicting both primary sentiment and emotional intensity—mapped across highly… See the full description on the dataset page: https://huggingface.co/datasets/abhiprd2000/Maithili_Sentiment_8K.maithili-instruction-tuningadaption-digital-payments-and-banking-terms-and-topics-hindi-marathi-bhojpuri-maithili
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-digital payments and Banking terms and topics- Hindi, Marathi, Bhojpuri, Maithili
This dataset contains question-and-answer pairs focused on personal finance and banking services in India, covering topics like UPI, net banking, tax payments, and government loan schemes. Each sample includes a user query followed by a detailed, step-by-step completion that provides actionable advice… See the full description on the dataset page: https://huggingface.co/datasets/sidddd625/adaption-digital-payments-and-banking-terms-and-topics-hindi-marathi-bhojpuri-maithili.Maithili-Corpus
Maithili Raw Corpus
Language: Maithili (मैथिली, ISO 639-3: mai)
License: CC-BY-4.
Size: 28,622 documents | 13.7M words | ~45.8M tokens
Format: JSONL (one paragraph per row)
Tags: unlabelled, low-resource, indic-nlp, monolingual, pretraining
Dataset Summary
A large, unlabelled corpus of written Maithili text for language model pretraining and unsupervised NLP research.
Property
Value
Documents
28,622
Total Words
13,664,375
Total Subword Tokens… See the full description on the dataset page: https://huggingface.co/datasets/kamal-018/Maithili-Corpus.maithili-forward-translation-datasetalpaca_maithili_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_maithili_taco.maithiliNewsDataFor access do mail on rockerritesh4@gmail.com or find us here https://x.com/Rocker_Ritesh
@misc{yadav2025maibertspeakmaithili,
title={Can maiBERT Speak for Maithili?},
author={Sumit Yadav and Raju Kumar Yadav and Utsav Maskey and Gautam Siddharth Kashyap Md Azizul Hoque and Ganesh Gautam},
year={2025},
eprint={2509.15048},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2509.15048},
}
Maithili_Poems
Maithili Poetry Dataset
Dataset Summary
This dataset is a curated corpus of Maithili poetry collected from multiple online repositories and digitized literary sources. It is normalized and structured for language modeling, tokenization experiments, and generative poetry tasks in Maithili.
Dataset Statistics
Metric
Value
Language
Maithili (mai)
File Size
~1.97 MB (Uncompressed)
Token Count
~0.45 Million
Word Count
~288,500
Line… See the full description on the dataset page: https://huggingface.co/datasets/kamal-018/Maithili_Poems.alpaca-maithili-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-maithili-cleaned.maithili_legal_datasetmaithili_poemdata_maithili_batch_v1_01.json.1maithili_Agrade_reasoning_v1_03raw_maithili_audio_dataIndia_Maithili_Speech_Recognition_Corpus
ID
King-ASR-664
Language
India Maithili
Duration
200 hours
Speakers
400 People
Parameters
16kHz, 16bits
Recording Device
Mobile
URL
https://dataoceanai.com/datasets/asr/india-maithili-speech-recognition-corpus-mobile/
maithili_datadata_maithili_batch_1.jsondata_maithili_batch_3.jsonmaithili_reasoning_batch_A1data_maithili_batch_A3.jsondata_maithili_Agrade_v1_02.jsondata_maithili_batch_2.jsondata_maithili_batch_v2_01.jsonmaithili_Agrade_reasoning_v1_02data_maithili_batch_A2.json
