maithili
Datasets
All datasets matching “maithili”raw_maithili_audio_data_newmaithili_syspin_female_tts_22050
Maithili TTS Dataset (IISc SYSPIN Female)
This is a Maithili female TTS dataset from the IISc SYSPIN project.
It has been converted to 22050 Hz (mono) for seamless use in TTS fine-tuning, following the same schema as Firoj112/nepali_openslr43_tts_22050.
Dataset Summary
Language: Maithili (mai)
Speaker: Spk0001 (Female)
Total Duration: ~59 hours 40 mins
Total Utterances: 34,412
Sampling Rate: 22050 Hz (Resampled from 48kHz)
Format: Mono channel, float32 PCM… See the full description on the dataset page: https://huggingface.co/datasets/Firoj112/maithili_syspin_female_tts_22050.Maithili_Sentiment_8K
Maithili 64K Dataset
🌾 Maithili Multi-Dimensional Sentiment Corpus
📌 Executive Summary
Standard sentiment analysis in Indian vernaculars relies on flat, one-dimensional classification. The Maithili Multi-Dimensional Sentiment Corpus (64,215 rows) introduces a high-resolution, socio-linguistically grounded architecture. It utilizes a dual-axis classification system—predicting both primary sentiment and emotional intensity—mapped across highly… See the full description on the dataset page: https://huggingface.co/datasets/abhiprd2000/Maithili_Sentiment_8K.Synthetic-Multispeaker-Maithili-Santalimaithili-instruction-tuningMaithili-Corpus
Maithili Raw Corpus
Language: Maithili (मैथिली, ISO 639-3: mai)
License: CC-BY-4.
Size: 28,622 documents | 13.7M words | ~45.8M tokens
Format: JSONL (one paragraph per row)
Tags: unlabelled, low-resource, indic-nlp, monolingual, pretraining
Dataset Summary
A large, unlabelled corpus of written Maithili text for language model pretraining and unsupervised NLP research.
Property
Value
Documents
28,622
Total Words
13,664,375
Total Subword Tokens… See the full description on the dataset page: https://huggingface.co/datasets/kamal-018/Maithili-Corpus.
