datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
maithili-instruction-tuningadaption-digital-payments-and-banking-terms-and-topics-hindi-marathi-bhojpuri-maithili
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-digital payments and Banking terms and topics- Hindi, Marathi, Bhojpuri, Maithili
This dataset contains question-and-answer pairs focused on personal finance and banking services in India, covering topics like UPI, net banking, tax payments, and government loan schemes. Each sample includes a user query followed by a detailed, step-by-step completion that provides actionable advice… See the full description on the dataset page: https://huggingface.co/datasets/sidddd625/adaption-digital-payments-and-banking-terms-and-topics-hindi-marathi-bhojpuri-maithili.Maithili-Corpus
Maithili Raw Corpus
Language: Maithili (मैथिली, ISO 639-3: mai)
License: CC-BY-4.
Size: 28,622 documents | 13.7M words | ~45.8M tokens
Format: JSONL (one paragraph per row)
Tags: unlabelled, low-resource, indic-nlp, monolingual, pretraining
Dataset Summary
A large, unlabelled corpus of written Maithili text for language model pretraining and unsupervised NLP research.
Property
Value
Documents
28,622
Total Words
13,664,375
Total Subword Tokens… See the full description on the dataset page: https://huggingface.co/datasets/kamal-018/Maithili-Corpus.data_maithili_batch_v1_01.json.1maithili_Agrade_reasoning_v1_03data_maithili_batch_1.jsondata_maithili_batch_3.jsonmaithili_reasoning_batch_A1data_maithili_batch_A3.jsondata_maithili_Agrade_v1_02.jsondata_maithili_batch_2.jsondata_maithili_batch_v2_01.jsonmaithili_Agrade_reasoning_v1_02data_maithili_batch_A2.json
