datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Rural_Women_Bhojpuri
Rural Bhojpuri ASR Dataset
Dataset Description
This dataset is curated to foster the development of inclusive Automatic Speech Recognition (ASR) systems, with a special focus on the underrepresented voices of rural Bhojpuri women. It contains audio clips in both Bhojpuri and Hindi, collected from real-world and synthetic sources, designed to train and evaluate ASR models that can accurately recognize diverse speech patterns.
This work is part of the research presented in… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Rural_Women_Bhojpuri.BhojpuriCorpus
Bhojpuri Corpus (BhojpuriCorpus) — Monolingual Pretraining Dataset for Bhojpuri (bho)
Overview
BhojpuriCorpus is a monolingual pretraining dataset for the Bhojpuri language (ISO 639-3: bho), containing 386,032 documents and approximately 24.97 Million estimated tokens. It is compiled from multiple public sources and preprocessed for vocabulary training and language model pretraining.
Motivation
BhojpuriCorpus was compiled to aggregate, clean, and… See the full description on the dataset page: https://huggingface.co/datasets/Satyam810/BhojpuriCorpus.bhojpuri
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/ankur02/bhojpuri.alpaca_bhojpuri_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_bhojpuri_taco.bhojpuri-lr-v3speech-qa-bhojpuri-hi-karyaasteria-bhojpuri-assamese-civic-qa
Asteria — Bhojpuri & Assamese Civic Q&A Dataset
A dataset of government scheme Q&A pairs in Bhojpuri and Assamese — two low-resource Indian languages.
Dataset Description
This dataset was collected by Asteria, an AI Agent built for the AI Agents Hackathon 2026. The agent helps rural Indian citizens access government welfare schemes by conversing in their native language.
Supported Languages
Bhojpuri (bho) — spoken by 50+ million people in Bihar, UP… See the full description on the dataset page: https://huggingface.co/datasets/Afuu-coder/asteria-bhojpuri-assamese-civic-qa.bhojpuriEnglish-Bhojpuri_Translation_Dataset
English-Bhojpuri Parallel Dataset
Overview
A cleaned and structured collection of parallel English-Bhojpuri sentence pairs in JSON Lines (.jsonl) format. Designed for low-resource machine translation tasks and fine-tuning models like:
mBART
mT5
MarianMT
Derived from diverse Bhojpuri media sources and reformatted for machine learning workflows.
Correct Format
Each line in your JSONL file must be:
{"translation": {"en": "English text", "bho": "Bhojpuri… See the full description on the dataset page: https://huggingface.co/datasets/nilayshenai/English-Bhojpuri_Translation_Dataset.adaption-digital-payments-and-banking-terms-and-topics-hindi-marathi-bhojpuri-maithili
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-digital payments and Banking terms and topics- Hindi, Marathi, Bhojpuri, Maithili
This dataset contains question-and-answer pairs focused on personal finance and banking services in India, covering topics like UPI, net banking, tax payments, and government loan schemes. Each sample includes a user query followed by a detailed, step-by-step completion that provides actionable advice… See the full description on the dataset page: https://huggingface.co/datasets/sidddd625/adaption-digital-payments-and-banking-terms-and-topics-hindi-marathi-bhojpuri-maithili.Rural_Women_Bhojpuri
Rural Bhojpuri ASR Dataset
Dataset Description
This dataset is curated to foster the development of inclusive Automatic Speech Recognition (ASR) systems, with a special focus on the underrepresented voices of rural Bhojpuri women. It contains audio clips in both Bhojpuri and Hindi, collected from real-world and synthetic sources, designed to train and evaluate ASR models that can accurately recognize diverse speech patterns.
This work is part of the research presented in… See the full description on the dataset page: https://huggingface.co/datasets/Devvrat024/Rural_Women_Bhojpuri.alpaca_data_cleaned_bhojpuri
Dataset Card for Dataset Name
This repository contains a translated version of the Alpaca-Cleaned dataset, originally provided by Yahma on Hugging Face. The dataset has been translated into Bhojpuri, a language spoken in the northern-eastern part of India and the Terai region of Nepal.
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
The Alpaca-Cleaned… See the full description on the dataset page: https://huggingface.co/datasets/SatyamDev/alpaca_data_cleaned_bhojpuri.translate-bhojpuri-hi-karyaBhojpuri-Behavioral-Corpus-8K
🚀 Bhojpuri Behavioral Corpus (Phase 2: Engineered Refinement)
⚠️ NOTICE: Phase 2 Refinement
This repository contains the Phase 2 Engineered Refinement. This Phase 2 is automatically refined specifically to prevent class collapse during fine-tuning.
📌 Executive Summary
The Bhojpuri Behavioral Corpus (Phase 2) is a 68,822-row, rigidly balanced dataset engineered to solve the inherent instability of low-resource language fine-tuning. Moving beyond noisy… See the full description on the dataset page: https://huggingface.co/datasets/abhiprd2000/Bhojpuri-Behavioral-Corpus-8K.alpaca-bhojpuri-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-bhojpuri-cleaned.xnli2.0_train_bhojpurixnli2.0_bhojpurialpaca_bhojpuri_instructionThis dataset has been created from SatyamDev/alpaca_data_cleaned_bhojpuri for instruction finetuning purpose.
bhojpuri_commentry_iplkreol-bhojpuri-ratingsbhojpuri_synthetic_dataset
