CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ai4bharat /Rural_Women_Bhojpuri Rural Bhojpuri ASR Dataset Dataset Description This dataset is curated to foster the development of inclusive Automatic Speech Recognition (ASR) systems, with a special focus on the underrepresented voices of rural Bhojpuri women. It contains audio clips in both Bhojpuri and Hindi, collected from real-world and synthetic sources, designed to train and evaluate ASR models that can accurately recognize diverse speech patterns. This work is part of the research presented in… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Rural_Women_Bhojpuri.audioautomatic-speech-recognition10K<n<100K6 likes270 downloads1y agoHugging Face02Satyam810 /BhojpuriCorpus Bhojpuri Corpus (BhojpuriCorpus) — Monolingual Pretraining Dataset for Bhojpuri (bho) Overview BhojpuriCorpus is a monolingual pretraining dataset for the Bhojpuri language (ISO 639-3: bho), containing 386,032 documents and approximately 24.97 Million estimated tokens. It is compiled from multiple public sources and preprocessed for vocabulary training and language model pretraining. Motivation BhojpuriCorpus was compiled to aggregate, clean, and… See the full description on the dataset page: https://huggingface.co/datasets/Satyam810/BhojpuriCorpus.texttext-generation100K<n<1M4 likes57 downloads2mo agoHugging Face03ankur02 /bhojpuri FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/ankur02/bhojpuri.automatic-speech-recognition10K<n<100K0 likes44 downloads2y agoHugging Face04saillab /alpaca_bhojpuri_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_bhojpuri_taco.text10K<n<100K0 likes43 downloads2y agoHugging Face05KonthouKabiAI /bhojpuri-lr-v3audio10K<n<100K2 likes35 downloads2y agoHugging Face061rsh /speech-qa-bhojpuri-hi-karyaaudion<1K0 likes27 downloads3y agoHugging Face07Afuu-coder /asteria-bhojpuri-assamese-civic-qa Asteria — Bhojpuri & Assamese Civic Q&A Dataset A dataset of government scheme Q&A pairs in Bhojpuri and Assamese — two low-resource Indian languages. Dataset Description This dataset was collected by Asteria, an AI Agent built for the AI Agents Hackathon 2026. The agent helps rural Indian citizens access government welfare schemes by conversing in their native language. Supported Languages Bhojpuri (bho) — spoken by 50+ million people in Bihar, UP… See the full description on the dataset page: https://huggingface.co/datasets/Afuu-coder/asteria-bhojpuri-assamese-civic-qa.tabular1K<n<10K0 likes26 downloads3mo agoHugging Face08kumarmanishiiit /bhojpuritexttext-generation1K<n<10K0 likes25 downloads1y agoHugging Face09nilayshenai /English-Bhojpuri_Translation_Dataset English-Bhojpuri Parallel Dataset Overview A cleaned and structured collection of parallel English-Bhojpuri sentence pairs in JSON Lines (.jsonl) format. Designed for low-resource machine translation tasks and fine-tuning models like: mBART mT5 MarianMT Derived from diverse Bhojpuri media sources and reformatted for machine learning workflows. Correct Format Each line in your JSONL file must be: {"translation": {"en": "English text", "bho": "Bhojpuri… See the full description on the dataset page: https://huggingface.co/datasets/nilayshenai/English-Bhojpuri_Translation_Dataset.texttranslation10K<n<100K2 likes20 downloads1y agoHugging Face10sidddd625 /adaption-digital-payments-and-banking-terms-and-topics-hindi-marathi-bhojpuri-maithili This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-digital payments and Banking terms and topics- Hindi, Marathi, Bhojpuri, Maithili This dataset contains question-and-answer pairs focused on personal finance and banking services in India, covering topics like UPI, net banking, tax payments, and government loan schemes. Each sample includes a user query followed by a detailed, step-by-step completion that provides actionable advice… See the full description on the dataset page: https://huggingface.co/datasets/sidddd625/adaption-digital-payments-and-banking-terms-and-topics-hindi-marathi-bhojpuri-maithili.text1K<n<10K0 likes20 downloads3mo agoHugging Face11Devvrat024 /Rural_Women_Bhojpuri Rural Bhojpuri ASR Dataset Dataset Description This dataset is curated to foster the development of inclusive Automatic Speech Recognition (ASR) systems, with a special focus on the underrepresented voices of rural Bhojpuri women. It contains audio clips in both Bhojpuri and Hindi, collected from real-world and synthetic sources, designed to train and evaluate ASR models that can accurately recognize diverse speech patterns. This work is part of the research presented in… See the full description on the dataset page: https://huggingface.co/datasets/Devvrat024/Rural_Women_Bhojpuri.audioautomatic-speech-recognition10K<n<100K0 likes18 downloads6mo agoHugging Face12SatyamDev /alpaca_data_cleaned_bhojpuri Dataset Card for Dataset Name This repository contains a translated version of the Alpaca-Cleaned dataset, originally provided by Yahma on Hugging Face. The dataset has been translated into Bhojpuri, a language spoken in the northern-eastern part of India and the Terai region of Nepal. This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description The Alpaca-Cleaned… See the full description on the dataset page: https://huggingface.co/datasets/SatyamDev/alpaca_data_cleaned_bhojpuri.texttranslation10K<n<100K1 likes16 downloads2y agoHugging Face131rsh /translate-bhojpuri-hi-karyaaudio1K<n<10K0 likes13 downloads3y agoHugging Face14abhiprd2000 /Bhojpuri-Behavioral-Corpus-8Kgated 🚀 Bhojpuri Behavioral Corpus (Phase 2: Engineered Refinement) ⚠️ NOTICE: Phase 2 Refinement This repository contains the Phase 2 Engineered Refinement. This Phase 2 is automatically refined specifically to prevent class collapse during fine-tuning. 📌 Executive Summary The Bhojpuri Behavioral Corpus (Phase 2) is a 68,822-row, rigidly balanced dataset engineered to solve the inherent instability of low-resource language fine-tuning. Moving beyond noisy… See the full description on the dataset page: https://huggingface.co/datasets/abhiprd2000/Bhojpuri-Behavioral-Corpus-8K.text10K<n<100K0 likes12 downloads6mo agoHugging Face15saillab /alpaca-bhojpuri-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-bhojpuri-cleaned.text10K<n<100K1 likes10 downloads2y agoHugging Face16Harsit /xnli2.0_train_bhojpuritabular100K<n<1M1 likes8 downloads3y agoHugging Face17Harsit /xnli2.0_bhojpuritext1K<n<10K0 likes5 downloads3y agoHugging Face18pksx01 /alpaca_bhojpuri_instructionThis dataset has been created from SatyamDev/alpaca_data_cleaned_bhojpuri for instruction finetuning purpose. text10K<n<100K1 likes5 downloads2y agoHugging Face19ojhasatwik /bhojpuri_commentry_iplaudion<1K0 likes2 downloads1y agoHugging Face20harshbheem /kreol-bhojpuri-ratingstextn<1K0 likes2 downloads4mo agoHugging Face21kumarmanishiiit /bhojpuri_synthetic_datasettexttext-generationn<1K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.