CoolFace
20 results

Bhojpuri

ai4bharat /Rural_Women_Bhojpuri Rural Bhojpuri ASR Dataset Dataset Description This dataset is curated to foster the development of inclusive Automatic Speech Recognition (ASR) systems, with a special focus on the underrepresented voices of rural Bhojpuri women. It contains audio clips in both Bhojpuri and Hindi, collected from real-world and synthetic sources, designed to train and evaluate ASR models that can accurately recognize diverse speech patterns. This work is part of the research presented in… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Rural_Women_Bhojpuri.audioautomatic-speech-recognition10K<n<100K6 likes293 downloads1y agoHugging Facesaillab /alpaca_bhojpuri_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_bhojpuri_taco.text10K<n<100K0 likes51 downloads2y agoHugging FaceSatyam810 /BhojpuriCorpus Bhojpuri Corpus (BhojpuriCorpus) — Monolingual Pretraining Dataset for Bhojpuri (bho) Overview BhojpuriCorpus is a monolingual pretraining dataset for the Bhojpuri language (ISO 639-3: bho), containing 386,032 documents and approximately 24.97 Million estimated tokens. It is compiled from multiple public sources and preprocessed for vocabulary training and language model pretraining. Motivation BhojpuriCorpus was compiled to aggregate, clean, and… See the full description on the dataset page: https://huggingface.co/datasets/Satyam810/BhojpuriCorpus.texttext-generation100K<n<1M4 likes47 downloads2mo agoHugging Faceankur02 /bhojpuri FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/ankur02/bhojpuri.automatic-speech-recognition10K<n<100K0 likes46 downloads2y agoHugging FaceKonthouKabiAI /bhojpuri-lr-v3audio10K<n<100K2 likes36 downloads2y agoHugging FaceAfuu-coder /asteria-bhojpuri-assamese-civic-qa Asteria — Bhojpuri & Assamese Civic Q&A Dataset A dataset of government scheme Q&A pairs in Bhojpuri and Assamese — two low-resource Indian languages. Dataset Description This dataset was collected by Asteria, an AI Agent built for the AI Agents Hackathon 2026. The agent helps rural Indian citizens access government welfare schemes by conversing in their native language. Supported Languages Bhojpuri (bho) — spoken by 50+ million people in Bihar, UP… See the full description on the dataset page: https://huggingface.co/datasets/Afuu-coder/asteria-bhojpuri-assamese-civic-qa.tabular1K<n<10K0 likes28 downloads3mo agoHugging Face