CoolFace
20 results

bhojpuri

ai4bharat /Rural_Women_Bhojpuri Rural Bhojpuri ASR Dataset Dataset Description This dataset is curated to foster the development of inclusive Automatic Speech Recognition (ASR) systems, with a special focus on the underrepresented voices of rural Bhojpuri women. It contains audio clips in both Bhojpuri and Hindi, collected from real-world and synthetic sources, designed to train and evaluate ASR models that can accurately recognize diverse speech patterns. This work is part of the research presented in… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Rural_Women_Bhojpuri.audioautomatic-speech-recognition10K<n<100K6 likes270 downloads1y agoHugging FaceSatyam810 /BhojpuriCorpus Bhojpuri Corpus (BhojpuriCorpus) — Monolingual Pretraining Dataset for Bhojpuri (bho) Overview BhojpuriCorpus is a monolingual pretraining dataset for the Bhojpuri language (ISO 639-3: bho), containing 386,032 documents and approximately 24.97 Million estimated tokens. It is compiled from multiple public sources and preprocessed for vocabulary training and language model pretraining. Motivation BhojpuriCorpus was compiled to aggregate, clean, and… See the full description on the dataset page: https://huggingface.co/datasets/Satyam810/BhojpuriCorpus.texttext-generation100K<n<1M4 likes57 downloads2mo agoHugging Faceankur02 /bhojpuri FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/ankur02/bhojpuri.automatic-speech-recognition10K<n<100K0 likes44 downloads2y agoHugging Facesaillab /alpaca_bhojpuri_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_bhojpuri_taco.text10K<n<100K0 likes43 downloads2y agoHugging FaceKonthouKabiAI /bhojpuri-lr-v3audio10K<n<100K2 likes35 downloads2y agoHugging Face1rsh /speech-qa-bhojpuri-hi-karyaaudion<1K0 likes27 downloads3y agoHugging Face