CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01saileshbro /nepali-cs-asr Nepali–English Code-Switched ASR A ~59-hour corpus of spontaneous Nepali–English code-switched speech clipped from publicly available STEM and CS lecture videos on YouTube. The dataset targets ASR model training and evaluation for code-switched (CS) Nepali–English speech — a variety commonly used in Nepali higher education and online tutoring, where teachers fluidly mix Nepali grammar with English technical vocabulary. v2 (2026-07) — the current revision. Splits are… See the full description on the dataset page: https://huggingface.co/datasets/saileshbro/nepali-cs-asr.audioautomatic-speech-recognition10K<n<100K1 likes882 downloads2mo agoHugging Face02sailu4 /lmaanaDataset LmaanaDataset Cleaned transcription data derived from the transcribed 100-gt-2.5 subset of atlasia/MoulSot-Full. Contents 81,603 non-empty transcription rows. original_text: source transcription. new_transcription: cleaned transcription with selected French loanwords restored to Latin script. training_text: recommended text field for training. review_status: indicates whether a candidate conversion was applied. audio/100-gt-2.5/: original audio-bearing Parquet… See the full description on the dataset page: https://huggingface.co/datasets/sailu4/lmaanaDataset.audioautomatic-speech-recognition10K<n<100K0 likes69 downloads13d agoHugging Face03Saient /two-minute-papers Dataset Card for "two-minute-papers" More Information needed textautomatic-speech-recognitionn<1K0 likes34 downloads9mo agoHugging Face04SaiyanSai /cleaned-asr-transcripts-hinglish cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/SaiyanSai/cleaned-asr-transcripts-hinglish.textautomatic-speech-recognition10K<n<100K0 likes16 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.