CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bingbangboom /cleaned-asr-transcriptstexttext-generation10K<n<100K1 likes21 downloads6mo agoHugging Face02SaiyanSai /cleaned-asr-transcripts-hinglish cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/SaiyanSai/cleaned-asr-transcripts-hinglish.textautomatic-speech-recognition10K<n<100K0 likes17 downloads4mo agoHugging Face03bingbangboom /cleaned-asr-transcripts-hinglish cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/cleaned-asr-transcripts-hinglish.textautomatic-speech-recognition10K<n<100K0 likes16 downloads5mo agoHugging Face04hudsongouge /podcast-transcripts-cleaned-phase2gated Podcast Transcripts Cleaned (Phase 2) Updated 2026-07-18T23-19-55Z UTC. Configs Config Rows Description episodes 4,304 Full-episode raw ASR → cleaned transcript (with episode_id / show_id) chunks 5,646 Per-chunk raw → cleaned pairs traces 287,271 Full LM traces (cleaner + evaluator): prompts, reasoning/CoT, tool outputs, pass/fail Cleaned pairs (episodes / chunks) Columns include episode_id, show_id, title, url, instruction… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/podcast-transcripts-cleaned-phase2.tabulartext-generation100K<n<1M1 likes3 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.