CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01japanese-asr /whisper_transcriptions.reazon_speech_all.wer_10.0.vectorized1M<n<10M0 likes87k downloads2y agoHugging Face02japanese-asr /whisper_transcriptions.reazon_speech_allaudio10M<n<100M16 likes46k downloads2y agoHugging Face03japanese-asr /whisper_transcriptions.mls.wer_10.0.vectorized1M<n<10M1 likes36k downloads2y agoHugging Face04japanese-asr /whisper_transcriptions.mls.wer_10.0audio1M<n<10M2 likes16k downloads2y agoHugging Face05distil-whisper /librispeech_long Dataset Card for "librispeech_long" More Information needed audion<1K4 likes11k downloads3y agoHugging Face06japanese-asr /whisper_transcriptions.reazonspeech.allaudio10M<n<100M4 likes6.8k downloads2y agoHugging Face07collabora /whisperspeech-librilightThis is a processed LibriLight dataset ready for training the WhisperSpeech models. See https://github.com/collabora/WhisperSpeech for more details. Quick start If you want to quickly train a basic WhisperSpeech model you can start by downloading the small subset: # magic includes to download only the small and validation data splits and the accompanying config files huggingface-cli download --repo-type dataset --include '*-small-*' '*small.dataset' '*-speakers*' --local-dir . --… See the full description on the dataset page: https://huggingface.co/datasets/collabora/whisperspeech-librilight.1 likes5.3k downloads3y agoHugging Face08distil-whisper /librispeech_asr-noise Dataset Card for "librispeech_asr-noise" More Information needed audio100K<n<1M2 likes4.5k downloads3y agoHugging Face09mei986 /whisperjav-wheels WhisperJAV Pre-built Wheels Pre-built Python wheels for WhisperJAV dependencies that are difficult to compile from source. Repository Structure whisperjav-wheels/ ├── llama-cpp-python/ │ ├── cu124/ # CUDA 12.4 wheels │ ├── cu121/ # CUDA 12.1 wheels (legacy) │ └── metal/ # Apple Silicon wheels └── README.md Available Wheels llama-cpp-python For local LLM translation support (whisperjav-translate --provider local).… See the full description on the dataset page: https://huggingface.co/datasets/mei986/whisperjav-wheels.0 likes3.8k downloads8mo agoHugging Face10japanese-asr /whisper_transcriptions.reazonspeech.all.wer_10.0audio1M<n<10M3 likes3.5k downloads2y agoHugging Face11distil-whisper /figuresimagen<1K2 likes3.3k downloads3y agoHugging Face12argmaxinc /whisperkit-evals WhisperKit WhisperKit is an on-device speech recognition framework for Apple Silicon: https://github.com/argmaxinc/WhisperKit For performance and accuracy benchmarks on real devices, please see: https://huggingface.co/spaces/argmaxinc/whisperkit-benchmarks 4 likes2.9k downloads2y agoHugging Face13distil-whisper /earnings22 Dataset Card for Earnings 22 Dataset Summary Earnings-22 provides a free-to-use benchmark of real-world, accented audio to bridge academic and industrial research. This dataset contains 125 files totalling roughly 119 hours of English language earnings calls from global countries. This dataset provides the full audios, transcripts, and accompanying metadata such as ticker symbol, headquarters country, and our defined "Language Region". Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/distil-whisper/earnings22.audio10K<n<100K19 likes2.7k downloads3y agoHugging Face14chenjoya /Live-WhisperX-526K Dataset Card for Live-WhisperX-526K Uses This dataset is used for the training of the LiveCC-7B-Instruct model. We only allow the use of this dataset for academic research and educational purposes. For OpenAI GPT-4o generated user prompts, we recommend users check the OpenAI Usage Policy. Project Page: https://showlab.github.io/livecc Paper: https://huggingface.co/papers/2504.16030 Data Sources After we finished the pre-training of LiveCC-7B-Base… See the full description on the dataset page: https://huggingface.co/datasets/chenjoya/Live-WhisperX-526K.video-text-to-text100K<n<1M11 likes2.6k downloads1mo agoHugging Face15distil-whisper /meanwhile Dataset Card for "meanwhile" This dataset consists of 64 segments from The Late Show with Stephen Colbert. This dataset was published as part of the Whisper release by OpenAI. See page 19 of the Whisper paper for details. audion<1K2 likes2.5k downloads3y agoHugging Face16Whispering-GPT /linustechtips-transcript-audio Dataset Card for "linustechtips" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Linus Tech Tips. Data Fields The dataset is composed by: id: Id of the youtube video. channel: Name of the… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/linustechtips-transcript-audio.audioautomatic-speech-recognitionn<1K4 likes2.3k downloads4y agoHugging Face17echodict /whisper.cpp whisper.cpp Stable: v1.8.1 / Roadmap High-performance inference of OpenAI's Whisper automatic speech recognition (ASR) model: Plain C/C++ implementation without dependencies Apple Silicon first-class citizen - optimized via ARM NEON, Accelerate framework, Metal and Core ML AVX intrinsics support for x86 architectures VSX intrinsics support for POWER architectures Mixed F16 / F32 precision Integer quantization support Zero memory allocations at runtime Vulkan support Support… See the full description on the dataset page: https://huggingface.co/datasets/echodict/whisper.cpp.0 likes2.1k downloads6mo agoHugging Face18Whispering-GPT /lex-fridman-podcast-transcript-audio Dataset Card for "lexFridmanPodcast-transcript-audio" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast. Data Fields The dataset is composed by: id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast-transcript-audio.audioautomatic-speech-recognitionn<1K0 likes2.1k downloads4y agoHugging Face19mesolitica /Malaysian-STT-Whisper Malaysian STT Whisper format Heavy postprocessing and post-translation to improve pseudolabeled Whisper Large V3. Also include word level timestamp. Postprocessing Check repetitive trigrams. Verify Voice Activity using Silero-VAD. Verify scores using Force Alignment. Post-translation We use mesolitica/nanot5-base-malaysian-translation-v2.1. Dataset involved Malaysian context v2 Singaporean context Indonesian context Mandarin audio Tamil audio… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-STT-Whisper.audioautomatic-speech-recognition10M<n<100M5 likes2k downloads1y agoHugging Face20japanese-asr /whisper_transcriptions.reazonspeech.all.wer_10.0.vectorized1M<n<10M0 likes1.6k downloads2y agoHugging Face21mandipgoswami /whisper-rirmega-bench Whisper-RIR-Mega: Paired Clean↔Reverberant Speech Robustness Benchmark Dataset Summary Whisper-RIR-Mega is a benchmark dataset of paired clean and reverberant speech for evaluating ASR robustness to room acoustics. Each sample consists of: audio_clean: Clean speech (LibriSpeech test-clean, 16 kHz) audio_reverb: Same utterance convolved with one RIR from RIR-Mega (v2) text_ref: Ground-truth transcript RIR metadata: rir_id, RT60, DRR, C50, etc. when available Technical… See the full description on the dataset page: https://huggingface.co/datasets/mandipgoswami/whisper-rirmega-bench.audio1K<n<10K1 likes1.5k downloads7mo agoHugging Face22argmaxinc /whisperkit-evals-dataset WhisperKit Evals Dataset Overview The WhisperKit Evals Dataset is a comprehensive collection of our speech recognition evaluation results, specifically designed to benchmark the performance of WhisperKit models across various devices and operating systems. This dataset provides detailed insights into performance and quality metrics, and model behavior under different conditions. Dataset Structure The dataset is organized into JSON files, each representing a… See the full description on the dataset page: https://huggingface.co/datasets/argmaxinc/whisperkit-evals-dataset.1 likes1.5k downloads11mo agoHugging Face23bofenghuang /stt-pseudo-labeled-whisper-large-v3-multilingualThis collection includes over 189,000 hours of speech-to-text data in seven languages: English, French, Spanish, Portuguese, Italian, German, and Dutch All segments were initially sorted by their IDs (timestamps). Adjacent segments from the same source were concatenated into 30-second chunks before being decoded using Whisper-Large-V3. The only exception was Common Voice, where segments were decoded individually before concatenation. In total, over 288,000 hours of audio data were collected… See the full description on the dataset page: https://huggingface.co/datasets/bofenghuang/stt-pseudo-labeled-whisper-large-v3-multilingual.4 likes1.4k downloads2y agoHugging Face24mesolitica /pseudolabel-malaysian-youtube-whisper-large-v3 Pseudolabel Malaysian Youtube videos using Whisper Large V3 Original dataset at https://huggingface.co/datasets/malaysia-ai/crawl-youtube, distributed pseudolabelled using 4x A100s script at https://github.com/mesolitica/malaysian-dataset/tree/master/speech-to-text-semisupervised/pseudolabel-whisper Each audio is 30 seconds. Each audio saved in 16k sample rate. audioautomatic-speech-recognition3 likes1.3k downloads3y agoHugging Face25japanese-asr /whisper_transcriptions.reazon_speech_all.wer_10.0audio1M<n<10M0 likes1k downloads2y agoHugging Face26mesolitica /pseudolabel-malaysian-youtube-whisper-large-v3-timestamp Pseudolabel Malaysian Youtube using Whisper Large V3 including Timestamp how to prepare the dataset wget https://huggingface.co/datasets/mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp/resolve/main/prepared-pseudolabel.jsonl huggingface-cli download --repo-type dataset \ --include 'output-audio-*.zip' \ --local-dir './' \ --max-workers 20 \ mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp wget… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp.audio1M<n<10M0 likes1k downloads1y agoHugging Face27japanese-asr /whisper_transcriptions.mlsaudio10M<n<100M1 likes968 downloads2y agoHugging Face28hlmshkr /mosaic-whisper-combinedAudio files from these links https://huggingface.co/datasets/mesolitica/pseudolabel-malaya-speech-stt-train-whisper-large-v3-timestamp https://huggingface.co/datasets/mesolitica/pseudolabel-imda-large-v3-timestamp https://huggingface.co/datasets/mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp https://huggingface.co/datasets/mesolitica/pseudolabel-indonesian-large-v3-timestamp https://huggingface.co/datasets/mesolitica/pseudolabel-nusantara-large-v3-timestamp text1M<n<10M0 likes872 downloads2y agoHugging Face29malaysia-ai /pseudolabel-dialects-youtube-whisper-large-v3 malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3 Pseudolabel malaysia-ai/malaysian-dialects-youtube using openai/whisper-large-v3 How to prepare the dataset huggingface-cli download --repo-type dataset \ --include '*.zip' \ --local-dir './' \ --max-workers 20 \ malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3 wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py python3… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3.audio1M<n<10M0 likes869 downloads1y agoHugging Face30fosple /german-asr-mixed-whisper Dataset Card Dataset Sources and Licensing This dataset is a mixture of several German and multilingual speech datasets. For each dataset, the license of the original author applies. Please consult the linked sources for detailed licensing information and terms of use. Dataset Name Original Source / Author Link TUDA-De German Speech Corpus LT Group at UHH / TU Darmstadt https://huggingface.co/datasets/uhhlt/Tuda-De Mozilla Common Voice Mozilla Foundation… See the full description on the dataset page: https://huggingface.co/datasets/fosple/german-asr-mixed-whisper.audioautomatic-speech-recognition1M<n<10M0 likes842 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.