CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01japanese-asr /whisper_transcriptions.reazon_speech_allaudio10M<n<100M16 likes46k downloads2y agoHugging Face02japanese-asr /whisper_transcriptions.reazonspeech.allaudio10M<n<100M4 likes6.8k downloads2y agoHugging Face03distil-whisper /librispeech_asr-noise Dataset Card for "librispeech_asr-noise" More Information needed audio100K<n<1M2 likes4.5k downloads3y agoHugging Face04japanese-asr /whisper_transcriptions.reazonspeech.all.wer_10.0audio1M<n<10M3 likes3.5k downloads2y agoHugging Face05distil-whisper /earnings22 Dataset Card for Earnings 22 Dataset Summary Earnings-22 provides a free-to-use benchmark of real-world, accented audio to bridge academic and industrial research. This dataset contains 125 files totalling roughly 119 hours of English language earnings calls from global countries. This dataset provides the full audios, transcripts, and accompanying metadata such as ticker symbol, headquarters country, and our defined "Language Region". Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/distil-whisper/earnings22.audio10K<n<100K19 likes2.7k downloads3y agoHugging Face06distil-whisper /meanwhile Dataset Card for "meanwhile" This dataset consists of 64 segments from The Late Show with Stephen Colbert. This dataset was published as part of the Whisper release by OpenAI. See page 19 of the Whisper paper for details. audion<1K2 likes2.5k downloads3y agoHugging Face07Whispering-GPT /linustechtips-transcript-audio Dataset Card for "linustechtips" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Linus Tech Tips. Data Fields The dataset is composed by: id: Id of the youtube video. channel: Name of the… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/linustechtips-transcript-audio.audioautomatic-speech-recognitionn<1K4 likes2.3k downloads4y agoHugging Face08Whispering-GPT /lex-fridman-podcast-transcript-audio Dataset Card for "lexFridmanPodcast-transcript-audio" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast. Data Fields The dataset is composed by: id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast-transcript-audio.audioautomatic-speech-recognitionn<1K0 likes2.1k downloads4y agoHugging Face09mesolitica /Malaysian-STT-Whisper Malaysian STT Whisper format Heavy postprocessing and post-translation to improve pseudolabeled Whisper Large V3. Also include word level timestamp. Postprocessing Check repetitive trigrams. Verify Voice Activity using Silero-VAD. Verify scores using Force Alignment. Post-translation We use mesolitica/nanot5-base-malaysian-translation-v2.1. Dataset involved Malaysian context v2 Singaporean context Indonesian context Mandarin audio Tamil audio… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-STT-Whisper.audioautomatic-speech-recognition10M<n<100M5 likes2k downloads1y agoHugging Face10mesolitica /pseudolabel-malaysian-youtube-whisper-large-v3-timestamp Pseudolabel Malaysian Youtube using Whisper Large V3 including Timestamp how to prepare the dataset wget https://huggingface.co/datasets/mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp/resolve/main/prepared-pseudolabel.jsonl huggingface-cli download --repo-type dataset \ --include 'output-audio-*.zip' \ --local-dir './' \ --max-workers 20 \ mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp wget… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp.audio1M<n<10M0 likes1k downloads1y agoHugging Face11japanese-asr /whisper_transcriptions.mlsaudio10M<n<100M1 likes968 downloads2y agoHugging Face12hlmshkr /mosaic-whisper-combinedAudio files from these links https://huggingface.co/datasets/mesolitica/pseudolabel-malaya-speech-stt-train-whisper-large-v3-timestamp https://huggingface.co/datasets/mesolitica/pseudolabel-imda-large-v3-timestamp https://huggingface.co/datasets/mesolitica/pseudolabel-malaysian-youtube-whisper-large-v3-timestamp https://huggingface.co/datasets/mesolitica/pseudolabel-indonesian-large-v3-timestamp https://huggingface.co/datasets/mesolitica/pseudolabel-nusantara-large-v3-timestamp text1M<n<10M0 likes872 downloads2y agoHugging Face13malaysia-ai /pseudolabel-dialects-youtube-whisper-large-v3 malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3 Pseudolabel malaysia-ai/malaysian-dialects-youtube using openai/whisper-large-v3 How to prepare the dataset huggingface-cli download --repo-type dataset \ --include '*.zip' \ --local-dir './' \ --max-workers 20 \ malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3 wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py python3… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3.audio1M<n<10M0 likes869 downloads1y agoHugging Face14fosple /german-asr-mixed-whisper Dataset Card Dataset Sources and Licensing This dataset is a mixture of several German and multilingual speech datasets. For each dataset, the license of the original author applies. Please consult the linked sources for detailed licensing information and terms of use. Dataset Name Original Source / Author Link TUDA-De German Speech Corpus LT Group at UHH / TU Darmstadt https://huggingface.co/datasets/uhhlt/Tuda-De Mozilla Common Voice Mozilla Foundation… See the full description on the dataset page: https://huggingface.co/datasets/fosple/german-asr-mixed-whisper.audioautomatic-speech-recognition1M<n<10M0 likes842 downloads6mo agoHugging Face15argmaxinc /whisperkit-test-dataaudion<1K0 likes693 downloads4mo agoHugging Face16japanese-asr /whisper_transcriptions.reazonspeech.large.wer_10.0audio1M<n<10M0 likes663 downloads3y agoHugging Face17jan-hq /instruction-convert-audio-whispervq-llama3.2-compresstext1M<n<10M0 likes655 downloads2y agoHugging Face18makaveli10 /indic-superb-whisperaudio1K<n<10K0 likes614 downloads3y agoHugging Face19distil-whisper /librispeech_asr-prompted Dataset Card for "librispeech_asr-prompted" More Information needed audio100K<n<1M0 likes608 downloads3y agoHugging Face20distil-whisper /tedlium-prompted Dataset Card for "tedlium-prompted" More Information needed audio100K<n<1M1 likes594 downloads3y agoHugging Face21skypro1111 /whisper-dataset-ytb-uk Dataset Card for Dataset Name This dataset is collected from youtube. audioautomatic-speech-recognition10K<n<100K2 likes582 downloads3y agoHugging Face22ivrit-ai /knesset-plenums-whisper-traininggated Dataset Card for ivrit.ai - Knesset Plenums Whisper Training This is a whisper-formatted version of the ivrit.ai Knesset Plenums dataset. This dataset was created by splitting long audio recordings, along with their respective transcriptions, into audio slices of 30 seconds or less. Each such slice represents one or more consecutive segments, along with timestamp token data and the previous slice's transcription. The code for this dataset preparation process is available on the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums-whisper-training.audiotext-to-speech100K<n<1M3 likes558 downloads10mo agoHugging Face23Scicom-intl /Whisper-Hallucination Whisper Hallucination and Repetition Probes This is a BENCHMARK. Every evaluation config is test — do not fine-tune on it. (The one exception is lexicon_synth, which is synthetic training material and ships its own train/test split. It is not one of the eight benchmark arms — see below.) Training on these clips invalidates every number you would then report. Build training data separately from the same source corpora, excluding the items listed in benchmark/exclusions.json in… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Whisper-Hallucination.audioautomatic-speech-recognition100K<n<1M0 likes442 downloads6h agoHugging Face24collabora /whisperspeech The WhisperSpeech Dataset This dataset contains data to train SPEAR TTS-like text-to-speech models that utilized semantic tokens derived from the OpenAI Whisper speech recognition model. We currently provide semantic and acoustic tokens for the LibriLight and LibriTTS datasets (English only). Acoustic tokens: 24kHz EnCodec 6kbps (8 quantizers) Semantic tokens: Whisper tiny VQ bottleneck trained on a subset of LibriLight Available LibriLight subsets: small/medium/large… See the full description on the dataset page: https://huggingface.co/datasets/collabora/whisperspeech.texttext-to-speech1K<n<10K19 likes417 downloads3y agoHugging Face25Diomande /bambara-whisper-featurestext100K<n<1M0 likes406 downloads5mo agoHugging Face26mesolitica /Malaysian-STT-Whisper-Stage2 Malaysian STT Whisper Stage 2 Extra dataset to compliment mesolitica/Malaysian-STT-Whisper. This dataset is stronger in confidence and suitable for second stage / annealing finetuning. how to prepare the dataset huggingface-cli download \ mesolitica/Malaysian-STT-Whisper-Stage2 \ --include "*.zip" \ --repo-type "dataset" \ --local-dir './' huggingface-cli download \ mesolitica/Malaysian-Multiturn-Chat-Assistant \ --include "*.zip" \ --exclude "voice/*.zip" \ --repo-type… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-STT-Whisper-Stage2.text10M<n<100M2 likes403 downloads1y agoHugging Face27uniiiii /Whisper-fine-tune-2audio1M<n<10M0 likes401 downloads2y agoHugging Face28bgstud /libri-proc-whisperaudio10K<n<100K0 likes394 downloads4y agoHugging Face29jan-hq /instruction-convert-audio-whispervq-llama3.2text1M<n<10M0 likes381 downloads2y agoHugging Face30jan-hq /instruction-convert-audio-whispervq-llama3.2-deduptext1M<n<10M0 likes370 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.