CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Whispering-GPT /linustechtips-transcript-audio Dataset Card for "linustechtips" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Linus Tech Tips. Data Fields The dataset is composed by: id: Id of the youtube video. channel: Name of the… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/linustechtips-transcript-audio.audioautomatic-speech-recognitionn<1K4 likes2.3k downloads4y agoHugging Face02Whispering-GPT /lex-fridman-podcast-transcript-audio Dataset Card for "lexFridmanPodcast-transcript-audio" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast. Data Fields The dataset is composed by: id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast-transcript-audio.audioautomatic-speech-recognitionn<1K0 likes2.1k downloads4y agoHugging Face03mesolitica /Malaysian-STT-Whisper Malaysian STT Whisper format Heavy postprocessing and post-translation to improve pseudolabeled Whisper Large V3. Also include word level timestamp. Postprocessing Check repetitive trigrams. Verify Voice Activity using Silero-VAD. Verify scores using Force Alignment. Post-translation We use mesolitica/nanot5-base-malaysian-translation-v2.1. Dataset involved Malaysian context v2 Singaporean context Indonesian context Mandarin audio Tamil audio… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-STT-Whisper.audioautomatic-speech-recognition10M<n<100M5 likes2k downloads1y agoHugging Face04mesolitica /pseudolabel-malaysian-youtube-whisper-large-v3 Pseudolabel Malaysian Youtube videos using Whisper Large V3 Original dataset at https://huggingface.co/datasets/malaysia-ai/crawl-youtube, distributed pseudolabelled using 4x A100s script at https://github.com/mesolitica/malaysian-dataset/tree/master/speech-to-text-semisupervised/pseudolabel-whisper Each audio is 30 seconds. Each audio saved in 16k sample rate. audioautomatic-speech-recognition3 likes1.3k downloads3y agoHugging Face05fosple /german-asr-mixed-whisper Dataset Card Dataset Sources and Licensing This dataset is a mixture of several German and multilingual speech datasets. For each dataset, the license of the original author applies. Please consult the linked sources for detailed licensing information and terms of use. Dataset Name Original Source / Author Link TUDA-De German Speech Corpus LT Group at UHH / TU Darmstadt https://huggingface.co/datasets/uhhlt/Tuda-De Mozilla Common Voice Mozilla Foundation… See the full description on the dataset page: https://huggingface.co/datasets/fosple/german-asr-mixed-whisper.audioautomatic-speech-recognition1M<n<10M0 likes842 downloads6mo agoHugging Face06distil-whisper /common_voice_13_0-timestamped Distil Whisper: Common Voice 13 With Timestamps This is a variant of the Common Voice 13 dataset, augmented to return the pseudo-labelled Whisper Transcriptions alongside the original dataset elements. The pseudo-labelled transcriptions were generated by labelling the input audio data with the Whisper large-v2 model with greedy sampling and timestamp prediction. For information on how the original dataset was curated, refer to the original dataset card. Standalone Usage… See the full description on the dataset page: https://huggingface.co/datasets/distil-whisper/common_voice_13_0-timestamped.automatic-speech-recognition0 likes783 downloads3y agoHugging Face07skypro1111 /whisper-dataset-ytb-uk Dataset Card for Dataset Name This dataset is collected from youtube. audioautomatic-speech-recognition10K<n<100K2 likes582 downloads3y agoHugging Face08ivrit-ai /knesset-plenums-whisper-traininggated Dataset Card for ivrit.ai - Knesset Plenums Whisper Training This is a whisper-formatted version of the ivrit.ai Knesset Plenums dataset. This dataset was created by splitting long audio recordings, along with their respective transcriptions, into audio slices of 30 seconds or less. Each such slice represents one or more consecutive segments, along with timestamp token data and the previous slice's transcription. The code for this dataset preparation process is available on the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums-whisper-training.audiotext-to-speech100K<n<1M3 likes558 downloads10mo agoHugging Face09anke01 /uyghur-whisper-finetune UyZh-FolkSpeech Whisper 微调数据集 维吾尔语-汉语平行语音数据集,适用于 Whisper 模型微调。 数据集来源 本数据集源自 UyZh-FolkSpeech,经过以下处理: 音频格式转换: M4A → WAV (16kHz, 单声道) 文本清洗: 移除不可见控制字符 (U+200E, U+200F 等) 数据集划分: train/val/test (80%/10%/10%) 目录重组: 按划分分文件夹存储 数据统计 划分 记录数 音频数 时长 train 6,348 6,348 407.60 分钟 validation 793 793 48.89 分钟 test 795 795 51.80 分钟 总计 7,936 7,936 508.29 分钟 内容分布 短句 (sentence): 3,812 条 词汇短语 (word): 4,124 条 说话人分布 每个文本条目由 4 位说话人… See the full description on the dataset page: https://huggingface.co/datasets/anke01/uyghur-whisper-finetune.audioautomatic-speech-recognition1K<n<10K0 likes530 downloads6mo agoHugging Face10Scicom-intl /Whisper-Hallucination Whisper Hallucination and Repetition Probes This is a BENCHMARK. Every evaluation config is test — do not fine-tune on it. (The one exception is lexicon_synth, which is synthetic training material and ships its own train/test split. It is not one of the eight benchmark arms — see below.) Training on these clips invalidates every number you would then report. Build training data separately from the same source corpora, excluding the items listed in benchmark/exclusions.json in… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Whisper-Hallucination.audioautomatic-speech-recognition100K<n<1M0 likes442 downloads6h agoHugging Face11Derur /whispers_win The best whispers for Windows Support me: Boosty or Donationalerts automatic-speech-recognition1 likes282 downloads1y agoHugging Face12Whispering-GPT /yannick-kilcher-transcript-audio Dataset Card for "yannic-kilcher-transcript-audio" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Yannic Kilcher. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Yannic Kilcher. Data Fields The dataset is composed by: id: Id of the youtube video. channel:… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/yannick-kilcher-transcript-audio.audioautomatic-speech-recognitionn<1K0 likes268 downloads4y agoHugging Face13flozi00 /german-asr-mixed-whisper Dataset Card Dataset Sources and Licensing This dataset is a mixture of several German and multilingual speech datasets. For each dataset, the license of the original author applies. Please consult the linked sources for detailed licensing information and terms of use. Dataset Name Original Source / Author Link TUDA-De German Speech Corpus LT Group at UHH / TU Darmstadt https://huggingface.co/datasets/uhhlt/Tuda-De Mozilla Common Voice Mozilla Foundation… See the full description on the dataset page: https://huggingface.co/datasets/flozi00/german-asr-mixed-whisper.audioautomatic-speech-recognition1M<n<10M4 likes193 downloads1y agoHugging Face14distil-whisper /tedliumThe TED-LIUM corpus is English-language TED talks, with transcriptions, sampled at 16kHz. It contains about 118 hours of speech.automatic-speech-recognition0 likes190 downloads3y agoHugging Face15latif98 /whisper-dari-16k Whisper Dari 16 kHz This dataset pairs 16 kHz WAV audio with Dari transcriptions and is arranged for the Hugging Face audiofolder loader. The columns are audio, transcription, and id. from datasets import load_dataset dataset = load_dataset("audiofolder", data_dir="whisper-dari/huggingface_dataset") print(dataset) The split sizes are 882 train, 111 validation, and 110 test samples. audioautomatic-speech-recognition1K<n<10K0 likes187 downloads2mo agoHugging Face16burakaydinofficial /Whispered Time-aligned multilingual ASR enrichment over Common Voice 17 A time-aligned, quality-scored enrichment layer over Common Voice 17 for 11 languages across 7 writing systems. Each row is one Common Voice clip with: the human transcript (the ground-truth target, used as-is), a whisper-large-v3 machine transcript (enrichment / agreement signal — not a replacement), word- and segment-level timestamps from MMS forced alignment of the human transcript, language-ID, WER/CER agreement… See the full description on the dataset page: https://huggingface.co/datasets/burakaydinofficial/Whispered.tabularautomatic-speech-recognition1M<n<10M0 likes182 downloads3mo agoHugging Face17Menlo /raw-speech-whispervq-v1 Dataset Overview This dataset contains over 2,4M English ASR samples, using: The a training set of parler-tts/mls_eng_10k Tokenized using WhisperVQ. Usage from datasets import load_dataset, Audio # Load Instruction Speech dataset dataset = load_dataset("homebrewltd/raw-speech-whispervq-v1",split='train') Dataset Fields Field Type Description tokens sequence Tokenized using Encodec text sequence Converted audio tokens Bias, Risks… See the full description on the dataset page: https://huggingface.co/datasets/Menlo/raw-speech-whispervq-v1.textautomatic-speech-recognition1M<n<10M0 likes160 downloads2y agoHugging Face18distil-whisper /common_voice_13_0 Distil Whisper: Common Voice 13 This is a variant of the Common Voice 13 dataset, augmented to return the pseudo-labelled Whisper Transcriptions alongside the original dataset elements. The pseudo-labelled transcriptions were generated by labelling the input audio data with the Whisper large-v2 model with greedy sampling. For information on how the original dataset was curated, refer to the original dataset card. Standalone Usage First, install the latest version… See the full description on the dataset page: https://huggingface.co/datasets/distil-whisper/common_voice_13_0.automatic-speech-recognition1 likes141 downloads3y agoHugging Face19Whispering-GPT /lex-fridman-podcast Dataset Card for "lexFridmanPodcast-transcript-audio" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast. Data Fields The dataset is composed by: id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast.textautomatic-speech-recognitionn<1K12 likes123 downloads3y agoHugging Face20erayyapagci /turkish-synthetic-whisper-rounds1-4.5-355h Turkish Synthetic Whisper Rounds 1–4.5 Archival release of the exact 199,590-record, 355.186-hour synthetic corpus used to fine-tune the final Round 4.5 Whisper Tiny and Base models. Each row in train.jsonl references both: training_audio: the exact clean or exactly-once postprocessed waveform used in training; and clean_audio: its original synthetic clean waveform. Audio is SHA-256 deduplicated and stored in deterministic tar.zst shards. Common Voice/FLEURS evaluation audio… See the full description on the dataset page: https://huggingface.co/datasets/erayyapagci/turkish-synthetic-whisper-rounds1-4.5-355h.automatic-speech-recognition0 likes96 downloads1mo agoHugging Face21DewiBrynJones /preprocessed-whisper-btb-cv-cvad-wlga-ca-2607 Dataset Card Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607 Revision: main Dataset Statistics Train Split Statistics Dataset Revision Split Duration (HH:MM:SS) Clips Words Words/Clip % DewiBrynJones/banc-trawsgrifiadau-bangor-2605 5bfe2d098c8486d97fac8be76d86ec9146435245 train 56:46:32 50,557 589,095 11.7 31.9 techiaith/corpws-clllc-wlga 5d00294c31c78b1d7937bb2c2bc6cc70bc18d410 clips 48:20:49 27,579… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607.tabularautomatic-speech-recognition100K<n<1M0 likes95 downloads1mo agoHugging Face22shaun3141 /jamaican-patwa-whisperA dataset of Jamaican Patwa audio recordings with transcriptions for training speech recognition models.automatic-speech-recognitionn<1K0 likes83 downloads1y agoHugging Face23distil-whisper /librispeech_asrLibriSpeech is a corpus of approximately 1000 hours of read English speech with sampling rate of 16 kHz, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned.87automatic-speech-recognition0 likes80 downloads3y agoHugging Face24ivrit-ai /crowd-recital-whisper-traininggated Dataset Card for ivrit.ai - Crowd Recital Dataset Details Dataset Description License The dataset is released under the ivrit.ai License, which enables broad research and commercial use. - Full license: https://www.ivrit.ai/en/the-license/ - FAQs: https://www.ivrit.ai/en/license-faqs/ Dataset Structure Data Fields Each example in the dataset contains: audio: An audio column containing: bytes: The audio data… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-whisper-training.audiotext-to-speech1K<n<10K3 likes79 downloads10mo agoHugging Face25distil-whisper /gigaspeech-lGigaSpeech is an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised and unsupervised training. Around 40,000 hours of transcribed audio is first collected from audiobooks, podcasts and YouTube, covering both read and spontaneous speaking styles, and a variety of topics, such as arts, science, sports, etc. A new forced alignment and segmentation pipeline is proposed to create sentence segments suitable for speech recognition training, and to filter out segments with low-quality transcription. For system training, GigaSpeech provides five subsets of different sizes, 10h, 250h, 1000h, 2500h, and 10000h. For our 10,000-hour XL training subset, we cap the word error rate at 4% during the filtering/validation stage, and for all our other smaller training subsets, we cap it at 0%. The DEV and TEST evaluation sets, on the other hand, are re-processed by professional human transcribers to ensure high transcription quality.automatic-speech-recognition2 likes72 downloads3y agoHugging Face26Whispering-GPT /yannic-kilcher-transcript Dataset Card for "yannic-kilcher-transcript" Dataset Summary This dataset is created by applying whisper to the videos of the Youtube channel Yannic Kilcher. The dataset was created a medium size whisper model. Languages Language: English Dataset Structure The dataset contains all the transcripts plus the audio of the different videos of Yannic Kilcher. Data Fields The dataset is composed by: id: Id of the youtube video. channel: Name… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/yannic-kilcher-transcript.textautomatic-speech-recognitionn<1K1 likes64 downloads4y agoHugging Face27mesolitica /pseudolabel-malaya-speech-stt-train-whisper-large-v3tabularautomatic-speech-recognition1M<n<10M1 likes64 downloads3y agoHugging Face28distil-whisper /ami-sdmThe AMI Meeting Corpus consists of 100 hours of meeting recordings. The recordings use a range of signals synchronized to a common timeline. These include close-talking and far-field microphones, individual and room-view video cameras, and output from a slide projector and an electronic whiteboard. During the meetings, the participants also have unsynchronized pens available to them that record what is written. The meetings were recorded in English using three different rooms with different acoustic properties, and include mostly non-native speakers. \nautomatic-speech-recognition1 likes62 downloads3y agoHugging Face29MadLook /arabic-whisper-multidialect Arabic Whisper Multi-Dialect ASR Dataset A comprehensive multi-dialect Arabic speech recognition dataset prepared for Whisper model fine-tuning. Dataset Description This dataset combines high-quality Arabic speech data from multiple dialects, specifically curated for fine-tuning OpenAI's Whisper models on Arabic speech recognition tasks. Dialects Included Modern Standard Arabic (MSA) - Formal Arabic used in media and formal contexts Egyptian Arabic (EGY) - The… See the full description on the dataset page: https://huggingface.co/datasets/MadLook/arabic-whisper-multidialect.audioautomatic-speech-recognition100K<n<1M2 likes62 downloads10mo agoHugging Face30distil-whisper /spgispeechThe SPGISpeech corpus is derived from company earnings calls manually transcribed by S&P Global, Inc. according to a pro- fessional style guide detailing conventions for capitalization, punctuation, denormalization of non-standard words and tran- scription of disfluencies in spontaneous speech. The basic unit of SPGISpeech is a pair consisting of a 5 to 15 second long 16 bit, 16kHz mono wav audio file and its transcription..automatic-speech-recognition0 likes59 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.