datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Whisper-Hallucination
Whisper Hallucination and Repetition Probes
This is a BENCHMARK. Every evaluation config is test — do not fine-tune on it.
(The one exception is lexicon_synth, which is synthetic training material and ships its
own train/test split. It is not one of the eight benchmark arms — see below.)
Training on these clips invalidates every number you would then report. Build training
data separately from the same source corpora, excluding the items listed in
benchmark/exclusions.json in… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Whisper-Hallucination.mixed-instruction-speech-whispervq-v3-fullvoxknesset-whisper-large-v3-ct2-inference
VoxKnesset × ivrit-ai/whisper-large-v3-ct2 — inference results
Transcriptions of ivrit-ai/VoxKnesset
produced by ivrit-ai/whisper-large-v3-ct2
(faster-whisper, float16, batched, language=he, VAD on), on 8× NVIDIA A40.
Columns: speaker metadata (from VoxKnesset), duration_s, reference_text
(official Knesset protocol), model_transcription, segments_json
(start/end/text), infer_time_s, split.
Current contents: 10-example pilot from the test split (data/results_10.parquet).
Full-run… See the full description on the dataset page: https://huggingface.co/datasets/Dolevabudi/voxknesset-whisper-large-v3-ct2-inference.mixed-instruction-speech-whispervq-v4Whispered
Time-aligned multilingual ASR enrichment over Common Voice 17
A time-aligned, quality-scored enrichment layer over Common Voice 17
for 11 languages across 7 writing systems. Each row is one Common Voice clip with:
the human transcript (the ground-truth target, used as-is),
a whisper-large-v3 machine transcript (enrichment / agreement signal — not a replacement),
word- and segment-level timestamps from MMS forced alignment of the human transcript,
language-ID, WER/CER agreement… See the full description on the dataset page: https://huggingface.co/datasets/burakaydinofficial/Whispered.unprocessed_dataset_whisper_finetuninginstruction-speech-WhisperVQ-Conversation-777kinstruction-speech-whispervq-v3-subset-2-filteredpreprocessed-whisper-btb-cv-cvad-wlga-ca-2607
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2605
5bfe2d098c8486d97fac8be76d86ec9146435245
train
56:46:32
50,557
589,095
11.7
31.9
techiaith/corpws-clllc-wlga
5d00294c31c78b1d7937bb2c2bc6cc70bc18d410
clips
48:20:49
27,579… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607.whisper-browser-benchmarks
whisper-browser-benchmarks
Measurements from a Whisper transcription pipeline running entirely inside a
browser tab: which audio and video containers the browser will actually decode,
how accurate the smallest usable Whisper size is on clean synthetic speech, how
long transcription takes relative to the length of the clip, what the first
load pulls over the wire, and what happens to clips longer than the model's
30-second window.
Everything here was measured, not quoted from a… See the full description on the dataset page: https://huggingface.co/datasets/ruanjiange/whisper-browser-benchmarks.instruction-speech-v1-WhisperVQ-Conversationinstruction-speech-whispervq-v2
Dataset Overview
This dataset contains nearly nearly 930,000 English speech instruction to text answer samples, using:
The combination of homebrewltd/instruction-speech-whispervq-v1 with 200,000 samples of audio-text transcription.
Tokenized using WhisperVQ.
Usage
from datasets import load_dataset, Audio
# Load Instruction Speech dataset
dataset = load_dataset("homebrewltd/instruction-speech-whispervq-v2",split='train')
Dataset Fields
Field
Type… See the full description on the dataset page: https://huggingface.co/datasets/Menlo/instruction-speech-whispervq-v2.instruction-speech-whispervq-v2-prompt-transcribepseudolabel-malaya-speech-stt-train-whisper-large-v3ap2-whisperbench
AP2-WhisperBench
The first AP2-specific benchmark for evaluating whisper attacks
on agent-mediated payment flows.
What is in the benchmark
File
Family
Size
Description
data/attacks/tier2_v1.json
Vault + Branded
100 (50+50)
v1 diversity-holdout phrasings; undefended ASR 32%/14%
data/attacks/tier2_v2.json
Vault + Branded
100 (50+50)
v2 intermediate phrasings; undefended ASR 60%/30%
data/attacks/tier2_v3.json
Vault + Branded
100 (50+50)
v3… See the full description on the dataset page: https://huggingface.co/datasets/anonymos-2321135/ap2-whisperbench.apple-speechanalyzer-vs-whisper-cpp-mac
Apple SpeechAnalyzer vs whisper.cpp on Mac
Four complete speech-recognition benchmark runs over the same deterministic
40-speaker LibriSpeech test-clean snapshot:
Engine
Model path
WER
CER
Repeated median post-speech latency
Repeated p95
Apple SpeechAnalyzer
progressiveTranscription on macOS 26.5
1.98%
1.02%
125–132 ms
194–201 ms
whisper.cpp server
1.8.4 · ggml-small.en
4.28%
1.79%
122–125 ms
152–161 ms
Every run completed 40/40 clips with no failures. Accuracy… See the full description on the dataset page: https://huggingface.co/datasets/researchaudio/apple-speechanalyzer-vs-whisper-cpp-mac.instruction-speech-v1.5-WhisperVQ-Conversationmixed-instruction-speech-whispervq-v3-full-phase2-3instruction-speech-whispervq-v1
Dataset Overview
This dataset contains nearly over 679,000 English speech instruction to text answer samples, using:
The combination of homebrewltd/instruction-speech-encodec-v1 and homebrewltd/instruction-speech-encodec-v1.5
Tokenized using WhisperVQ.
Usage
from datasets import load_dataset, Audio
# Load Instruction Speech dataset
dataset = load_dataset("homebrewltd/raw-speech-whispervq-v1",split='train')
Dataset Fields
Field
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/Menlo/instruction-speech-whispervq-v1.WhisperPreprocessedTraincd_hparam_search_whisper_th_megaspeech_v3mixed-noise-instruction-speech-whispervq-v3KUET_Whispers_Dataset
📘 Dataset Card: KUET Whispers
🧾 Overview
KUET Whispers is a curated dataset of anonymous, emotionally expressive posts collected from HazyBoard—a student-run, anonymous confessions platform at Khulna University of Engineering & Technology (KUET) in Bangladesh. The dataset captures natural, informal, and highly engaging text data in Bangla, English, and code-mixed formats, making it suitable for a range of NLP tasks including:
Sentiment and emotion analysis
Code-mixed… See the full description on the dataset page: https://huggingface.co/datasets/Sanjidh090/KUET_Whispers_Dataset.creation-garden-whisper-inversion
basedlsg/creation-garden-whisper-inversion
Experimental data for WHISPER inversion control across multiple seeds, studying multi-agent coordination.
instruction-speech-whispervq-v3-subset-2stt_latencia_whisper_distillargev2WhisperModel_Aistt_latencia_whisper_distillargeprocessed_dataset_whisper_enwhisper_base_layer_features_val
