datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
whisper-hallucinations
Whisper Hallucinations on Noise
Dataset Summary
This dataset lists common hallucinations from OpenAI Whisper when the input has no speech.
We build it from a noise-only corpus.
We run Whisper on noise clips.
We collect any non-empty text that Whisper outputs.
We deduplicate phrases and count how often they occur.
Use it to test, detect, and reduce non-speech hallucinations.
Motivation
ASR models often output text on silence or noise.
These false hits harm UX… See the full description on the dataset page: https://huggingface.co/datasets/sachaarbonel/whisper-hallucinations.whisper_transcriptions_greedywhisper-browser-benchmarks
whisper-browser-benchmarks
Measurements from a Whisper transcription pipeline running entirely inside a
browser tab: which audio and video containers the browser will actually decode,
how accurate the smallest usable Whisper size is on clean synthetic speech, how
long transcription takes relative to the length of the clip, what the first
load pulls over the wire, and what happens to clips longer than the model's
30-second window.
Everything here was measured, not quoted from a… See the full description on the dataset page: https://huggingface.co/datasets/ruanjiange/whisper-browser-benchmarks.apple-speechanalyzer-vs-whisper-cpp-mac
Apple SpeechAnalyzer vs whisper.cpp on Mac
Four complete speech-recognition benchmark runs over the same deterministic
40-speaker LibriSpeech test-clean snapshot:
Engine
Model path
WER
CER
Repeated median post-speech latency
Repeated p95
Apple SpeechAnalyzer
progressiveTranscription on macOS 26.5
1.98%
1.02%
125–132 ms
194–201 ms
whisper.cpp server
1.8.4 · ggml-small.en
4.28%
1.79%
122–125 ms
152–161 ms
Every run completed 40/40 clips with no failures. Accuracy… See the full description on the dataset page: https://huggingface.co/datasets/researchaudio/apple-speechanalyzer-vs-whisper-cpp-mac.whisper_transcriptions_greedy_timestamped2k-whisperwhisper_transcriptions_token_idswhisper_bn_inference_lpKUET_Whispers_Dataset
📘 Dataset Card: KUET Whispers
🧾 Overview
KUET Whispers is a curated dataset of anonymous, emotionally expressive posts collected from HazyBoard—a student-run, anonymous confessions platform at Khulna University of Engineering & Technology (KUET) in Bangladesh. The dataset captures natural, informal, and highly engaging text data in Bangla, English, and code-mixed formats, making it suitable for a range of NLP tasks including:
Sentiment and emotion analysis
Code-mixed… See the full description on the dataset page: https://huggingface.co/datasets/Sanjidh090/KUET_Whispers_Dataset.WhisperModel_Aiwhisper_train_guj_eng_pseudolabelledwhisperv3p_inference_lpwhisper_regression_AverageCSR_2columnsKUET_whispers_Dataset
📘 Dataset Card: KUET Whispers
🧾 Overview
KUET Whispers is a curated dataset of anonymous, emotionally expressive posts collected from HazyBoard—a student-run, anonymous confessions platform at Khulna University of Engineering & Technology (KUET) in Bangladesh. The dataset captures natural, informal, and highly engaging text data in Bangla, English, and code-mixed formats, making it suitable for a range of NLP tasks including:
Sentiment and emotion analysis
Code-mixed… See the full description on the dataset page: https://huggingface.co/datasets/sanjidh90/KUET_whispers_Dataset.200entries_Whisper_GT_AverageCSR100entries_whisper_regression_AverageCSR_2columns_100entriesATCOSIM_dictation_by_finetuned_whisper_smallwhisper_8_avg_ground_truth_scores_2columnswhisper_4_selfprom_scores_regression_2columnsWhisper-small_ACTOSIM_Test_datacloud-whisper-datasetWhisper-small_ACTOSIM_Train_dataalgerian-whisper-results2022218009_PitchScore_Whisper_IRwhisperv3_inference_lpwhisper_trainwhisper_testwhisperseg-dataset-efplwhisperseg-efpl-v2
