datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
whisper-datapseudolabel-malaya-speech-stt-train-whisper-large-v3ap2-whisperbench
AP2-WhisperBench
The first AP2-specific benchmark for evaluating whisper attacks
on agent-mediated payment flows.
What is in the benchmark
File
Family
Size
Description
data/attacks/tier2_v1.json
Vault + Branded
100 (50+50)
v1 diversity-holdout phrasings; undefended ASR 32%/14%
data/attacks/tier2_v2.json
Vault + Branded
100 (50+50)
v2 intermediate phrasings; undefended ASR 60%/30%
data/attacks/tier2_v3.json
Vault + Branded
100 (50+50)
v3… See the full description on the dataset page: https://huggingface.co/datasets/anonymos-2321135/ap2-whisperbench.whisper-transcripts-the-vergeannotations_creators:
machine-generated
language:
en
language_creators:
crowdsourced
license: []
multilinguality:
monolingual
paperswithcode_id: wikitext-2
pretty_name: Whisper-Transcripts
size_categories:
1M<n<10M
source_datasets:
original
tags: []
task_categories:
text-generation
fill-mask
task_ids:
language-modeling
masked-language-modeling
Whisper-V3-Turbo-Jsonlcd_hparam_search_whisper_th_megaspeech_v3creation-garden-whisper-inversion
basedlsg/creation-garden-whisper-inversion
Experimental data for WHISPER inversion control across multiple seeds, studying multi-agent coordination.
dataset-whisper-carnavalarabic-whisper-correction
Arabic ASR Post-Correction Dataset
Dataset for correcting common errors in Whisper ASR outputs for Arabic.
Features
Input: Noisy ASR transcriptions
Target: Corrected Modern Standard Arabic (MSA)
filtered-pseudolabel-malaysian-youtube-whisper-large-v3WhisperHintswhisper-finetune-audio_test2wikivideo-asr-whisperDCS_WhisperTrainingDatasetwhisper-finetune-audio_test3PTT-RKG_Whisper_Fine-Tunetiny_sherlock_whisper_snac_combined
