datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
whisper-dataap2-whisperbench
AP2-WhisperBench
The first AP2-specific benchmark for evaluating whisper attacks
on agent-mediated payment flows.
What is in the benchmark
File
Family
Size
Description
data/attacks/tier2_v1.json
Vault + Branded
100 (50+50)
v1 diversity-holdout phrasings; undefended ASR 32%/14%
data/attacks/tier2_v2.json
Vault + Branded
100 (50+50)
v2 intermediate phrasings; undefended ASR 60%/30%
data/attacks/tier2_v3.json
Vault + Branded
100 (50+50)
v3… See the full description on the dataset page: https://huggingface.co/datasets/anonymos-2321135/ap2-whisperbench.pseudolabel-malaya-speech-stt-train-whisper-large-v3Whisper-V3-Turbo-Jsonlwhisper-transcripts-the-vergeannotations_creators:
machine-generated
language:
en
language_creators:
crowdsourced
license: []
multilinguality:
monolingual
paperswithcode_id: wikitext-2
pretty_name: Whisper-Transcripts
size_categories:
1M<n<10M
source_datasets:
original
tags: []
task_categories:
text-generation
fill-mask
task_ids:
language-modeling
masked-language-modeling
creation-garden-whisper-inversion
basedlsg/creation-garden-whisper-inversion
Experimental data for WHISPER inversion control across multiple seeds, studying multi-agent coordination.
dataset-whisper-carnavalarabic-whisper-correction
Arabic ASR Post-Correction Dataset
Dataset for correcting common errors in Whisper ASR outputs for Arabic.
Features
Input: Noisy ASR transcriptions
Target: Corrected Modern Standard Arabic (MSA)
whisper-finetune-audio_test2wikivideo-asr-whisperWhisperHintswhisper-finetune-audio_test3filtered-pseudolabel-malaysian-youtube-whisper-large-v3DCS_WhisperTrainingDatasettiny_sherlock_whisper_snac_combinedPTT-RKG_Whisper_Fine-Tune
