datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reprocessed_singapore_national_speech_corpus
Dataset Card for Reprocessed National Speech Corpus
NOTE: This is an Reprocessed version KaraKaraWitch from Recursal.The official download can be found here.
Dataset Details
Dataset Description
Dataset Description:
The National Speech Corpus (NSC) is the first large-scale Singapore English corpus, sponsored by the Info-communications and Media Development Authority (IMDA) of Singapore. The objective is to serve as a primary resource of open speech data for… See the full description on the dataset page: https://huggingface.co/datasets/recursal/reprocessed_singapore_national_speech_corpus.audioset-humans-reprocessedadam-degenerate-reproMelodic_pattern_reproduction_performances_gradinggradio-bug-reproductionrepro-aiff-83876test
advwave-repro-ext
advwave-repro-ext
AdvWave adversarial audio vs Qwen2-Audio-7B-Instruct on the 520 AdvBench prompts, with clip length = 6.54s.
Files
Path
Contents
question/
520 adversarial wavs, 16 kHz mono, advwave_advbench_{id}_a1.wav.
test_250.json
Merged manifest, 520 records — the canonical file SARSteer loads.
test_250.shard0of2.json
Odd IDs (260).
test_250.shard1of4.json / test_250.shard3of4.json
Even IDs, split 130/130 across 2 GPUs.
*.craft.log… See the full description on the dataset page: https://huggingface.co/datasets/neiv06/advwave-repro-ext.
