datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voice-authenticity-datasetNeapolitan-Spoken-Corpus
Neapolitan Spoken Corpus (NSC)
A corpus of read Neapolitan speech for ASR evaluation, with a validated
Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters,
metric implementations, per-clip results, and error annotations.
This release supersedes the earlier 141-clip single-speaker version of this
repository. The earlier release corresponds to Speaker S1 of the present
corpus; the old audioData/ and transcripts.csv are replaced by
data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nsc-author/Neapolitan-Spoken-Corpus.PitchBench
PitchBench
A benchmark for testing what audio / acoustic signals Audio Language Models (ALMs) do
and don't understand. PitchBench probes pitch perception across 29 controlled
experiments — single-pitch ID, onsets/offsets, chords, sequences, contour, audio
effects, and polyphonic streams.
Each row is one (audio, question, answer) triple: a short WAV stimulus, the
question (prompt*) asked of the model, and the ground-truth answer fields
(experiment-specific column names).… See the full description on the dataset page: https://huggingface.co/datasets/pitchbench-authors/PitchBench.rde-neste-authPrivate-Sound-Databaseseven-voiceAI_HOME_VOICE
