datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
semantic-vad-eot
Semantic-VAD EOT
End-of-turn (semantic VAD) turns built from word-level forced alignments, schema-compatible
with livekit/eot-bench-data.
Each row is one user turn: an audio clip (16 kHz mp3), its words, and ordered
silence_spans. Per the eot-bench convention the last silence span is the true
end-of-turn (eot); earlier spans are mid-turn hold pauses (labels positional, not stored).
Splits
For every data type, all shards except the last form the train base; that… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/semantic-vad-eot.asr-semantic-probe-eng
ASR Semantic Probing Dataset (English)
Synthetic English audio dataset for probing whether ASR encoder representations
encode semantic category information beyond acoustic features. Constructed for
mechanistic interpretability studies of speech recognition models.
Splits
This dataset is released as a single unsplit collection. Downstream users are
expected to define their own train/test splits based on the experimental design.
For probing experiments where speaker… See the full description on the dataset page: https://huggingface.co/datasets/soaring0616/asr-semantic-probe-eng.SemanticTimbreDataset
Dataset Card for the Semantic Timbre Dataset
The Semantic Timbre Dataset is a dataset of electric guitar audio samples labelled with 19 timbre descriptors at gradually increasing timbre magnitude.
Dataset Details
The Semantic Timbre Dataset contains 275,310 audio files of monophonic electric guitar sounds.
These 275,310 audio files are grouped into 19 semantic timbre descriptors which describe the timbral characteristics of each sound.
The 19 timbre descriptors are:… See the full description on the dataset page: https://huggingface.co/datasets/JoeCameron1/SemanticTimbreDataset.semantic-vad-eot
Semantic VAD / EoT — Sampled Test Set (tchiayan/semantic-vad-eot)
This dataset is a small, sampled derivative of Scicom-intl/semantic-vad-eot, created for quick, lightweight evaluation and benchmarking of Semantic VAD / end-of-turn (EoT) detection models without needing to stream or download the full source dataset.
Only the test split of the source dataset is used. For each subset (language / domain), 100 examples are randomly sampled from a shuffle buffer, so this is intended… See the full description on the dataset page: https://huggingface.co/datasets/tchiayan/semantic-vad-eot.bark-wave-semanticreconstructed_semantic_audiosSemanticTextualSimilarity_SpokenSTSProsody_Semantic_Mismatch
