datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
t2a-mommy
t2a-mommy
Female-voice ASMR corpus for the text2asmr project.
Previously published as aoxo/audios2.
Companion repos: aoxo/t2a-daddy (male voice),
aoxo/t2a-audios-v1 (the original v1 corpus).
Layout
path
what
<creator>/<title>.m4a
source audio, 48 kHz AAC, one folder per creator
<creator>/<title>.json
word-level Whisper large-v3 alignment ([] = skipped: near-silent or undecodable)
labels/qwen3omni.jsonl
non-speech ontology labels for gap clips… See the full description on the dataset page: https://huggingface.co/datasets/aoxo/t2a-mommy.t2a-daddy
t2a-daddy
Male-voice ASMR corpus for the text2asmr project.
Previously published as aoxo/audios3.
Companion repos: aoxo/t2a-mommy (female voice),
aoxo/t2a-audios-v1 (the original v1 corpus).
Layout
path
what
<creator>/<title>.m4a
source audio, 48 kHz AAC, one folder per creator
<creator>/<title>.json
word-level Whisper large-v3 alignment ([] = skipped: near-silent or undecodable)
labels/qwen3omni.jsonl
non-speech ontology labels for gap clips… See the full description on the dataset page: https://huggingface.co/datasets/aoxo/t2a-daddy.t2a-audios-v1
t2a-audios-v1
The original text2asmr corpus (previously aoxo/audios): 48 kHz stereo ASMR audio with word-level
alignments, used for the v1 generator (Chatterbox speech LoRA, Stable Audio Open trigger LoRA) and as
the source for the reconstructed trigger ontology.
Superseded for ontology work by aoxo/t2a-mommy and
aoxo/t2a-daddy, which are larger, creator-attributed
and split by voice.
path
what
<id>.m4a
source audio, 48 kHz
<id>.json
word-level alignment + silence… See the full description on the dataset page: https://huggingface.co/datasets/aoxo/t2a-audios-v1.Urbansound8K_t2a
Dataset Card for "Urbansound8K_t2a"
More Information needed
MACS_t2a
Dataset Card for "MACS_t2a"
More Information needed
spoken-squad-t2agigaspeech_t2alegemma_t2t_data_544kFLARE-1k-Unified-T2VAFLARE-1k-Audio-T2VAsounddescs_t2a
SoundDescsT2ARetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Natural language description for different audio sources from the BBC Sound Effects webpage.
Task category
Any2AnyRetrieval (text-to-audio)
Domains
Encyclopaedic, Written
Reference
IEEE Transactions on Multimedia
Source datasets:
mteb/sounddescs_t2a
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/sounddescs_t2a.audiocaps_t2a
AudioCapsT2ARetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Natural language description for any kind of audio in the wild.
Task category
t2a
Domains
Encyclopaedic, Written
Reference
https://audiocaps.github.io/
Source datasets:
mteb/audiocaps_t2a
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("AudioCapsT2ARetrieval")
evaluator = mteb.MTEB([task])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/audiocaps_t2a.t2adata3clotho_t2a_v2
ClothoT2ARetrieval.v2
An MTEB dataset
Massive Text Embedding Benchmark
An audio captioning dataset containing audio clips from the Freesound platform and their corresponding captions. Version 2 removes empty-string queries. For more information see #5062
Task category
Any2AnyRetrieval (text-to-audio)
Domains
Encyclopaedic, Written
Reference
Clotho: An Audio Captioning Dataset
Source datasets:
mteb/Clotho
mteb/Clotho
How to evaluate on this task… See the full description on the dataset page: https://huggingface.co/datasets/lxercode/clotho_t2a_v2.legemma_t2t_data_544k_fullt2a-triggerst2a_audio_ldm2SongDescriber-T2Alegemma_t2t_data_544k_full_trimjl_corpus_t2aMusicCaps_t2a
Dataset Card for "MusicCaps_t2a"
More Information needed
LPMusicCapsMTT_t2a
LPMusicCapsMTTT2ARetrieval
An MTEB dataset
Massive Text Embedding Benchmark
LLM-generated pseudo captions for 10-second music clips from the MagnaTagATune dataset. Captions were produced by prompting a large language model with the human-annotated tags of each clip, giving four differently-styled captions per clip. Complements MusicCaps, whose captions are human-written and whose audio comes from AudioSet.
Task category
Any2AnyRetrieval (text-to-audio)
Domains
Music… See the full description on the dataset page: https://huggingface.co/datasets/hubxrt/LPMusicCapsMTT_t2a.LibriTTS_t2a
Dataset Card for "LibriTTS_t2a"
More Information needed
FLARE-1k-Unified-T2VAFLARE-1k-Audio-T2VAt2a_stable_audio_openv14-fake-vcapv-t2aEmoV_DB_t2a
Dataset Card for "EmoV_DB_t2a"
More Information needed
corpus5-t2inserts-sample-200-20260916spoken-squad-t2a
