t2a
Datasets
All datasets matching “t2a”t2a-mommy
t2a-mommy
Female-voice ASMR corpus for the text2asmr project.
Previously published as aoxo/audios2.
Companion repos: aoxo/t2a-daddy (male voice),
aoxo/t2a-audios-v1 (the original v1 corpus).
Layout
path
what
<creator>/<title>.m4a
source audio, 48 kHz AAC, one folder per creator
<creator>/<title>.json
word-level Whisper large-v3 alignment ([] = skipped: near-silent or undecodable)
labels/qwen3omni.jsonl
non-speech ontology labels for gap clips… See the full description on the dataset page: https://huggingface.co/datasets/aoxo/t2a-mommy.atlas-32-turn-level-actor-critic
32. A turn-level actor-critic derived from the value of computation
1. Question and links
Read this first. The reading copy of this directory is t2ance/atlas-experiments under 32-turn-level-actor-critic/; the saved training steps and the per-token training arrays are only in the Hugging Face repository t2ance/atlas-32-turn-level-actor-critic.
Does a critic that predicts the return at the start of each turn, and is supervised there alone, learn on the… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-32-turn-level-actor-critic.atlas-30-openmathreasoning-genselect-training
30. Does Qwen3.5-9B learn to use candidates on OpenMathReasoning?
1. Question and links
Trained by reinforcement learning on rows of eight OpenMathReasoning GenSelect candidates with one to seven of them correct, does Qwen3.5-9B's accuracy with eight candidates on held-out rows of the same distribution rise above its untrained accuracy and above the majority vote of the eight?
Report source:… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-30-openmathreasoning-genselect-training.t2a-daddy
t2a-daddy
Male-voice ASMR corpus for the text2asmr project.
Previously published as aoxo/audios3.
Companion repos: aoxo/t2a-mommy (female voice),
aoxo/t2a-audios-v1 (the original v1 corpus).
Layout
path
what
<creator>/<title>.m4a
source audio, 48 kHz AAC, one folder per creator
<creator>/<title>.json
word-level Whisper large-v3 alignment ([] = skipped: near-silent or undecodable)
labels/qwen3omni.jsonl
non-speech ontology labels for gap clips… See the full description on the dataset page: https://huggingface.co/datasets/aoxo/t2a-daddy.atlas-31-strengthening-candidate-verification-under-rl
31. Strengthening candidate verification under reinforcement learning
1. Question and links
Read this first. The reading copy of this directory is t2ance/atlas-experiments under 31-strengthening-candidate-verification-under-rl/; the saved training steps and the per-token training arrays are on the Hugging Face repository t2ance/atlas-31-strengthening-candidate-verification-under-rl only.
How can reinforcement learning make the orchestrator's comparing and… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-31-strengthening-candidate-verification-under-rl.t2a-audios-v1
t2a-audios-v1
The original text2asmr corpus (previously aoxo/audios): 48 kHz stereo ASMR audio with word-level
alignments, used for the v1 generator (Chatterbox speech LoRA, Stable Audio Open trigger LoRA) and as
the source for the reconstructed trigger ontology.
Superseded for ontology work by aoxo/t2a-mommy and
aoxo/t2a-daddy, which are larger, creator-attributed
and split by voice.
path
what
<id>.m4a
source audio, 48 kHz
<id>.json
word-level alignment + silence… See the full description on the dataset page: https://huggingface.co/datasets/aoxo/t2a-audios-v1.
