datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tts-datagen
GPT-OSS 120B native reasoning traces for TTS Datagen
Summary
This dataset contains 2,865 synthetic competitive-programming questions,
45,840 independently sampled GPT-OSS 120B solutions (16 per question), and 50
verified test cases per question (143,250 test cases total). Each solution
preserves the model's native reasoning trace separately from its final answer.
The reasoning was returned by MetaGen's native Dialog Completion interface as
dialog reasoning… See the full description on the dataset page: https://huggingface.co/datasets/harman/tts-datagen.agent-sft-stitch-zh-tts-taste-codec-chat-sample
Gemma 4 E2B Taste-S multi-turn codec SFT
This dataset contains 37,362 complete Traditional Chinese agent
dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers
229,434 synthesized speech segments, approximately
520.5 hours of audio before codec extraction.
Every assistant speech segment is represented without Gemma native audio tags:
<SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY>
The first assistant output starts immediately with <SAY>.
[SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.MED-TTS
MED-TTS
MED-TTS (Multi-Emotion and Duration-annotated Text dataset for TTS) is a multilingual emotional transition dataset designed for expressive speech, audio, and multimodal generation research. Unlike conventional emotion datasets that assign a single static emotion label to an utterance, MED-TTS focuses on dynamic emotion flow within an utterance through segment-level annotations.
The dataset contains Chinese and English text samples annotated with:
utterance-level emotion… See the full description on the dataset page: https://huggingface.co/datasets/Chanson-0803/MED-TTS.LDM-TTS-Base-SFT-19K
LDM-TTS-Base-SFT-19K
Supervised fine-tuning (SFT) corpus for Large Discovery Models (LDM): a dataset that
distils an acquisition-guided, test-time search policy into a language-model proposer so
that a single forward pass emulates a full model-based optimization loop.
Dataset Summary
An LDM couples three components in a recurrent generate → select → evaluate → update
loop: an LLM that proposes candidate experiments, a probabilistic surrogate that maps
observations… See the full description on the dataset page: https://huggingface.co/datasets/Yangtze-ailab/LDM-TTS-Base-SFT-19K.TTS_Astro_data
Burmese Mahabote Astrology
မြန်မာဗေဒင်ပညာအတွက် Dataset ရအောင် စတင်စမ်းသပ်ခြင်းဖြစ်ပါသည်။
Astro Dataset v1.0.0:
မဟာဘုတ်ဋ္ဌာနနှင့် ဘွားဇာတာအဟော
Notes:
လောလောဆယ် Dataset များကို ကောင်းစွာ မပြင်ဆင်ရသေးပါ။ ဟောစာတမ်းအတွက် သေချာမွမ်းမံပြင်ဆင်ရပါဦးမည်။
Author
@Guru-ThutaSann
License
Apache License 2.0
