datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-sft-stitch-zh-tts-taste-codec-chat-sample
Gemma 4 E2B Taste-S multi-turn codec SFT
This dataset contains 37,362 complete Traditional Chinese agent
dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers
229,434 synthesized speech segments, approximately
520.5 hours of audio before codec extraction.
Every assistant speech segment is represented without Gemma native audio tags:
<SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY>
The first assistant output starts immediately with <SAY>.
[SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.agent-sft-stitch-zh
Agent-STITCH-S 中文版 (agent-sft-stitch-zh)
由 voidful/agent-sft 轉換而成的台灣繁體中文 Agent-STITCH-S 語音助理合成資料:模型一邊對使用者說話(<SAY>)、一邊私下推理([SOPR])、一邊呼叫工具(<TOOL_CALL>)的 speak-while-acting 軌跡。
產製流程
從 agent-sft(309,322 筆)篩出具完整 tool-call 鏈(user → tool_call → tool_result → final answer)的對話;多輪對話的早前輪次保留為 context 供改寫模型 grounding。
用 google/gemma-4-26B-A4B-it 把每筆改寫成台灣繁體中文的 speech-first STITCH-S 軌跡:先安全開場 → [SOPR] 推理 → <TOOL_CALL> → 等待語音(不得洩漏 pending 結果)→ <TOOL_RESULT> → 逐步整合 → 最終口語答覆… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh.tw-pretrain
