yilele/synth-qa-taste-codec-chat
Synthetic QA Taste-S Codec Chat 18571 single-turn Traditional Chinese QA utterances with synthesized speech, 21.6 hours of audio before codec extraction. Assistant speech is represented as: <SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY> Each text token is followed by its 16 Taste-S FSQ codes (codebooks a..p). Configurations default — messages (user question + assistant <SAY> speech), audio, and answer text. Statistics Utterances: 18571… See the full description on the dataset page: https://huggingface.co/datasets/yilele/synth-qa-taste-codec-chat.
Synthetic QA Taste-S Codec Chat
18571 single-turn Traditional Chinese QA utterances with synthesized speech, 21.6 hours of audio before codec extraction.
Assistant speech is represented as:
<SAY> texttoken <acode> <bcode> ... <pcode> ... </SAY>
Each text token is followed by its 16 Taste-S FSQ codes (codebooks a..p).
Configurations
default — messages (user question + assistant <SAY> speech), audio, and answer text.
Statistics
- Utterances: 18571 (train 18012 / validation 270 / test 289)
- Audio: 21.6 hours (16 kHz)
- Codec tokens: 5771440
Codec Vocabulary
Codec token ids follow 262144 + codebook_index * 6561 + code (codebook_index 0..15 for a..p, code 0..6560): 104,976 codec tokens over the gemma-4-E4B base vocabulary.
