CoolFace
Datasetpublic

voidful/agent-sft-stitch-zh-tts-taste-codec-sample

agent-sft-stitch-zh-tts Taste-S codec sample Ten accepted synthesized clips sampled from voidful/agent-sft-stitch-zh-tts, encoded with andybi7676/taste-s-en-zhtw-small-gemma4. Extraction follows IntelliGen's stage1_extract_taste.py: the 24 kHz source audio is resampled to 16 kHz and converted to 80-bin cool-whisper features. The encoder is conditioned on the external transcript tokenized with the Gemma 4 tokenizer, with streaming disabled. codec_indices has shape… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-sample.

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes17downloads
Dataset Card

agent-sft-stitch-zh-tts Taste-S codec sample

Ten accepted synthesized clips sampled from voidful/agent-sft-stitch-zh-tts, encoded with andybi7676/taste-s-en-zhtw-small-gemma4.

Extraction follows IntelliGen's stage1_extract_taste.py: the 24 kHz source audio is resampled to 16 kHz and converted to 80-bin cool-whisper features. The encoder is conditioned on the external transcript tokenized with the Gemma 4 tokenizer, with streaming disabled.

codec_indices has shape [codec_num_frames, 16]. A frame corresponds to one conditioned text token, not a fixed-rate audio frame. Each of the 16 FSQ codebooks has 6,561 possible indices (0..6560). This repository is only a 10-row validation sample before full-dataset extraction.