somu9/nonverbal-tts-filtered-tokens
MintTTS Pre-tokenized Audio Tokens Pre-extracted audio codec tokens for TTS training. Source Dataset: somu9/nonverbal-tts-filtered Codec: MOSS-Audio-Tokenizer-Nano Codec sample rate: 48,000 Hz (stereo) Frame rate: 12.5 Hz (1 frame = 80ms) Stats Metric Value Total samples 6,082 Total audio hours 15.6h Codebooks 16 Avg frames/sample 115.3 Avg duration 9.2s Format JSONL file (manifest.jsonl) where each line is:… See the full description on the dataset page: https://huggingface.co/datasets/somu9/nonverbal-tts-filtered-tokens.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face