somu9/nonverbal-tts-filtered-tokens
MintTTS Pre-tokenized Audio Tokens Pre-extracted audio codec tokens for TTS training. Source Dataset: somu9/nonverbal-tts-filtered Codec: MOSS-Audio-Tokenizer-Nano Codec sample rate: 48,000 Hz (stereo) Frame rate: 12.5 Hz (1 frame = 80ms) Stats Metric Value Total samples 6,082 Total audio hours 15.6h Codebooks 16 Avg frames/sample 115.3 Avg duration 9.2s Format JSONL file (manifest.jsonl) where each line is:… See the full description on the dataset page: https://huggingface.co/datasets/somu9/nonverbal-tts-filtered-tokens.
17
No card is published for this repository, or it could not be fetched from Hugging Face right now.
