somu9/expresso-conversational-tokens
MintTTS Pre-tokenized Audio Tokens Pre-extracted audio codec tokens for TTS training. Source Dataset: somu9/expresso-conversational Codec: MOSS-Audio-Tokenizer-Nano Codec sample rate: 48,000 Hz (stereo) Frame rate: 12.5 Hz (1 frame = 80ms) Stats Metric Value Total samples 29,487 Total audio hours 27.8h Codebooks 16 Avg frames/sample 42.5 Avg duration 3.4s Format JSONL file (manifest.jsonl) where each line is:… See the full description on the dataset page: https://huggingface.co/datasets/somu9/expresso-conversational-tokens.
08
No card is published for this repository, or it could not be fetched from Hugging Face right now.
