datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
everyayah-with-tajweed-tokenslibritts-bpe-tokens
libritts-bpe-tokens
To learn about Trelis Enterprise Voice Services, see Trelis.com/voice-ai-services.
GPT-2 BPE tokens of LibriTTS-R text_normalized transcripts. Each utterance is terminated with the EOS token (50256). Tokens are in column token_ids (list[int]), vocab=50,257.
Splits
Mirrors the source LibriTTS-R splits (filtered by parler-tts; total ≈ 538 h):
split
utterances
hours
train.clean.100
~32 k
~53 h
train.clean.360
~112 k
~218 h
train.other.500… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/libritts-bpe-tokens.libritts-snac-tokens
libritts-snac-tokens
To learn about Trelis Enterprise Voice Services, see Trelis.com/voice-ai-services.
LibriTTS-R encoded with hubertsiuzdak/snac_24khz (hierarchical RVQ, 3 levels at 12 / 24 / 48 fps, 4,096 entries each).
Orpheus-style interleave per 1/12-sec audio frame: [L0[t], L1[2t], L1[2t+1], L2[4t], L2[4t+1], L2[4t+2], L2[4t+3]]. 7 tokens per audio frame, 84 fps flat.
Offset vocab 12,288: L0 in [0, 4096), L1 in [4096, 8192), L2 in [8192, 12288). Decode with level = token //… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/libritts-snac-tokens.everyayah-mapped-to-tajweed-tokensljspeech-tokens-v2librispeech-hubert-discrete-tokens
Dataset Card for "librispeech-hubert-discrete-tokens"
More Information needed
fleurs-hubert-discrete-tokens
Dataset Card for "fleurs-hubert-discrete-tokens"
More Information needed
musiccaps-mot-tokens
MusicCaps Pre-Encoded Tokens for Mixture-of-Transformers (MoT)
Dataset Description
This dataset contains pre-encoded audio tokens from the MusicCaps dataset,
processed through Meta's MusicGen EnCodec tokenizer for use in Mixture-of-Transformers (MoT) training.
Dataset Summary
5,233 music clips encoded as discrete tokens
4 codebook layers from MusicGen's EnCodec
~500 tokens per 10-second clip
Compressed from ~12GB audio to 82MB tokens
Ready for multimodal… See the full description on the dataset page: https://huggingface.co/datasets/LeeHarrold/musiccaps-mot-tokens.libritts-mimi-tokens
libritts-mimi-tokens
To learn about Trelis Enterprise Voice Services, see Trelis.com/voice-ai-services.
LibriTTS-R encoded with kyutai/mimi (RVQ codec, 8 codebooks x 2,048 entries, 12.5 frames/sec). Two streams per row:
codes_semantic (list[uint32], 12.5 fps, vocab 2,048) — codebook 0 only (WavLM-distilled, content-aligned).
codes_all_flat (list[uint32], 100 fps, offset vocab 16,384) — all 8 codebooks interleaved per frame, with codebook k mapped to [k*2048, (k+1)*2048) so a flat… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/libritts-mimi-tokens.processed_discrete_tokensspanish-s-lenition-tokens
