datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tokenise_spanish_datasetDataset tokenizado para TTS en español. Incluye audio, transcripción, emoción, y códigos SNAC.
everyayah-with-tajweed-tokensspanish_tokenizedenglish_tokenizedaudio-tokenizer-demo
Audio Tokenizer Demo Dataset
Audio samples processed by various neural audio tokenizers for comparison.
Tokenizers
Column
Tokens/sec
Codebooks
cosyvoice2
25
1
glm4voice
12.5
1
mimoaudio
6.25
8
neucodec
50
1
wavtokenizer
40
1
xcodec2
50
1
unicodec
75
1
bicodec
50
1 (Semantic) + 1 (Global Fixed 35 tokens)
flexicodec
<100
1(FSQ)+7(RVQ)
varstok
<40
1
MOSS
400
32
H-Codec-1.5
<200
4(Acoustic) + 4(Semantic)
russian_sample-dataset-tokenised_Orpheus_TTSlibritts-bpe-tokens
libritts-bpe-tokens
To learn about Trelis Enterprise Voice Services, see Trelis.com/voice-ai-services.
GPT-2 BPE tokens of LibriTTS-R text_normalized transcripts. Each utterance is terminated with the EOS token (50256). Tokens are in column token_ids (list[int]), vocab=50,257.
Splits
Mirrors the source LibriTTS-R splits (filtered by parler-tts; total ≈ 538 h):
split
utterances
hours
train.clean.100
~32 k
~53 h
train.clean.360
~112 k
~218 h
train.other.500… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/libritts-bpe-tokens.JLSpeech_tokenizedThis dataset uses focal codec to tokenize audio in JL speech dataset.
common_voice_22_0_toki_pona_parquet
Common Voice 22.0 - Toki Pona Subset!
My own Parquet conversion of Toki Pona's subset of Fsicoli's reupload of Common Voice 22 so we don't have to downgrade to Datasets 3.6 anymore!
Why?
Because the original dataset required Hugging Face Datasets 3.6 or older because it has Python code and it's in TAR shards.
This is in Parquet and works with any recent version of Hugging Face Datasets!
Details
Dataset Structure
DatasetDict({… See the full description on the dataset page: https://huggingface.co/datasets/MihaiPopa-1/common_voice_22_0_toki_pona_parquet.libritts-snac-tokens
libritts-snac-tokens
To learn about Trelis Enterprise Voice Services, see Trelis.com/voice-ai-services.
LibriTTS-R encoded with hubertsiuzdak/snac_24khz (hierarchical RVQ, 3 levels at 12 / 24 / 48 fps, 4,096 entries each).
Orpheus-style interleave per 1/12-sec audio frame: [L0[t], L1[2t], L1[2t+1], L2[4t], L2[4t+1], L2[4t+2], L2[4t+3]]. 7 tokens per audio frame, 84 fps flat.
Offset vocab 12,288: L0 in [0, 4096), L1 in [4096, 8192), L2 in [8192, 12288). Decode with level = token //… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/libritts-snac-tokens.va_tokenized
Model Card for Model ID
Model Details
Model Description
Developed by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Model type: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Finetuned from model [optional]: [More Information Needed]
Model Sources [optional]
Repository: [More Information Needed]
Paper… See the full description on the dataset page: https://huggingface.co/datasets/fguryel/va_tokenized.everyayah-mapped-to-tajweed-tokensexperiment-audio-tokenizerljspeech-tokens-v2miscellaneous_yt_chunked_tokenized
miscellaneous_yt_chunked_48k_tokenized
This is a gated Uzbek tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_tokenized.librispeech-hubert-discrete-tokens
Dataset Card for "librispeech-hubert-discrete-tokens"
More Information needed
emilia-token
Emilia EN Pocket Mimi continuous latents
This gated repository contains the English Emilia training data representation
used by the LatentTTS experiments in this project. Audio was encoded offline
with the continuous Gaussian Mimi speech VAE used by Pocket TTS. The files are
intended to let an authorized researcher reproduce latent-domain training
without encoding the source audio again.
Access and licensing
This is a derived representation of the
official Emilia… See the full description on the dataset page: https://huggingface.co/datasets/Gong1212/emilia-token.fleurs-hubert-discrete-tokens
Dataset Card for "fleurs-hubert-discrete-tokens"
More Information needed
audio-tokenizer-demo
Audio Tokenizer Demo Dataset
Audio samples processed by various neural audio tokenizers for comparison.
Tokenizers
Column
Tokens/sec
Codebooks
cosyvoice2
25
1
glm4voice
12.5
1
mimoaudio
6.25
8
neucodec
50
1
wavtokenizer
40
1
xcodec2
50
1
unicodec
75
1
bicodec
50
1 (Semantic) + 1 (Global Fixed 35 tokens)
flexicodec
<100
1(FSQ)+7(RVQ)
varstok
<40
1
MOSS
400
32
dMel_tokenized_lj_speech_c80_sr16_hop400IWSE-InstructionBasedSpeechEdit-llasa_tokenizemusiccaps-mot-tokens
MusicCaps Pre-Encoded Tokens for Mixture-of-Transformers (MoT)
Dataset Description
This dataset contains pre-encoded audio tokens from the MusicCaps dataset,
processed through Meta's MusicGen EnCodec tokenizer for use in Mixture-of-Transformers (MoT) training.
Dataset Summary
5,233 music clips encoded as discrete tokens
4 codebook layers from MusicGen's EnCodec
~500 tokens per 10-second clip
Compressed from ~12GB audio to 82MB tokens
Ready for multimodal… See the full description on the dataset page: https://huggingface.co/datasets/LeeHarrold/musiccaps-mot-tokens.omni_chunked_tokenized
omni_chunked_tokenized
This is a gated Uzbek tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized speech training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked_tokenized.morphemically_tokenised_english_cdscv_chunked_tokenized
cv_chunked_tokenized
This is a gated Uzbek tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_tokenized.yt4_chunked_tokenized
yt4_chunked_48k_tokenized
This is a gated Russian tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_chunked_tokenized.audiobook_chunked_tokenized
audiobook_chunked_tokenized
This is a gated Uzbek tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized speech training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_tokenized.espeech_podcasts_chunked_tokenized
espeech_podcasts_chunked_tokenized
This is a gated Russian tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_tokenized.yt3_chunked_tokenized
yt3_chunked_48k_tokenized
This is a gated Russian tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_tokenized.yt2_chunked_tokenized
yt2_chunked_48k_tokenized
This is a gated Russian tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_tokenized.
