datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
audio_pretrain_tokenized_w2v2_7audio_pretrain_tokenized_w2v2_14AudioTokenized_Orpheus_TTS_Hindiaudio_pretrain_tokenized_w2v2_13audio_pretrain_tokenized_w2v2_11audio_youtube_chunked_tokenized
audio_youtube_chunked_tokenized
This is a gated Uzbek tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_chunked_tokenized.audio_pretrain_tokenized_w2v2_5audio_pretrain_tokenized_w2v2_8audio_pretrain_tokenized_w2v2_1audio_pretrain_tokenized_w2v2_9audio_pretrain_tokenized_w2v2_10audio_pretrain_tokenized_w2v2_6audio_youtube_chunked_qwen3_tts_12hz_tokenized
audio_youtube_chunked Qwen3-TTS 12Hz tokenized
This dataset contains Qwen3-TTS 12Hz tokenized JSONL shards for audio_youtube_chunked from the OmniVoice NFA-aligned all-dataset preparation run.
It is separate from the existing non-Qwen audio_youtube_chunked_tokenized artifacts. This repo contains tokenized JSONL records with audio_codes and metadata; raw audio files are not included.
Tokenizer: Qwen/Qwen3-TTS-Tokenizer-12Hz
Rows: 559475
Shards: 18
Published: 2026-05-28T20:57:59Z… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_chunked_qwen3_tts_12hz_tokenized.audio_pretrain_tokenized_w2v2_2qjc-audio-tokenizedaudio_pretrain_tokenized_w2v2_12Genshin-Audio-Tokenized-40Hzaudio_pretrain_tokenized_w2v2_4nano4m-audio-tokenizedaudio_pretrain_tokenized_w2v2_3
