instinct
Datasets
All datasets matching “instinct”qwen3-tts-12hz-tokenized-v1
Qwen3-TTS 12 Hz Tokenized Corpora
This public, manually gated repository archives Qwen3-TTS 12 Hz, 16-codebook
training records for the Instinct TTS preparation pipeline. It stores tokenized
JSONL shards, immutable source-repository mappings, per-file hashes, audit
reports, and filterable dataset/split provenance.
The current completed Instinct tranche contains 5,069,906 rows in 635 shards.
Audio is not duplicated here: reference_audio_ref and target_audio_ref
resolve through… See the full description on the dataset page: https://huggingface.co/datasets/instinct1912/qwen3-tts-12hz-tokenized-v1.audiobook_chunked_tts_train
audiobook_chunked_tts_train
This is a gated Uzbek TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS training workflows.… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_tts_train.audio_youtube_chunked_tts_train
audio_youtube_chunked_tts_train
This is a gated Uzbek TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_chunked_tts_train.espeech_podcasts_chunked_tts_train
espeech_podcasts_chunked_tts_train
This is a gated Russian TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_tts_train.zy_chunked_tts_train
zy_chunked_tts_train
This is a gated Uzbek TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS training workflows.
Derived… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/zy_chunked_tts_train.multilingual-tts-voice-dataset
Multilingual TTS Voice Dataset
Multilingual speech and structured voice-control data for text-to-speech research and training.
The collection covers English, Kazakh, Kyrgyz, Russian, Tajik, Turkish, Turkmen, Uzbek, and mixed-language speech. Access requests are reviewed manually.
Configurations
audio: utterances with embedded audio.
prompt_specs: structured text and delivery specifications.
clone_pairs: same-speaker reference and target pairs.
voice_profiles:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/multilingual-tts-voice-dataset.
