datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen3-tts-12hz-tokenized-v1
Qwen3-TTS 12 Hz Tokenized Corpora
This public, manually gated repository archives Qwen3-TTS 12 Hz, 16-codebook
training records for the Instinct TTS preparation pipeline. It stores tokenized
JSONL shards, immutable source-repository mappings, per-file hashes, audit
reports, and filterable dataset/split provenance.
The current completed Instinct tranche contains 5,069,906 rows in 635 shards.
Audio is not duplicated here: reference_audio_ref and target_audio_ref
resolve through… See the full description on the dataset page: https://huggingface.co/datasets/instinct1912/qwen3-tts-12hz-tokenized-v1.audiobook_chunked_tts_train
audiobook_chunked_tts_train
This is a gated Uzbek TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS training workflows.… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_tts_train.audio_youtube_chunked_tts_train
audio_youtube_chunked_tts_train
This is a gated Uzbek TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_chunked_tts_train.espeech_podcasts_chunked_tts_train
espeech_podcasts_chunked_tts_train
This is a gated Russian TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_tts_train.zy_chunked_tts_train
zy_chunked_tts_train
This is a gated Uzbek TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS training workflows.
Derived… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/zy_chunked_tts_train.multilingual-tts-voice-dataset
Multilingual TTS Voice Dataset
Multilingual speech and structured voice-control data for text-to-speech research and training.
The collection covers English, Kazakh, Kyrgyz, Russian, Tajik, Turkish, Turkmen, Uzbek, and mixed-language speech. Access requests are reviewed manually.
Configurations
audio: utterances with embedded audio.
prompt_specs: structured text and delivery specifications.
clone_pairs: same-speaker reference and target pairs.
voice_profiles:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/multilingual-tts-voice-dataset.yt4_chunked_speech_restorised_tts_train
yt4_chunked_speech_restorised_tts_train
This is a gated Russian TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_chunked_speech_restorised_tts_train.education_chunked_tts_train
education_chunked_tts_train
This is a gated Uzbek TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS training workflows.… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/education_chunked_tts_train.yt_chunked_speech_restorised_tts_train
yt_chunked_speech_restorised_tts_train
This is a gated Russian TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt_chunked_speech_restorised_tts_train.yt3_chunked_speech_restorised_tts_train
yt3_chunked_speech_restorised_tts_train
This is a gated Russian TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_speech_restorised_tts_train.tbp_chunked_tts_train
tbp_chunked_tts_train
This is a gated Russian TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS training workflows.… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked_tts_train.yt2_chunked_speech_restorised_tts_train
yt2_chunked_speech_restorised_tts_train
This is a gated Russian TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_speech_restorised_tts_train.yt1_chunked_speech_restorised_tts_train
yt1_chunked_speech_restorised_tts_train
This is a gated Russian TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt1_chunked_speech_restorised_tts_train.Xerxes-Instruct-700K
Dataset Card for "Xerxes-Instruct-700K"
Description
Xerxes, named after a Persian King renowned for his wisdom and strategic prowess, is an amalgamation of four distinct datasets. This dataset has been curated to cater to the burgeoning needs of natural language processing tasks, particularly in the domain of conversation modeling and comprehension.
The dataset encompasses conversations sourced from a variety of sources, ranging from generative models to real-world… See the full description on the dataset page: https://huggingface.co/datasets/Instinct-AI/Xerxes-Instruct-700K.miscellaneous_yt_chunked_speech_restorised_tts_train
miscellaneous_yt_chunked_speech_restorised_tts_train
This is a gated Uzbek TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_speech_restorised_tts_train.instinct-data
Instinct Dataset
This repository contains the next-edit data used to train and evaluate Continue's state-of-the-art open Next Edit model, Instinct. The splits are given by language, with Typescript being the original, and other languages bootstrapped synthetically off of the Typescript data. For more information on the dataset, please refer to our blog post. We additionally have code available on GitHub.
raw_instinctThis dataset is from the paper "Raw Instinct: Trust Your Classifiers and Skip the Conversion" (also on arXiV: https://arxiv.org/abs/2403.14439)
It is also available on Kaggle: https://www.kaggle.com/datasets/mathiasviborg/raw-instinct
This dataset is meant to compare RAW and RGB inputs to deep learning classifiers.
All images are of grains of different rice, and one class has its own folder with training, validation and test data.
RAW files come in 40x40 16-bit .npy files, RGB files come in… See the full description on the dataset page: https://huggingface.co/datasets/vapaau/raw_instinct.cv_chunked_tts_train
cv_chunked_tts_train
This is a gated Uzbek TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_tts_train.default_voices_chunked_tts_train
default_voices_chunked_tts_train
This is a gated Uzbek TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_tts_train.miscellaneous_yt_chunked_tokenized
miscellaneous_yt_chunked_48k_tokenized
This is a gated Uzbek tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_tokenized.omni_chunked_tts_train
omni_chunked_tts_train
This is a gated Uzbek TTS training dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Prepared for TTS training workflows.… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked_tts_train.raw-instinctive-biasomni_chunked_tokenized
omni_chunked_tokenized
This is a gated Uzbek tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized speech training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked_tokenized.cv_chunked_tokenized
cv_chunked_tokenized
This is a gated Uzbek tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_tokenized.yt4_chunked_tokenized
yt4_chunked_48k_tokenized
This is a gated Russian tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_chunked_tokenized.audiobook_chunked_tokenized
audiobook_chunked_tokenized
This is a gated Uzbek tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized speech training… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_tokenized.yt3_chunked_tokenized
yt3_chunked_48k_tokenized
This is a gated Russian tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_tokenized.yt2_chunked_tokenized
yt2_chunked_48k_tokenized
This is a gated Russian tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_tokenized.espeech_podcasts_chunked_tokenized
espeech_podcasts_chunked_tokenized
This is a gated Russian tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_tokenized.yt1_chunked_tokenized
yt1_chunked_48k_tokenized
This is a gated Russian tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: ru (Russian)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt1_chunked_tokenized.
