datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
default_voices_chunked_speech_restorised_tts_train_clone_pairs_raw
default_voices_chunked_speech_restorised_tts_train_clone_pairs_raw
This is a gated Uzbek raw TTS clone-pair plan for default_voices_chunked_speech_restorised.
It contains speaker-reference and target metadata rows used to build tokenized clone-pair training shards.
Contents
clone_pair_plan.jsonl: one reference/target pair per row
clone_pair_plan_summary.json: upload-time summary
Dataset Summary
Field
Value
Source dataset key… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_speech_restorised_tts_train_clone_pairs_raw.yt3_chunked_speech_restorised_tts_train_clone_pairs
yt3_chunked_speech_restorised_tts_train_clone_pairs
This is a gated Russian TTS training clone-pair dataset.
It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows.
Language
Primary language: ru (Russian)
Contents
audios/shard-*.tar: tokenized audio shards
txts/shard-*.jsonl: per-example metadata and text fields
data.lst: repository-relative shard manifest
tokenized_dataset_summary.json: upload-time… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_speech_restorised_tts_train_clone_pairs.yt2_chunked_speech_restorised_tts_train_clone_pairs
yt2_chunked_speech_restorised_tts_train_clone_pairs
This is a gated Russian TTS training clone-pair dataset.
It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows.
Language
Primary language: ru (Russian)
Contents
audios/shard-*.tar: tokenized audio shards
txts/shard-*.jsonl: per-example metadata and text fields
data.lst: repository-relative shard manifest
tokenized_dataset_summary.json: upload-time… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_speech_restorised_tts_train_clone_pairs.tbp_chunked_speech_restorised_tts_train_clone_pairs
tbp_chunked_speech_restorised_tts_train_clone_pairs
This is a gated Russian TTS training clone-pair dataset.
It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows.
Language
Primary language: ru (Russian)
Contents
audios/shard-*.tar: tokenized audio shards
txts/shard-*.jsonl: per-example metadata and text fields
data.lst: repository-relative shard manifest
tokenized_dataset_summary.json: upload-time… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked_speech_restorised_tts_train_clone_pairs.espeech_podcasts_chunked_speech_restorised_tts_train_clone_pairs_raw
espeech_podcasts_chunked_speech_restorised_tts_train_clone_pairs_raw
This is a gated Russian raw TTS clone-pair plan for espeech_podcasts_chunked_speech_restorised.
It contains speaker-reference and target metadata rows used to build tokenized clone-pair training shards.
Contents
clone_pair_plan.jsonl: one reference/target pair per row
clone_pair_plan_summary.json: upload-time summary
Dataset Summary
Field
Value
Source dataset key… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_speech_restorised_tts_train_clone_pairs_raw.audiobook_chunked_speech_restorised_tts_train_clone_pairs
audiobook_chunked_speech_restorised_tts_train_clone_pairs
This is a gated Uzbek TTS training clone-pair dataset.
It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows.
Language
Primary language: uz (Uzbek)
Contents
audios/shard-*.tar: tokenized audio shards
txts/shard-*.jsonl: per-example metadata and text fields
data.lst: repository-relative shard manifest
tokenized_dataset_summary.json:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_speech_restorised_tts_train_clone_pairs.espeech_podcasts_chunked_speech_restorised_tts_train_clone_pairs
espeech_podcasts_chunked_speech_restorised_tts_train_clone_pairs
This is a gated Russian TTS training clone-pair dataset.
It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows.
Language
Primary language: ru (Russian)
Contents
audios/shard-*.tar: tokenized audio shards
txts/shard-*.jsonl: per-example metadata and text fields
data.lst: repository-relative shard manifest… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_speech_restorised_tts_train_clone_pairs.yt4_chunked_speech_restorised_tts_train_clone_pairs
yt4_chunked_speech_restorised_tts_train_clone_pairs
This is a gated Russian TTS training clone-pair dataset.
It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows.
Language
Primary language: ru (Russian)
Contents
audios/shard-*.tar: tokenized audio shards
txts/shard-*.jsonl: per-example metadata and text fields
data.lst: repository-relative shard manifest
tokenized_dataset_summary.json: upload-time… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_chunked_speech_restorised_tts_train_clone_pairs.all15_speaker_deduped_tts_train_clone_pairs_raw
All-15 Speaker-Deduped TTS Train Clone Pairs Raw
This dataset contains raw metadata rows for speaker-deduped TTS voice-clone training pairs. It does not contain audio bytes. Rows point back to source audio records and include reference/target metadata, language, dataset, tier, and precomputed speaker-similarity fields from the mining pipeline.
Contents
data/train/distinct_speaker_clone_pair_plan.jsonl.gz: all survivor rows.
data/by_dataset/*.jsonl.gz: the same… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/all15_speaker_deduped_tts_train_clone_pairs_raw.tbp_chunked_speech_restorised_tts_train_clone_pairs_raw
tbp_chunked_speech_restorised_tts_train_clone_pairs_raw
This is a gated Russian raw TTS clone-pair plan for tbp_chunked_speech_restorised.
It contains speaker-reference and target metadata rows used to build tokenized clone-pair training shards.
Contents
clone_pair_plan.jsonl: one reference/target pair per row
clone_pair_plan_summary.json: upload-time summary
Dataset Summary
Field
Value
Source dataset key
tbp_chunked_speech_restorised… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked_speech_restorised_tts_train_clone_pairs_raw.default_voices_chunked_speech_restorised_tts_train_clone_pairs
default_voices_chunked_speech_restorised_tts_train_clone_pairs
This is a gated Uzbek TTS training clone-pair dataset.
It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows.
Language
Primary language: uz (Uzbek)
Contents
audios/shard-*.tar: tokenized audio shards
txts/shard-*.jsonl: per-example metadata and text fields
data.lst: repository-relative shard manifest
tokenized_dataset_summary.json:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_speech_restorised_tts_train_clone_pairs.miscellaneous_yt_chunked_speech_restorised_tts_train_clone_pairs_raw
miscellaneous_yt_chunked_speech_restorised_tts_train_clone_pairs_raw
This is a gated Uzbek raw TTS clone-pair plan for miscellaneous_yt_chunked_speech_restorised.
It contains speaker-reference and target metadata rows used to build tokenized clone-pair training shards.
Contents
clone_pair_plan.jsonl: one reference/target pair per row
clone_pair_plan_summary.json: upload-time summary
Dataset Summary
Field
Value
Source dataset key… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_speech_restorised_tts_train_clone_pairs_raw.yt3_chunked_speech_restorised_tts_train_clone_pairs_raw
yt3_chunked_speech_restorised_tts_train_clone_pairs_raw
This is a gated Russian raw TTS clone-pair plan for yt3_chunked_speech_restorised.
It contains speaker-reference and target metadata rows used to build tokenized clone-pair training shards.
Contents
clone_pair_plan.jsonl: one reference/target pair per row
clone_pair_plan_summary.json: upload-time summary
Dataset Summary
Field
Value
Source dataset key
yt3_chunked_speech_restorised… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_speech_restorised_tts_train_clone_pairs_raw.yt_chunked_speech_restorised_tts_train_clone_pairs
yt_chunked_speech_restorised_tts_train_clone_pairs
This is a gated Russian TTS training clone-pair dataset.
It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows.
Language
Primary language: ru (Russian)
Contents
audios/shard-*.tar: tokenized audio shards
txts/shard-*.jsonl: per-example metadata and text fields
data.lst: repository-relative shard manifest
tokenized_dataset_summary.json: upload-time… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt_chunked_speech_restorised_tts_train_clone_pairs.yt1_chunked_speech_restorised_tts_train_clone_pairs_raw
yt1_chunked_speech_restorised_tts_train_clone_pairs_raw
This is a gated Russian raw TTS clone-pair plan for yt1_chunked_speech_restorised.
It contains speaker-reference and target metadata rows used to build tokenized clone-pair training shards.
Contents
clone_pair_plan.jsonl: one reference/target pair per row
clone_pair_plan_summary.json: upload-time summary
Dataset Summary
Field
Value
Source dataset key
yt1_chunked_speech_restorised… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt1_chunked_speech_restorised_tts_train_clone_pairs_raw.yt2_chunked_speech_restorised_tts_train_clone_pairs_raw
yt2_chunked_speech_restorised_tts_train_clone_pairs_raw
This is a gated Russian raw TTS clone-pair plan for yt2_chunked_speech_restorised.
It contains speaker-reference and target metadata rows used to build tokenized clone-pair training shards.
Contents
clone_pair_plan.jsonl: one reference/target pair per row
clone_pair_plan_summary.json: upload-time summary
Dataset Summary
Field
Value
Source dataset key
yt2_chunked_speech_restorised… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_speech_restorised_tts_train_clone_pairs_raw.yt4_chunked_speech_restorised_tts_train_clone_pairs_raw
yt4_chunked_speech_restorised_tts_train_clone_pairs_raw
This is a gated Russian raw TTS clone-pair plan for yt4_chunked_speech_restorised.
It contains speaker-reference and target metadata rows used to build tokenized clone-pair training shards.
Contents
clone_pair_plan.jsonl: one reference/target pair per row
clone_pair_plan_summary.json: upload-time summary
Dataset Summary
Field
Value
Source dataset key
yt4_chunked_speech_restorised… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_chunked_speech_restorised_tts_train_clone_pairs_raw.miscellaneous_yt_chunked_speech_restorised_tts_train_clone_pairs
miscellaneous_yt_chunked_speech_restorised_tts_train_clone_pairs
This is a gated Uzbek TTS training clone-pair dataset.
It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows.
Language
Primary language: uz (Uzbek)
Contents
audios/shard-*.tar: tokenized audio shards
txts/shard-*.jsonl: per-example metadata and text fields
data.lst: repository-relative shard manifest
tokenized_dataset_summary.json:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_speech_restorised_tts_train_clone_pairs.yt1_chunked_speech_restorised_tts_train_clone_pairs
yt1_chunked_speech_restorised_tts_train_clone_pairs
This is a gated Russian TTS training clone-pair dataset.
It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows.
Language
Primary language: ru (Russian)
Contents
audios/shard-*.tar: tokenized audio shards
txts/shard-*.jsonl: per-example metadata and text fields
data.lst: repository-relative shard manifest
tokenized_dataset_summary.json: upload-time… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt1_chunked_speech_restorised_tts_train_clone_pairs.audiobook_chunked_speech_restorised_tts_train_clone_pairs_raw
audiobook_chunked_speech_restorised_tts_train_clone_pairs_raw
This is a gated Uzbek raw TTS clone-pair plan for audiobook_chunked_speech_restorised.
It contains speaker-reference and target metadata rows used to build tokenized clone-pair training shards.
Contents
clone_pair_plan.jsonl: one reference/target pair per row
clone_pair_plan_summary.json: upload-time summary
Dataset Summary
Field
Value
Source dataset key… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_speech_restorised_tts_train_clone_pairs_raw.yt_chunked_speech_restorised_tts_train_clone_pairs_raw
yt_chunked_speech_restorised_tts_train_clone_pairs_raw
This is a gated Russian raw TTS clone-pair plan for yt_chunked_speech_restorised.
It contains speaker-reference and target metadata rows used to build tokenized clone-pair training shards.
Contents
clone_pair_plan.jsonl: one reference/target pair per row
clone_pair_plan_summary.json: upload-time summary
Dataset Summary
Field
Value
Source dataset key
yt_chunked_speech_restorised… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt_chunked_speech_restorised_tts_train_clone_pairs_raw.
