datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
libritts-r-mimi-latentsmultilingual-speech-text-corpusmunch-1-latent-NEW-parquet
🎙️ Urdu TTS Latent Dataset — munch-1-latent-NEW-parquet
Pre-computed DACVAE latent representations for 51,021 Urdu utterances, ready for TTS model training. No audio decoding required at training time — load the dataset, reshape the binary blob, and train.
Source
Field
Value
Source audio
Humair332/Urdu-munch-1
Codec
Aratako/Semantic-DACVAE-Japanese-32dim
Codec sample rate
48,000 Hz
Encoder hop size
1,920 samples
Latent frame rate
25.0 Hz
Latent dim… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/munch-1-latent-NEW-parquet.latent-space-train
latent-space-train
Speech dataset prepared with Trelis Studio.
Statistics
Metric
Value
Source files
1
Train samples
8
Total duration
3.4 minutes
Columns
Column
Type
Description
audio
Audio
Audio segment (16kHz) - speech only, silence stripped via VAD
text
string
Plain transcription (no timestamps) - backwards compatible
text_ts
string
Transcription WITH Whisper timestamp tokens (e.g., `<
start_time
string
Segment start in… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/latent-space-train.latent-space-validation
latent-space-validation
Speech dataset prepared with Trelis Studio.
Statistics
Metric
Value
Source files
1
Validation samples
9
Total duration
3.3 minutes
Columns
Column
Type
Description
audio
Audio
Audio segment (16kHz) - speech only, silence stripped via VAD
text
string
Plain transcription (no timestamps) - backwards compatible
text_ts
string
Transcription WITH Whisper timestamp tokens (e.g., `<
start_time
string
Segment… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/latent-space-validation.latent-space-train-from-txt
latent-space-train-from-txt
Speech dataset prepared with Trelis Studio.
Statistics
Metric
Value
Source files
1
Train samples
9
Total duration
3.4 minutes
Columns
Column
Type
Description
audio
Audio
Audio segment (16kHz)
text
string
Transcription text
start_time
string
Segment start (HH:MM:SS.mmm)
end_time
string
Segment end (HH:MM:SS.mmm)
word_timestamps
list
Word-level timestamps
source_file
string
Original audio… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/latent-space-train-from-txt.latent-space-train-sample
latent-space-train-sample
Speech dataset prepared with Trelis Studio.
Statistics
Metric
Value
Source files
1
Train samples
9
Total duration
3.4 minutes
Columns
Column
Type
Description
audio
Audio
Audio segment (16kHz)
text
string
Transcription text
start_time
string
Segment start (HH:MM:SS.mmm)
end_time
string
Segment end (HH:MM:SS.mmm)
word_timestamps
list
Word-level timestamps
source_file
string
Original audio filename… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/latent-space-train-sample.GigaS2S-1000
