datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
S2T_Korean_limit_silent_spacelatent-space-train
latent-space-train
Speech dataset prepared with Trelis Studio.
Statistics
Metric
Value
Source files
1
Train samples
8
Total duration
3.4 minutes
Columns
Column
Type
Description
audio
Audio
Audio segment (16kHz) - speech only, silence stripped via VAD
text
string
Plain transcription (no timestamps) - backwards compatible
text_ts
string
Transcription WITH Whisper timestamp tokens (e.g., `<
start_time
string
Segment start in… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/latent-space-train.whale-songs-w-spaceslatent-space-validation
latent-space-validation
Speech dataset prepared with Trelis Studio.
Statistics
Metric
Value
Source files
1
Validation samples
9
Total duration
3.3 minutes
Columns
Column
Type
Description
audio
Audio
Audio segment (16kHz) - speech only, silence stripped via VAD
text
string
Plain transcription (no timestamps) - backwards compatible
text_ts
string
Transcription WITH Whisper timestamp tokens (e.g., `<
start_time
string
Segment… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/latent-space-validation.latent-space-train-from-txt
latent-space-train-from-txt
Speech dataset prepared with Trelis Studio.
Statistics
Metric
Value
Source files
1
Train samples
9
Total duration
3.4 minutes
Columns
Column
Type
Description
audio
Audio
Audio segment (16kHz)
text
string
Transcription text
start_time
string
Segment start (HH:MM:SS.mmm)
end_time
string
Segment end (HH:MM:SS.mmm)
word_timestamps
list
Word-level timestamps
source_file
string
Original audio… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/latent-space-train-from-txt.TongaASR_Space_Exampleslatent-space-train-sample
latent-space-train-sample
Speech dataset prepared with Trelis Studio.
Statistics
Metric
Value
Source files
1
Train samples
9
Total duration
3.4 minutes
Columns
Column
Type
Description
audio
Audio
Audio segment (16kHz)
text
string
Transcription text
start_time
string
Segment start (HH:MM:SS.mmm)
end_time
string
Segment end (HH:MM:SS.mmm)
word_timestamps
list
Word-level timestamps
source_file
string
Original audio filename… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/latent-space-train-sample.Space-Marine-2
