datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Latin-Audio
Dataset Summary
Vox Classica is a Latin speech corpus of ~73 hours of audio, segmented into short audio clips by sentence. Vox Classica is a large-scale, ML-ready dataset of human-read Classical Latin. It was designed to address the absence of a publicly available human-read Latin corpus large enough for model training.
Alignment and curation: Kaiyuan Zhao
Language: Latin (Classical)
Uses
This dataset is built for training and evaluating speech processing models… See the full description on the dataset page: https://huggingface.co/datasets/Ken-Z/Latin-Audio.libritts-r-mimi-latentslatam-spanish-speech-orpheus-tts-24khz
LATAM Spanish High-Quality Speech Dataset (24kHz - Orpheus TTS Ready)
Dataset Description
This dataset contains approximately 24 hours of high-quality speech audio in Latin American Spanish, specifically prepared for Text-to-Speech (TTS) applications like OrpheusTTS, which require a 24kHz sampling rate.
The audio files are derived from the Crowdsourced high-quality speech datasets made by Google and were obtained via OpenSLR. The original recordings were high-quality… See the full description on the dataset page: https://huggingface.co/datasets/GianDiego/latam-spanish-speech-orpheus-tts-24khz.uyghur-cv-latinwhisper-dari-16k
Whisper Dari 16 kHz
This dataset pairs 16 kHz WAV audio with Dari transcriptions and is arranged
for the Hugging Face audiofolder loader. The columns are audio,
transcription, and id.
from datasets import load_dataset
dataset = load_dataset("audiofolder", data_dir="whisper-dari/huggingface_dataset")
print(dataset)
The split sizes are 882 train, 111 validation, and 110 test samples.
multilingual-speech-text-corpusLatin-Audio
Dataset Summary
Vox Classica is a Latin speech corpus of ~73 hours of audio, segmented into short audio clips by sentence. Vox Classica is a large-scale, ML-ready dataset of human-read Classical Latin. It was designed to address the absence of a publicly available human-read Latin corpus large enough for model training.
Alignment and curation: Kaiyuan Zhao
Language: Latin (Classical)
Uses
This dataset is built for training and evaluating speech processing models for… See the full description on the dataset page: https://huggingface.co/datasets/reesjon9/Latin-Audio.munch-1-latent-NEW-parquet
🎙️ Urdu TTS Latent Dataset — munch-1-latent-NEW-parquet
Pre-computed DACVAE latent representations for 51,021 Urdu utterances, ready for TTS model training. No audio decoding required at training time — load the dataset, reshape the binary blob, and train.
Source
Field
Value
Source audio
Humair332/Urdu-munch-1
Codec
Aratako/Semantic-DACVAE-Japanese-32dim
Codec sample rate
48,000 Hz
Encoder hop size
1,920 samples
Latent frame rate
25.0 Hz
Latent dim… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/munch-1-latent-NEW-parquet.google-latam-spanish-uniform-vad
Google LATAM Spanish — Uniform VAD and 0.5 s Edges
This public derivative contains 15,016 Latin American Spanish utterances from
the Google crowdsourced TTS datasets packaged by ylacombe. Female and male
audio were freshly exported from the same pinned upstream revisions and passed
through exactly the same processing pipeline.
Configurations
Configuration
Train
Validation
Total
Hours including edge padding
argentina-female
3,542
379
3,921
4.181… See the full description on the dataset page: https://huggingface.co/datasets/groxaxo/google-latam-spanish-uniform-vad.google-latam-spanish-boundary-normalized
Google LATAM Spanish Boundary-Normalized Audio
Female Spanish speech from the following upstream datasets:
Argentina: ylacombe/google-argentinian-spanish
Chile: ylacombe/google-chilean-spanish
Colombia: ylacombe/google-colombian-spanish
Attribution and Thanks
Many thanks to ylacombe for publishing
and maintaining the original Argentinian, Chilean, and Colombian Spanish
datasets. The recordings, transcripts, speaker labels, and original dataset
structure come… See the full description on the dataset page: https://huggingface.co/datasets/groxaxo/google-latam-spanish-boundary-normalized.LATAM-High-Fidelity-ASR
Dataset Overview
This dataset contains high-quality conversational audio samples curated for Automatic Speech Recognition tasks in Spanish variants and Portugese.
The dataset includes:
Paired audio + transcripts
Natural, non-scripted conversational speech
Single Speaker & Dual-speaker interactions
Audio Specifications
Sampling Rate: 16 kHz – 24 kHz
Bit Depth: 16-bit
Audio Type: Non-scripted conversational speech
Supported Languages
Language… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/LATAM-High-Fidelity-ASR.latvialatrobe-ca-asr-exact-v2
La Trobe CA-ASR Exact-Boundary Dataset
This dataset contains the processed La Trobe Corpus of Spoken Australian
English examples used for the principal ASR adaptation experiment accompanying
Keeping the Ums: Conversation-Analysis-Informed Verbatim ASR for Web Speech.
Source and licence
The source is the La Trobe Corpus of Spoken Australian English (Mullan,
2002), La Trobe University, https://doi.org/10.26181/23089559. The source
collection is distributed under the… See the full description on the dataset page: https://huggingface.co/datasets/stcoats/latrobe-ca-asr-exact-v2.yt_data_24-06-24_latinlatin-america-realmms-tts-uig-script_latin-UQSpeechmms-tts-uig-script_latin-UQSpeech7Latvian-Speech-Dataset
Latvian Dataset Metadata
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Latvian (la)
🏷️ Tags
Audio, Speech, Speech Recognition, ML, Machine, Machine Learning, Latvian
📦 Size Category
n < 1K
latin-american-difflatin-american-ttsmms-tts-uig-script_latin-UQSpeech6eleven_labs_datase_latinSG-RESOURCE-AUDIOlatin_music11labs_07-08-24_latinTTS-dataset-Manipur-latin
TTS-dataset-Manipur-latin
Dataset Description
This dataset comprises a collection of Manipuri (Romanized) speech audio recordings paired with their corresponding text transcriptions. It is designed to support research and development in Text-to-Speech (TTS) systems for the Manipuri language, specifically using Romanized script for text input.
Languages
This dataset is primarily in Manipuri (ISO 639-3: mni) and uses the Latin script for its text component.… See the full description on the dataset page: https://huggingface.co/datasets/DayanandaThokchom/TTS-dataset-Manipur-latin.latvian-speech-datasetlatent-space-train
latent-space-train
Speech dataset prepared with Trelis Studio.
Statistics
Metric
Value
Source files
1
Train samples
8
Total duration
3.4 minutes
Columns
Column
Type
Description
audio
Audio
Audio segment (16kHz) - speech only, silence stripped via VAD
text
string
Plain transcription (no timestamps) - backwards compatible
text_ts
string
Transcription WITH Whisper timestamp tokens (e.g., `<
start_time
string
Segment start in… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/latent-space-train.11_labs_24_06_24_latinmms-tts-uig-script_latin-UQSpeech2
