datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Latin-Audio
Dataset Summary
Vox Classica is a Latin speech corpus of ~73 hours of audio, segmented into short audio clips by sentence. Vox Classica is a large-scale, ML-ready dataset of human-read Classical Latin. It was designed to address the absence of a publicly available human-read Latin corpus large enough for model training.
Alignment and curation: Kaiyuan Zhao
Language: Latin (Classical)
Uses
This dataset is built for training and evaluating speech processing models… See the full description on the dataset page: https://huggingface.co/datasets/Ken-Z/Latin-Audio.latam-spanish-speech-orpheus-tts-24khz
LATAM Spanish High-Quality Speech Dataset (24kHz - Orpheus TTS Ready)
Dataset Description
This dataset contains approximately 24 hours of high-quality speech audio in Latin American Spanish, specifically prepared for Text-to-Speech (TTS) applications like OrpheusTTS, which require a 24kHz sampling rate.
The audio files are derived from the Crowdsourced high-quality speech datasets made by Google and were obtained via OpenSLR. The original recordings were high-quality… See the full description on the dataset page: https://huggingface.co/datasets/GianDiego/latam-spanish-speech-orpheus-tts-24khz.whisper-dari-16k
Whisper Dari 16 kHz
This dataset pairs 16 kHz WAV audio with Dari transcriptions and is arranged
for the Hugging Face audiofolder loader. The columns are audio,
transcription, and id.
from datasets import load_dataset
dataset = load_dataset("audiofolder", data_dir="whisper-dari/huggingface_dataset")
print(dataset)
The split sizes are 882 train, 111 validation, and 110 test samples.
Latin-Audio
Dataset Summary
Vox Classica is a Latin speech corpus of ~73 hours of audio, segmented into short audio clips by sentence. Vox Classica is a large-scale, ML-ready dataset of human-read Classical Latin. It was designed to address the absence of a publicly available human-read Latin corpus large enough for model training.
Alignment and curation: Kaiyuan Zhao
Language: Latin (Classical)
Uses
This dataset is built for training and evaluating speech processing models for… See the full description on the dataset page: https://huggingface.co/datasets/reesjon9/Latin-Audio.LATAM-High-Fidelity-ASR
Dataset Overview
This dataset contains high-quality conversational audio samples curated for Automatic Speech Recognition tasks in Spanish variants and Portugese.
The dataset includes:
Paired audio + transcripts
Natural, non-scripted conversational speech
Single Speaker & Dual-speaker interactions
Audio Specifications
Sampling Rate: 16 kHz – 24 kHz
Bit Depth: 16-bit
Audio Type: Non-scripted conversational speech
Supported Languages
Language… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/LATAM-High-Fidelity-ASR.latvian-text
Latvian text dataset
Data set of latvian language texts. Intended for use in AI tool development, like speech recognition or spellcheckers
Data sources used
Latvian Wikisource articles - https://wikisource.org/wiki/Category:Latvian
Literary works of Rainis - https://repository.clarin.lv/repository/xmlui/handle/20.500.12574/41
Latvian Wikipedia articles - https://huggingface.co/datasets/joelito/EU_Wikipedias
European Parliament Proceedings Parallel Corpus -… See the full description on the dataset page: https://huggingface.co/datasets/RaivisDejus/latvian-text.latrobe-ca-asr-exact-v2
La Trobe CA-ASR Exact-Boundary Dataset
This dataset contains the processed La Trobe Corpus of Spoken Australian
English examples used for the principal ASR adaptation experiment accompanying
Keeping the Ums: Conversation-Analysis-Informed Verbatim ASR for Web Speech.
Source and licence
The source is the La Trobe Corpus of Spoken Australian English (Mullan,
2002), La Trobe University, https://doi.org/10.26181/23089559. The source
collection is distributed under the… See the full description on the dataset page: https://huggingface.co/datasets/stcoats/latrobe-ca-asr-exact-v2.Latvian-Speech-Dataset
Latvian Dataset Metadata
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Latvian (la)
🏷️ Tags
Audio, Speech, Speech Recognition, ML, Machine, Machine Learning, Latvian
📦 Size Category
n < 1K
LatinAccents
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/jstack32/LatinAccents.
