datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Latin-Audio
Dataset Summary
Vox Classica is a Latin speech corpus of ~73 hours of audio, segmented into short audio clips by sentence. Vox Classica is a large-scale, ML-ready dataset of human-read Classical Latin. It was designed to address the absence of a publicly available human-read Latin corpus large enough for model training.
Alignment and curation: Kaiyuan Zhao
Language: Latin (Classical)
Uses
This dataset is built for training and evaluating speech processing models… See the full description on the dataset page: https://huggingface.co/datasets/Ken-Z/Latin-Audio.uyghur-cv-latinLatin-Audio
Dataset Summary
Vox Classica is a Latin speech corpus of ~73 hours of audio, segmented into short audio clips by sentence. Vox Classica is a large-scale, ML-ready dataset of human-read Classical Latin. It was designed to address the absence of a publicly available human-read Latin corpus large enough for model training.
Alignment and curation: Kaiyuan Zhao
Language: Latin (Classical)
Uses
This dataset is built for training and evaluating speech processing models for… See the full description on the dataset page: https://huggingface.co/datasets/reesjon9/Latin-Audio.yt_data_24-06-24_latinlatin-america-realmms-tts-uig-script_latin-UQSpeechmms-tts-uig-script_latin-UQSpeech7latin-american-difflatin-american-ttsmms-tts-uig-script_latin-UQSpeech6eleven_labs_datase_latinlatin_music11labs_07-08-24_latinTTS-dataset-Manipur-latin
TTS-dataset-Manipur-latin
Dataset Description
This dataset comprises a collection of Manipuri (Romanized) speech audio recordings paired with their corresponding text transcriptions. It is designed to support research and development in Text-to-Speech (TTS) systems for the Manipuri language, specifically using Romanized script for text input.
Languages
This dataset is primarily in Manipuri (ISO 639-3: mni) and uses the Latin script for its text component.… See the full description on the dataset page: https://huggingface.co/datasets/DayanandaThokchom/TTS-dataset-Manipur-latin.11_labs_24_06_24_latinmms-tts-uig-script_latin-UQSpeech2mms-tts-uig-script_latin-UQSpeech4ta-latinyt_data_11-08-24_latinmms-tts-uig-script_latin-UQSpeech3LatinYoutubeThis is a dataset with text/audio pairs of Classical Latin extracted from youtube videos from the channels Scorpio Martianus, LATINITIUS and Musa Pedestris
elevenlabs_questions_gpt_responses_1_latinmms-tts-uig-script_latin-UQSpeech5elevenlabs_exclamations_gpt_responses_1_latinsubset_african_latin_arabic
