datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
liepa-3
LIEPA-3 — Lithuanian Speech Corpus
Didysis lietuvių kalbos garsynas (LIEPA-3)
Dataset Summary
LIEPA-3 is a large, open corpus of Lithuanian speech (~10,000 hours,
~7.5 million audio files) built for automatic speech recognition (ASR),
text-to-speech (TTS) and linguistic research. It spans read, spontaneous,
phonetically-annotated and dialectal speech recorded under a wide range of
conditions (studio, dictaphone, radio, TV, telephone, audiobooks).
Official… See the full description on the dataset page: https://huggingface.co/datasets/meldynamics/liepa-3.liepa-2
Dataset Card for LIEPA-2
Dataset Summary
The LIEPA-2 dataset is a large-scale annotated speech corpus for the Lithuanian language, developed under the project "Development of Services Controlled by Lithuanian Speech" (LIEPA-2). It is a phonetically representative, structured collection of data (audio recordings and annotations) designed for scientific research in speech technologies and the development of electronic services.
Total Duration: 1000 hours
Access:… See the full description on the dataset page: https://huggingface.co/datasets/meldynamics/liepa-2.liepa-asr
Liepa ASR Dataset
Lithuanian Automatic Speech Recognition (ASR) dataset from the LIEPA project (Lietuvių šnekos garsynas LIEPA) developed at Vilnius University.It provides a phonetically representative corpus for ASR and TTS research, capturing diverse speakers and recording styles.
Dataset Summary
The Liepa ASR dataset contains speech recordings and their transcriptions, designed for both speech recognition and speech synthesis research.
Total speakers: 376 (248… See the full description on the dataset page: https://huggingface.co/datasets/meldynamics/liepa-asr.
