datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Luxemburgish_Press_Conferences_GovLuxembourgish-Speech-Dataset
Luxembourgish Dataset Metadata
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Luxembourgish (lb)
🏷️ Tags
Audio, Speech, Speech Recognition, Machine, Machine Learning, ML
📦 Size Category
n < 1K
YodaLingua-Luxembourgish-Extended
YodaLingua-Luxembourgish-Extended
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Luxembourgish-Extended portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
21,687 audio–transcription pairs
Total duration
75 hours
Speakers
2,674 distinct speakers
Audio format… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Luxembourgish-Extended.
