CoolFace
Datasetpublic

POTOMITAN/luxembourgish-corpus

CorpusLux: Luxembourgish Speech Transcriptions Dataset of Luxembourgish speech transcriptions from government press conferences (2025-2026), formatted as speaker-text pairs (inspired by POTOMITAN/PawolKreyol-gfc). Dataset Structure Data Splits Train: 125 turns (80%) Validation: 16 turns (10%) Test: 16 turns (10%) Total: 157 turns from 6 documents Features Feature Type Description Source string Speaker name (e.g., "Gilles… See the full description on the dataset page: https://huggingface.co/datasets/POTOMITAN/luxembourgish-corpus.

sourceHugging Faceunknownupdated 1mo agoView on Hugging Face
1likes24downloads
4 commits on main
a9713761mo ago

Add new entry for "Ons Heemecht" to train.jsonl

medhicorneille
dc7c5eb1mo ago

Delete dataset_infos.json

medhicorneille
01e71d31mo ago

init

medhicorneille
d634eca1mo ago

initial commit

medhicorneille