datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kfo-luxury-hospitality-corpus
Americas Great Resorts: Canonical Reference Repository
Maintainer: Andrew Paul, Founder and Managing Director, Americas Great ResortsOrganization: Americas Great Resorts (americasgreatresorts.net)Published: May 2026Last Updated: September 24, 2026
Hugging Face Dataset: Version 1.31
Dataset card version: 1.31Built: September 24, 2026Source branch: Americas-Great-Resorts/AGR mainGitHub release: v1.10Release commit: 1a3f5ebSource snapshot date: September 24… See the full description on the dataset page: https://huggingface.co/datasets/Americas-Great-Resorts/kfo-luxury-hospitality-corpus.LuxInstruct
LuxInstruct
Dataset Summary
LuxInstruct is the first large-scale cross-lingual instruction tuning dataset for Luxembourgish, introduced in LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish (Philippy et al., 2025).
It addresses the lack of high-quality instruction–response data for low-resource languages by avoiding direct machine translation into Luxembourgish. Instead, it leverages aligned data from English, French, and German to generate natural… See the full description on the dataset page: https://huggingface.co/datasets/fredxlpy/LuxInstruct.luxembourgish-parliamentary-corpus
Luxembourgish Parliamentary Corpus (2023–2028)
A provenance-documented, speaker-attributed corpus of Luxembourg's parliamentary
proceedings, built from the official session reports (comptes rendus /
"D'Chamberblietchen") of the Chambre des Députés, legislature 2023–2028.
Luxembourgish (Lëtzebuergesch) is a documented low-resource language: the
Luxembourgish Wikipedia holds roughly 64,000 articles and most large language
models perform poorly in it for lack of training material.… See the full description on the dataset page: https://huggingface.co/datasets/Decima-Data/luxembourgish-parliamentary-corpus.
