datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PawolKreyol-gfc
Analyse Lexicographique du Kreyòl Guadeloupéen
Métadonnées du Corpus
Date de génération : 06 November 2025 à 20:55
Version du pipeline : 3.0 - Pipeline Unique
Source des données : Dataset POTOMITAN/PawolKreyol-gfc (Hugging Face)
Nombre de textes : 427
Tokens totaux : 22,058
Types lexicaux : 3,680
1. Corpus et Échantillonnage
1.1 Taille et Couverture
Total des tokens : 239,808
Types lexicaux uniques : 3,680
Type-Token Ratio (TTR) :… See the full description on the dataset page: https://huggingface.co/datasets/POTOMITAN/PawolKreyol-gfc.luxembourgish-corpus
CorpusLux: Luxembourgish Speech Transcriptions
Dataset of Luxembourgish speech transcriptions from government press conferences (2025-2026),
formatted as speaker-text pairs (inspired by POTOMITAN/PawolKreyol-gfc).
Dataset Structure
Data Splits
Train: 125 turns (80%)
Validation: 16 turns (10%)
Test: 16 turns (10%)
Total: 157 turns from 6 documents
Features
Feature
Type
Description
Source
string
Speaker name (e.g., "Gilles… See the full description on the dataset page: https://huggingface.co/datasets/POTOMITAN/luxembourgish-corpus.serana_skyrim_elderscrollspotong_rumput_robot.jsonpotong_rumput_advanced.jsonpotong_rumput_basic.json
