gieljnssns/vlaams
Flemish Dataset The vocabulary foundation is organized by linguistic categories (pronouns, verbs, nouns, adjectives) with over 200 core Flemish words. Data Types synthetic: 280,000 examples (35.0%) pattern: 160,000 examples (20.0%) conversation: 120,000 examples (15.0%) qa: 96,000 examples (12.0%) translation: 64,000 examples (8.0%) grammar: 40,000 examples (5.0%) definition: 18,000 examples (2.2%) story: 12,000 examples (1.5%) culture: 6,000 examples (0.8%)… See the full description on the dataset page: https://huggingface.co/datasets/gieljnssns/vlaams.
Flemish Dataset
The vocabulary foundation is organized by linguistic categories (pronouns, verbs, nouns, adjectives) with over 200 core Flemish words.
Data Types
- synthetic: 280,000 examples (35.0%)
- pattern: 160,000 examples (20.0%)
- conversation: 120,000 examples (15.0%)
- qa: 96,000 examples (12.0%)
- translation: 64,000 examples (8.0%)
- grammar: 40,000 examples (5.0%)
- definition: 18,000 examples (2.2%)
- story: 12,000 examples (1.5%)
- culture: 6,000 examples (0.8%)
- proverb: 4,000 examples (0.5%)
Citation
If you find this work relevant or helpful to your work, please kindly cite it:
@misc{flemishdataset,
title={Flemish Dataset},
author={Finbarrs Oketunji},
year={2025}
}Copyright
(c) Copyright 2025 Finbarrs Oketunji. All Rights Reserved.
