CoolFace
Datasetpublic

gieljnssns/vlaams

Flemish Dataset The vocabulary foundation is organized by linguistic categories (pronouns, verbs, nouns, adjectives) with over 200 core Flemish words. Data Types synthetic: 280,000 examples (35.0%) pattern: 160,000 examples (20.0%) conversation: 120,000 examples (15.0%) qa: 96,000 examples (12.0%) translation: 64,000 examples (8.0%) grammar: 40,000 examples (5.0%) definition: 18,000 examples (2.2%) story: 12,000 examples (1.5%) culture: 6,000 examples (0.8%)… See the full description on the dataset page: https://huggingface.co/datasets/gieljnssns/vlaams.

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
0likes12downloads
Dataset Card

Flemish Dataset

The vocabulary foundation is organized by linguistic categories (pronouns, verbs, nouns, adjectives) with over 200 core Flemish words.

Data Types

  • —synthetic: 280,000 examples (35.0%)
  • —pattern: 160,000 examples (20.0%)
  • —conversation: 120,000 examples (15.0%)
  • —qa: 96,000 examples (12.0%)
  • —translation: 64,000 examples (8.0%)
  • —grammar: 40,000 examples (5.0%)
  • —definition: 18,000 examples (2.2%)
  • —story: 12,000 examples (1.5%)
  • —culture: 6,000 examples (0.8%)
  • —proverb: 4,000 examples (0.5%)

Citation

If you find this work relevant or helpful to your work, please kindly cite it:

@misc{flemishdataset,
  title={Flemish Dataset}, 
  author={Finbarrs Oketunji},
  year={2025}
}

Copyright

(c) Copyright 2025 Finbarrs Oketunji. All Rights Reserved.