datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
staatsblad-synth-nl
Synthetic Dutch from the Belgisch Staatsblad
Diverse, fluent Dutch pretraining text synthesized from
guust-franssens/belgisch-staatsblad
(CC0, Belgian official-gazette filings). Adds Belgium/Flanders coverage to Dutch LM pretraining
mixes, where clean Belgian-Dutch prose is otherwise scarce.
The source text is noisy OCR from scanned PDFs, but its metadata (company, juridical form,
act type, city, date) is clean. A local LLM (google/gemma-2-9b-it)
"launders" the OCR + metadata… See the full description on the dataset page: https://huggingface.co/datasets/pdelobelle/staatsblad-synth-nl.fineweb-dutch-edu-mt
FineWeb-Edu Dutch Machine Translated
Machine-translated Dutch text dataset derived from the FineWeb-Edu corpus.
Dataset Details
Source: HuggingFaceFW/fineweb-edu (sample-10BT subset)
Translation: English → Dutch using Unbabel/Tower-Plus-9B
Size: Up to 1.5M samples
Format: Translated text with original metadata
Schema
text: Machine-translated Dutch text
id: Original sample identifier from FineWeb-Edu
url: Source URL
Quality Notice
⚠️ This… See the full description on the dataset page: https://huggingface.co/datasets/pdelobelle/fineweb-dutch-edu-mt.nemotron-dutch-mt
Nemotron Post-Training Dataset (Dutch Translation)
Machine-translated Dutch version of NVIDIA's Nemotron Post-Training Dataset, specifically the chat conversations.
Dataset Details
Source: nvidia/Nemotron-Post-Training-Dataset-v2 (chat split)
Translation: English → Dutch using Unbabel/Tower-Plus-9B
Size: 445,287 conversations with 1,327,548 total messages
Format: Conversational data with original structure preserved
Dataset Statistics
Total… See the full description on the dataset page: https://huggingface.co/datasets/pdelobelle/nemotron-dutch-mt.
