CoolFace
Datasetpublic

pdelobelle/staatsblad-synth-nl

Synthetic Dutch from the Belgisch Staatsblad Diverse, fluent Dutch pretraining text synthesized from guust-franssens/belgisch-staatsblad (CC0, Belgian official-gazette filings). Adds Belgium/Flanders coverage to Dutch LM pretraining mixes, where clean Belgian-Dutch prose is otherwise scarce. The source text is noisy OCR from scanned PDFs, but its metadata (company, juridical form, act type, city, date) is clean. A local LLM (google/gemma-2-9b-it) "launders" the OCR + metadata… See the full description on the dataset page: https://huggingface.co/datasets/pdelobelle/staatsblad-synth-nl.

sourceHugging Facecc0-1.0updated 2mo agoView on Hugging Face
0likes478downloads
2 commits on main
c0c9e842mo ago

Add synthetic Dutch from Belgisch Staatsblad (~49M tokens, 185640 texts)

pdelobelle
bac35d12mo ago

initial commit

pdelobelle