datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NL-DIR
Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark (CVPR 2025)
Dataset Introduction
NL-DIR consists of 41,795 document images with a diverse range of document types, and each image corresponds to 5 high-quality fine-grained semantic queries, which are generated and evaluated through large language models in conjunction with manual verification, resulting 200K+ queries in total. Following an 8:1:1 ratio partition, the dataset is divided into… See the full description on the dataset page: https://huggingface.co/datasets/nianbing/NL-DIR.babylm-nld
BabyLM Dataset
Dataset Description
This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: nld
Script: Latn
Tier: 100M
Byte Premium Factor: 1.051606
Size (MB): 569.49
Expected Size (MB): 571.02
Number of Documents: 304,611
Total Tokens: 109,885,564
Tokenizer: separate by whitespace
Tokens Per Category
child-books: 4,576,823 tokens
child-directed-speech: 3,304,756… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-nld.NLDnl_dc_cls
Dataset Card for "nl_dc_cls"
More Information needed
BabyBabel-nld-10Mnld_Latn-sampleBabyBabel-nld-33MBabyBabel-nld-20Mbabylm-eng-nld-50-50-stratified
BabyLM English–Dutch 50/50 Stratified
A bilingual training corpus for the Multilingual BabyLM project, combining 50% of the English and Dutch BabyBabelLM corpora.
Sampling is stratified by category: each source category is sampled independently to preserve the original category proportions within each language.
Token counts
Language
Tokens
Share
English (eng)
49,481,353
47.4%
Dutch (nld)
54,953,487
52.6%
Total104,434,840
100%
Category… See the full description on the dataset page: https://huggingface.co/datasets/adzcai/babylm-eng-nld-50-50-stratified.BabyBabel-nld-50MBabyBabel-nld-held-outfinepdfs-nld-stats
Statistics for HuggingFaceFW/finepdfs-edu (Dutch (Latin script))
Aggregate statistics computed using Polars streaming on the HuggingFaceFW/finepdfs-edu dataset.
Performance
Processed 847,361 documents in 125.88 seconds.
Step
Time
Global stats
34.19s
Language stats
31.81s
Extractor stats
29.46s
Dump stats
30.42s
Total
125.88s
Speed comes from Polars only reading metadata columns (not the text column),
thanks to Parquet's columnar format and lazy… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/finepdfs-nld-stats.nl-de_top_cs_devBabyBabel-nld-60MBabyBabel-nld-30MBabyBabel-nld-66Mnl-de_non_top_cs_trainBabyBabel-nld-40Mnldv
Dataset Card for "nldv"
More Information needed
nl-de_non_top_cs_devcommonvoice_male_1_nldBabyBabel-nld-80Mstory-summarization-demonl-de_top_cs_trainsentiment-demoThis is a demo dataset...
Sentiment Demo
Installation
...
Usage
...
License
...
Contributions
...
BabyBabel-nld-70MBabyBabel-nld-100Mgoldfish-nld-latn-100mblangid-nld_Latncommonvoice_male_2_nld
