datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DNANet_2p5pMixture_PPF6C_2024
2p5p Mixture DNA Research dataset
This dataset repository contains RFU signal reading (.hid) files and their corresponding person mixture labels (.txt) used in the DNANet paper and code.
The data consist of DNA sample mixtures of 2 to 5 persons, of which the mixtures composition is known,
allowing for training on actual ground-truth data for DNA annotation tools such as DNANet
If you use this dataset in your research please cite it appropriately:
@ARTICLE{Benschop2019,
title… See the full description on the dataset page: https://huggingface.co/datasets/NetherlandsForensicInstitute/DNANet_2p5pMixture_PPF6C_2024.s2orc-citation-pairs-translated-nlThis is a Dutch version of the S2ORC: The Semantic Scholar Open Research Corpus. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
ipfs_netherlands_laws_ir
Netherlands legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_netherlands_laws (revision 659c8fa0db188d5dd9624dff49474065c9cd3f6e) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Netherlands prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_netherlands_laws_ir.NFI_FARED_IMUThis is the README file for the dataset Netherlands Forensic Institute: Forensic Activity Recognition Dataset (NFI_FARED), published as a part of the paper "Hi-OSCAR: Hierarchical Open-set Classifier for Human Activity Recognition.". Two forms of data were collected: Digital Traces from iPhones worn on the subjects' bodies, and raw sensor signals from body-worn Inertial Measurement Units (IMUs). This dataset and README refers to the IMU data. The Digital Trace data is available here.
NFI_FARED… See the full description on the dataset page: https://huggingface.co/datasets/NetherlandsForensicInstitute/NFI_FARED_IMU.Dutch-Judiciary-Court-Cases-Netherlands-Rechtspraak-Vector-V3ipfs_netherlands_laws
Netherlands In-Force National Law Corpus (BWB)
Research snapshot of in-force national legislation from the Dutch
Basiswettenbestand (BWB), published via KOOP / wetten.overheid.nl.
Not legal advice. The official gazette (Staatsblad / authentic source)
prevails over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-02
Source
BWB (KOOP)
Collector
scrapers/collect_bwb.py
In-force national instruments
18,626
Language
Dutch (nl)
Jurisdiction… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_netherlands_laws.wetwijzer_netherlands_legal_corpus
WetWijzer Netherlands Legal Corpus
Hugging Face target: justicedao/wetwijzer_netherlands_legal_corpus.
This unified dataset bundles the quality-audited WetWijzer Netherlands legal corpus stack in one repository for frontend retrieval. It preserves the existing compatibility repositories and does not replace or delete them.
Contents
Laws: 4,999
Articles: 89,737
CID index rows: 94,736
Vector mapping rows: 94,736
BM25 document rows: 94,736
BM25 term rows: 120,521… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/wetwijzer_netherlands_legal_corpus.netherlands-power-market
Netherlands Power Market Data
Part of a country-split collection of historical electricity market data gathered for the elpriser.org price-forecasting project. See the index dataset for the other countries: Denmark, Germany, Norway, Sweden, Finland, Netherlands.
License & attribution
CC BY 4.0. Source: ENTSO-E Transparency Platform (transparency.entsoe.eu).
Contents
NL bidding zone, ~2018-09/10 → present:
File
Description… See the full description on the dataset page: https://huggingface.co/datasets/Elpriser/netherlands-power-market.wiki-atomic-edits-translated-nlThis is a Dutch version of the Wiki Atomic Edits dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
wikipedia-questions
Dutch Synthetic Questions for Wikipedia Articles
A selection of synthetically generated questions and keywords for (chunks of) Wikipedia articles.
This dataset can be used to train sentence embedding models.
Source dataset
The dataset is based on the wikimedia/wikipedia dataset, 20231101.nl subset.
Recipe
Generation was done using the following general recipe:
Filter out short articles (<768 characters) to remove many automatically generated stubs.
Split up… See the full description on the dataset page: https://huggingface.co/datasets/NetherlandsForensicInstitute/wikipedia-questions.ipfs_netherlands_laws
IPFS Netherlands Laws
Hugging Face target: justicedao/ipfs_netherlands_laws.
This dataset packages Netherlands law records with deterministic IPFS Content IDs. Each row includes a cid and content_address; article rows also include the parent law_cid.
This is a quality-audited catalog-backed Netherlands snapshot from official Dutch government sources. It is not the full Dutch legal corpus: the persistent catalog contains 42,956 discovered BWBR identifiers, of which 5,000 are… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_netherlands_laws.ipfs_netherlands_laws_knowledge_graph
IPFS Netherlands Laws Knowledge Graph
Hugging Face target: justicedao/ipfs_netherlands_laws_knowledge_graph.
JSON-LD graph and node/edge tables whose identities are IPFS content addresses.
This graph currently has 94736 nodes and 89737 edges from the paired CID dataset. Source scrape date: 2026-06-27T14:04:23.964996. Full BWB discovery inventory found 42,956 unique BWBR identifiers from official SRU discovery; this paired snapshot contains 5,000 completed identifiers and must… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_netherlands_laws_knowledge_graph.flickr30k-captions-translated-nlThis is a Dutch version of the Flickr30k captions dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. For more information about the use of this dataset please refer to the flicker terms of use
simplewiki-translated-nlThis is a Dutch version of the SimpleWiki text simplification dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
stackexchange-duplicate-questions-translated-nlThis is a Dutch version of the Stackexchange duplicate questions dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
allnli-translated-nlThe AIINLI dataset is a combination of the SNLI and the MultiNLI corpora. Which we have auto-translated
into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
netherlands-laws-nl-normalized
Netherlands Laws (Dutch, Normalized)
Hugging Face target: justicedao/netherlands-laws-nl-normalized.
This package is a normalized version of the Netherlands laws scrape output.
This is a capped Netherlands scrape, not the full Dutch corpus. The scrape used max_documents=100, parsed 151 law record(s), and discovered 626 unique official BWBR law document(s) before applying the cap. Documents failed: 0.
This refresh includes parser coverage improvements for older/French heading… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/netherlands-laws-nl-normalized.Dutch-Disciplinary-Court-Cases-Netherlands-Tuchtrechtmsmarco-translated-nlThis is a Dutch version of the MS MARCO dataset.
Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model,
specifically the huggingface implementation.
A newer translation of this dataset using LLMs is available at NetherlandsForensicInstitute/msmarco-nl.
squad-nl-v2.0
SQuAD-NL v2.0 for Sentence Transformers
The SQuAD-NL v2.0 dataset (on Hugging Face: GroNLP/squad-nl-v2.0), modified for use in Sentence Transformers as a dataset of type "Pair with Similarity Score".
Score
We added an extra column score to the original dataset.
The value of score is 1.0 if the question has an answer in the context (no matter where), and 0.0 if there are no answers in the context.
The allows the evaluation of embedding models that aim to pair queries… See the full description on the dataset page: https://huggingface.co/datasets/NetherlandsForensicInstitute/squad-nl-v2.0.quora-duplicates-translated-nlThis is a Dutch version of the Quora Duplicates dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. For more information about the use of this dataset please refer to the Quora Terms of Service.
altlex-translated-nlThis is a Dutch version of the AltLex dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
Mr.Porter.Product.prices.Netherlands
Mr Porter web scraped data
About the website
The fashion retail industry in EMEA, specifically in the Netherlands, is a well-established and diverse marketplace encompassing a broad spectrum of brands and products. The ever-evolving landscape has witnessed a shift from traditional high-street stores to online platforms like Mr Porter, a trend accelerated by the pandemic. A vital player in the menswear luxury online retailer sector, Mr Porter has paved the way for… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Mr.Porter.Product.prices.Netherlands.ipfs_netherlands_laws_bm25_index
IPFS Netherlands Laws BM25 Index
Hugging Face target: justicedao/ipfs_netherlands_laws_bm25_index.
Sparse BM25 document and postings tables keyed by source CID.
This index covers 94736 documents and 120521 terms from the paired CID dataset. Source scrape date: 2026-06-27T14:04:23.964996. Full BWB discovery inventory found 42,956 unique BWBR identifiers from official SRU discovery; this paired snapshot contains 5,000 completed identifiers and must not be described as the full… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_netherlands_laws_bm25_index.sentence-compression-translated-nlThis is a Dutch version of the Sentence Compression dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
coco-captions-translated-nlThis is a Dutch version of the Coco captions dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
ipfs_netherlands_laws_vector_index
IPFS Netherlands Laws Vector Index
Hugging Face target: justicedao/ipfs_netherlands_laws_vector_index.
Dense vector mapping keyed by source CID, with FAISS and TF-IDF/SVD artifacts.
This index covers 94736 rows from the paired CID dataset. Source scrape date: 2026-06-27T14:04:23.964996. Full BWB discovery inventory found 42,956 unique BWBR identifiers from official SRU discovery; this paired snapshot contains 5,000 completed identifiers and must not be described as the full… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_netherlands_laws_vector_index.MajorTom-Netherlandshello_netherlandsmsmarco-nl
MS MARCO NL
This is a machine translation of the MS MARCO dataset.
This dataset can be used to train sentence embedding models.
In contrast to our previous translation,
an LLM (GPT-4o mini) was used for the translation.
This results in generally higher translation quality.
Source dataset
The dataset is based on the MS MARCO dataset.
Model
We used a deployment of GPT-4o mini using the Microsoft Azure OpenAI APIs.
Prompt
The following… See the full description on the dataset page: https://huggingface.co/datasets/NetherlandsForensicInstitute/msmarco-nl.
