CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NetherlandsForensicInstitute /vuurwerkverkenner-development-data NFI Fireworks Development Dataset for the "Vuurwerkverkenner" Application The Netherlands Forensic Institute (NFI) Fireworks development dataset consists of scans of fireworks wrappers from fireworks that were investigated in casework in the Netherlands from 2010 onwards. Artificially created snippets are available for all wrappers, and for a subset of the wrappers photographs of actual fireworks snippets (pieces of the wrapper post-detonation) are included. Data… See the full description on the dataset page: https://huggingface.co/datasets/NetherlandsForensicInstitute/vuurwerkverkenner-development-data.image10K<n<100K0 likes1.9k downloads3mo agoHugging Face02NetherlandsForensicInstitute /DNANet_2p5pMixture_PPF6C_2024 2p5p Mixture DNA Research dataset This dataset repository contains RFU signal reading (.hid) files and their corresponding person mixture labels (.txt) used in the DNANet paper and code. The data consist of DNA sample mixtures of 2 to 5 persons, of which the mixtures composition is known, allowing for training on actual ground-truth data for DNA annotation tools such as DNANet If you use this dataset in your research please cite it appropriately: @ARTICLE{Benschop2019, title… See the full description on the dataset page: https://huggingface.co/datasets/NetherlandsForensicInstitute/DNANet_2p5pMixture_PPF6C_2024.textn<1K1 likes1.7k downloads5mo agoHugging Face03NetherlandsForensicInstitute /s2orc-citation-pairs-translated-nlThis is a Dutch version of the S2ORC: The Semantic Scholar Open Research Corpus. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. textsentence-similarity10M<n<100M0 likes260 downloads2y agoHugging Face04EinzzCookie /netherlands-trending-tiktok-videos-verified-users-tiktok [Region] Trending TikTok Videos – Verified Users Dataset name: EinzzCookie/[region]-trending-tiktok-videos-verified-users-tiktok This dataset contains metadata of currently trending TikTok videos posted by verified accounts in the [Region] region. Source Collected by EinzzCookie / TikTrackTelegram: @einzzcookie If you use this data, please credit the original source:https://einzzcookie.org/ License MIT Contents The dataset typically… See the full description on the dataset page: https://huggingface.co/datasets/EinzzCookie/netherlands-trending-tiktok-videos-verified-users-tiktok.other10K<n<100K1 likes208 downloads20d agoHugging Face05justicedao /ipfs_netherlands_laws_ir Netherlands legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_netherlands_laws (revision 659c8fa0db188d5dd9624dff49474065c9cd3f6e) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of Netherlands prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No legal… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_netherlands_laws_ir.tabulartext-retrieval1M<n<10M0 likes206 downloads2d agoHugging Face06NetherlandsForensicInstitute /NFI_FARED_IMUThis is the README file for the dataset Netherlands Forensic Institute: Forensic Activity Recognition Dataset (NFI_FARED), published as a part of the paper "Hi-OSCAR: Hierarchical Open-set Classifier for Human Activity Recognition.". Two forms of data were collected: Digital Traces from iPhones worn on the subjects' bodies, and raw sensor signals from body-worn Inertial Measurement Units (IMUs). This dataset and README refers to the IMU data. The Digital Trace data is available here. NFI_FARED… See the full description on the dataset page: https://huggingface.co/datasets/NetherlandsForensicInstitute/NFI_FARED_IMU.image10M<n<100M1 likes197 downloads9mo agoHugging Face07NetherlandsForensicInstitute /vuurwerkverkenner-application-data Vuurwerkverkenner This dataset is utilized by the Vuurwerkverkenner application to link fragments of exploded (heavy) fireworks to their originating firework types. You can explore the application at www.vuurwerkverkenner.nl. The dataset includes various firework types examined in casework by the Netherlands Forensic Institute. Categories Firework wrappers that closely resemble each other visually may be grouped into categories. Typically, a wrapper stands alone… See the full description on the dataset page: https://huggingface.co/datasets/NetherlandsForensicInstitute/vuurwerkverkenner-application-data.image1K<n<10K1 likes186 downloads3mo agoHugging Face08UniDataPro /netherlands-license-plate-dataset License Plate Recognition Dataset Dataset contains 73,400+ high-resolution images of vehicle license plates captured across diverse real-world conditions in the Netherlands. Designed to advance research and development in automatic license plate recognition (ALPR), optical character recognition (OCR), and intelligent traffic management systems. With this dataset, researchers and developers can build and refine solutions for traffic analysis, transportation management, and… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/netherlands-license-plate-dataset.imageobject-detectionn<1K1 likes172 downloads10d agoHugging Face09vGassen /Dutch-Judiciary-Court-Cases-Netherlands-Rechtspraak0 likes132 downloads1y agoHugging Face10vGassen /Dutch-Judiciary-Court-Cases-Netherlands-Rechtspraak-Vector-V3text100K<n<1M0 likes111 downloads1y agoHugging Face11NetherlandsForensicInstitute /rppg-deepfake-detectionvideon<1K0 likes89 downloads28d agoHugging Face12endomorphosis /ipfs_netherlands_laws Netherlands In-Force National Law Corpus (BWB) Research snapshot of in-force national legislation from the Dutch Basiswettenbestand (BWB), published via KOOP / wetten.overheid.nl. Not legal advice. The official gazette (Staatsblad / authentic source) prevails over this corpus. Snapshot Field Value Snapshot date 2026-09-02 Source BWB (KOOP) Collector scrapers/collect_bwb.py In-force national instruments 18,626 Language Dutch (nl) Jurisdiction… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_netherlands_laws.texttext-retrieval10K<n<100K0 likes75 downloads22d agoHugging Face13justicedao /wetwijzer_netherlands_legal_corpus WetWijzer Netherlands Legal Corpus Hugging Face target: justicedao/wetwijzer_netherlands_legal_corpus. This unified dataset bundles the quality-audited WetWijzer Netherlands legal corpus stack in one repository for frontend retrieval. It preserves the existing compatibility repositories and does not replace or delete them. Contents Laws: 4,999 Articles: 89,737 CID index rows: 94,736 Vector mapping rows: 94,736 BM25 document rows: 94,736 BM25 term rows: 120,521… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/wetwijzer_netherlands_legal_corpus.tabulartext-retrieval100K<n<1M0 likes66 downloads3mo agoHugging Face14Anatolii2026 /netherlands-house-prices-by-city-funda-2026 Netherlands House Prices by City — Funda.nl, August 2026 Median asking prices for homes on sale in the Netherlands, by city, collected on 31 August 2026. What is inside File Rows What it is funda-nl-market-summary-2026-08-31.csv 67 One row per city with at least 5 listings: median price, median price per m², median living space, median rooms funda-nl-schema-sample.csv 59 A small sample of individual listings, to show the field structure… See the full description on the dataset page: https://huggingface.co/datasets/Anatolii2026/netherlands-house-prices-by-city-funda-2026.n<1K0 likes51 downloads23d agoHugging Face15Elpriser /netherlands-power-market Netherlands Power Market Data Part of a country-split collection of historical electricity market data gathered for the elpriser.org price-forecasting project. See the index dataset for the other countries: Denmark, Germany, Norway, Sweden, Finland, Netherlands. License & attribution CC BY 4.0. Source: ENTSO-E Transparency Platform (transparency.entsoe.eu). Contents NL bidding zone, ~2018-09/10 → present: File Description… See the full description on the dataset page: https://huggingface.co/datasets/Elpriser/netherlands-power-market.text1M<n<10M0 likes50 downloads3mo agoHugging Face16NetherlandsForensicInstitute /NFI_FARED_Digital_TracesThis is the README file for the dataset Netherlands Forensic Institute: Forensic Activity Recognition Dataset (NFI_FARED). Two forms of data were collected: Digital Traces from iPhones worn on the subjects' bodies, and raw sensor signals from body-worn Inertial Measurement Units (IMUs). This dataset and README refers to the Digital Trace data. The IMU data is available here. Published as a part of the paper "Forensic Activity Classification Using Digital Traces from iPhones: A Machine… See the full description on the dataset page: https://huggingface.co/datasets/NetherlandsForensicInstitute/NFI_FARED_Digital_Traces.0 likes44 downloads10mo agoHugging Face17NetherlandsForensicInstitute /wiki-atomic-edits-translated-nlThis is a Dutch version of the Wiki Atomic Edits dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. textsentence-similarity10M<n<100M1 likes43 downloads2y agoHugging Face18NetherlandsForensicInstitute /wikipedia-questions Dutch Synthetic Questions for Wikipedia Articles A selection of synthetically generated questions and keywords for (chunks of) Wikipedia articles. This dataset can be used to train sentence embedding models. Source dataset The dataset is based on the wikimedia/wikipedia dataset, 20231101.nl subset. Recipe Generation was done using the following general recipe: Filter out short articles (<768 characters) to remove many automatically generated stubs. Split up… See the full description on the dataset page: https://huggingface.co/datasets/NetherlandsForensicInstitute/wikipedia-questions.textfeature-extraction10M<n<100M3 likes41 downloads1y agoHugging Face19justicedao /ipfs_netherlands_laws IPFS Netherlands Laws Hugging Face target: justicedao/ipfs_netherlands_laws. This dataset packages Netherlands law records with deterministic IPFS Content IDs. Each row includes a cid and content_address; article rows also include the parent law_cid. This is a quality-audited catalog-backed Netherlands snapshot from official Dutch government sources. It is not the full Dutch legal corpus: the persistent catalog contains 42,956 discovered BWBR identifiers, of which 5,000 are… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_netherlands_laws.tabulartext-retrieval100K<n<1M0 likes38 downloads3mo agoHugging Face20justicedao /ipfs_netherlands_laws_knowledge_graph IPFS Netherlands Laws Knowledge Graph Hugging Face target: justicedao/ipfs_netherlands_laws_knowledge_graph. JSON-LD graph and node/edge tables whose identities are IPFS content addresses. This graph currently has 94736 nodes and 89737 edges from the paired CID dataset. Source scrape date: 2026-06-27T14:04:23.964996. Full BWB discovery inventory found 42,956 unique BWBR identifiers from official SRU discovery; this paired snapshot contains 5,000 completed identifiers and must… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_netherlands_laws_knowledge_graph.text100K<n<1M0 likes37 downloads3mo agoHugging Face21NetherlandsForensicInstitute /flickr30k-captions-translated-nlThis is a Dutch version of the Flickr30k captions dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. For more information about the use of this dataset please refer to the flicker terms of use textsentence-similarity100K<n<1M0 likes35 downloads3y agoHugging Face22VillaLabs /hello_netherlands_5000 likes35 downloads1y agoHugging Face23NetherlandsForensicInstitute /simplewiki-translated-nlThis is a Dutch version of the SimpleWiki text simplification dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. textsentence-similarity100K<n<1M0 likes32 downloads2y agoHugging Face24NetherlandsForensicInstitute /stackexchange-duplicate-questions-translated-nlThis is a Dutch version of the Stackexchange duplicate questions dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. textsentence-similarity100K<n<1M0 likes32 downloads2y agoHugging Face25NetherlandsForensicInstitute /allnli-translated-nlThe AIINLI dataset is a combination of the SNLI and the MultiNLI corpora. Which we have auto-translated into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. textsentence-similarity100K<n<1M0 likes31 downloads2y agoHugging Face26justicedao /netherlands-laws-nl-normalized Netherlands Laws (Dutch, Normalized) Hugging Face target: justicedao/netherlands-laws-nl-normalized. This package is a normalized version of the Netherlands laws scrape output. This is a capped Netherlands scrape, not the full Dutch corpus. The scrape used max_documents=100, parsed 151 law record(s), and discovered 626 unique official BWBR law document(s) before applying the cap. Documents failed: 0. This refresh includes parser coverage improvements for older/French heading… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/netherlands-laws-nl-normalized.tabulartext-retrieval1K<n<10K0 likes30 downloads3mo agoHugging Face27vGassen /Dutch-Disciplinary-Court-Cases-Netherlands-Tuchtrechttext10K<n<100K0 likes29 downloads1y agoHugging Face28vGassen /Dutch-Judiciary-Court-Cases-Netherlands-Rechtspraak-Relevant-Metadata0 likes26 downloads1y agoHugging Face29NetherlandsForensicInstitute /msmarco-translated-nlThis is a Dutch version of the MS MARCO dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. A newer translation of this dataset using LLMs is available at NetherlandsForensicInstitute/msmarco-nl. textsentence-similarity100K<n<1M1 likes24 downloads1y agoHugging Face30NetherlandsForensicInstitute /squad-nl-v2.0 SQuAD-NL v2.0 for Sentence Transformers The SQuAD-NL v2.0 dataset (on Hugging Face: GroNLP/squad-nl-v2.0), modified for use in Sentence Transformers as a dataset of type "Pair with Similarity Score". Score We added an extra column score to the original dataset. The value of score is 1.0 if the question has an answer in the context (no matter where), and 0.0 if there are no answers in the context. The allows the evaluation of embedding models that aim to pair queries… See the full description on the dataset page: https://huggingface.co/datasets/NetherlandsForensicInstitute/squad-nl-v2.0.textsentence-similarity100K<n<1M2 likes24 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.