datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vuurwerkverkenner-development-data
NFI Fireworks Development Dataset for the "Vuurwerkverkenner" Application
The Netherlands Forensic Institute (NFI) Fireworks development dataset consists of scans of fireworks wrappers from
fireworks that were investigated in casework in the Netherlands from 2010 onwards. Artificially created snippets
are available for all wrappers, and for a subset of the wrappers photographs of actual fireworks snippets (pieces of the
wrapper post-detonation) are included.
Data… See the full description on the dataset page: https://huggingface.co/datasets/NetherlandsForensicInstitute/vuurwerkverkenner-development-data.rhizomorphic-networks-data
Rhizomorphic Networks — Data and Analysis Outputs
This dataset contains the experimental image data, segmentation outputs, and
downstream analysis results used for the quantitative analysis of
Armillaria gallica rhizomorphic networks.
The directory structure is organised according to the main stages of the analysis
pipeline:
Rhizomorphic Networks/
│
├── 01_raw inputs/
│ ├── control/
│ ├── furnace/
│ └── nutrient density/
│
├── 02_segmentation outputs/
│ ├── raw… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/rhizomorphic-networks-data.netjuunosusume
Bangumi Image Base of Net-juu No Susume
This is the image base of bangumi Net-juu No Susume, we detected 40 characters, 4334 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/netjuunosusume.sri-lankan-album-stamp-detection
Sri Lankan Album Stamp Detection Dataset
This dataset contains annotated Sri Lankan stamp album page images prepared for stamp detection using YOLO-based object detection models.
It is used for the Stamp AI Model 1 stamp detector, where the goal is to detect individual stamps from full album page images.
Contents
raw_pages/ — original album page images
yolo_dataset/ — YOLO-formatted detection dataset
The YOLO dataset contains image files, label files, and a… See the full description on the dataset page: https://huggingface.co/datasets/nethsith/sri-lankan-album-stamp-detection.sri-lankan-stamp-matcher-dataset
Sri Lankan Stamp Matcher Dataset
Source data for the Stamp AI Model 2 stamp-matching project.
Contents
reviewed_crops/ — reviewed single-stamp crop images
metadata-final.xlsx — master metadata file
The Excel file contains the review status, exact stamp grouping, design-family grouping, face value, overprint details, condition, and capture-quality notes.
accepted, pending, and excluded statuses are retained so the dataset can be reviewed and regenerated… See the full description on the dataset page: https://huggingface.co/datasets/nethsith/sri-lankan-stamp-matcher-dataset.tcga-ov-multiomics-network-derived-results
TCGA-OV Multiomics Network Derived Results
This dataset contains derived, publication-ready outputs from a reproducible TCGA-OV multi-omics network analysis pipeline.
Current status
Primary manuscript target: Journal of Biomedical Informatics
Preferred bundle: manuscript/journal_of_biomedical_informatics/
Current JBI main manuscript status:
required statement-of-significance table included
main-paper combined tables/figures reduced to a compliant <=8
sequential in-text… See the full description on the dataset page: https://huggingface.co/datasets/hssling/tcga-ov-multiomics-network-derived-results.Network_Defense_Symmetric_Competitive102,400,000 timesteps, Multi-Agent Reinforcement Learning
Total Environment Steps= 10 parallel environments × 7,000 episodes ×2,048 steps= 102400000 Timesteps
-The Red Agent’s goal is to discover vulnerabilities, elevate privileges, compromise assets, and maintain persistence. Its action space can be modeled after phases of the
MITRE ATT&CK framework.
-The Blue Agent’s goal is to maintain system availability, reduce the attack surface, detect malicious behavior… See the full description on the dataset page: https://huggingface.co/datasets/privateboss/Network_Defense_Symmetric_Competitive.Network_Defense_Symmetric_Competitive102,400,000 timesteps, Multi-Agent Reinforcement Learning
Total Environment Steps= 10 parallel environments × 7,000 episodes ×2,048 steps= 102400000 Training Timesteps
-The Red Agent’s goal is to discover vulnerabilities, elevate privileges, compromise assets, and maintain persistence. Its action space can be modeled after phases of the
MITRE ATT&CK framework.
-The Blue Agent’s goal is to maintain system availability, reduce the attack surface, detect malicious… See the full description on the dataset page: https://huggingface.co/datasets/TorontoMetropolitanUniversity/Network_Defense_Symmetric_Competitive.NFI_FARED_IMUThis is the README file for the dataset Netherlands Forensic Institute: Forensic Activity Recognition Dataset (NFI_FARED), published as a part of the paper "Hi-OSCAR: Hierarchical Open-set Classifier for Human Activity Recognition.". Two forms of data were collected: Digital Traces from iPhones worn on the subjects' bodies, and raw sensor signals from body-worn Inertial Measurement Units (IMUs). This dataset and README refers to the IMU data. The Digital Trace data is available here.
NFI_FARED… See the full description on the dataset page: https://huggingface.co/datasets/NetherlandsForensicInstitute/NFI_FARED_IMU.netherlands-license-plate-dataset
License Plate Recognition Dataset
Dataset contains 73,400+ high-resolution images of vehicle license plates captured across diverse real-world conditions in the Netherlands. Designed to advance research and development in automatic license plate recognition (ALPR), optical character recognition (OCR), and intelligent traffic management systems.
With this dataset, researchers and developers can build and refine solutions for traffic analysis, transportation management, and… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/netherlands-license-plate-dataset.netflix_datasetContinuous-Ethnicity-Face-Recognition
Dataset Card for Ethnicity Fairness in a Continuous Space
Dataset Details
Dataset Description
This dataset provides the training images (from BalancedFace and GlobalFace produced by BUPT) used in the paper: "Balancing Beyond Discrete Categories: Continuous Demographic Labels for Fair Face Recognition".
These have been curated to be balanced in a continuous ethnicity space, following three different strategies: Protocol A, Protocol B and Protocol C.… See the full description on the dataset page: https://huggingface.co/datasets/netopedro/Continuous-Ethnicity-Face-Recognition.ROVR-Open-Dataset
ROVR Open Dataset
Introduction
Welcome to the ROVR Open Dataset repository! This dataset is designed to empower autonomous driving and robotics research by providing rich, real-world data captured from ADAS cameras and LiDAR sensors. The dataset spans 50+ countries with over 20 million kilometers of driving data, making it ideal for training and developing advanced AI algorithms for depth estimation, object detection, and semantic segmentation.… See the full description on the dataset page: https://huggingface.co/datasets/ROVR-Network/ROVR-Open-Dataset.3danimation-datasetvuurwerkverkenner-application-data
Vuurwerkverkenner
This dataset is utilized by the Vuurwerkverkenner application to link fragments of exploded (heavy) fireworks to their
originating firework types. You can explore the application
at www.vuurwerkverkenner.nl. The dataset includes various firework types examined in
casework by the Netherlands Forensic Institute.
Categories
Firework wrappers that closely resemble each other visually may be grouped into categories. Typically, a wrapper stands
alone… See the full description on the dataset page: https://huggingface.co/datasets/NetherlandsForensicInstitute/vuurwerkverkenner-application-data.ART-Netdemo
TreeOfLife-200M Embeddings
Work in progress. This dataset and card are under active development.
Pre-computed image embeddings for 200M+ images from the TreeOfLife-200M dataset, sorted by taxonomic hierarchy for efficient filtered access.
This repository hosts embedding configs for TreeOfLife-200M. Each config corresponds to a different embedding model and/or precision. Currently available: BioCLIP 2 (float16). Additional configs (e.g., BioCLIP 2.5 Huge) will be added as new… See the full description on the dataset page: https://huggingface.co/datasets/netzhang/demo.dream-network-environment-cards
Dream Network — Environment Location Cards (20)
20 fully-annotated environment reference cards — the locations of the Dream Network, each a self-contained 1:1 card: a cinematic establishing view plus alternate views and a full worldbuilding stat panel rendered into the image.
Companion to dream-network-player-cards (the characters) and unhinged-cast-20 (their turnaround sheets). Together they form a complete production bible: who the characters are, what they look like, and… See the full description on the dataset page: https://huggingface.co/datasets/TheMindExpansionNetwork/dream-network-environment-cards.dream-network-player-cards
Dream Network — Character Player Cards (20)
20 fully-annotated character reference cards for the Dream Network cast. Each 1:1 card is a complete, self-contained character reference: full-body turnaround, expression lineup, and a personality stat panel — all rendered into the image itself.
This is the companion to unhinged-cast-20 (which holds the multi-view turnaround sheets). Those sheets show what they look like; these cards explain who they are.
What every card… See the full description on the dataset page: https://huggingface.co/datasets/TheMindExpansionNetwork/dream-network-player-cards.mushroom-net
Mushroom Net: Large-Scale Normalized CLIP-Style Dataset
Overview
Mushroom Net is a comprehensive, human-curated dataset of mushroom images extracted from 53 unique source archives (~69 GB). It is designed to train high-performance vision models (CLIP, ResNet, ViT) for mushroom classification, edibility prediction, and object detection.
Total Images: 227,230
Resolution: 224x224 RGB (Normalized)
Unique Classes: 2,350+ Normalized Taxonomies
Classification Heads: 4 (Species… See the full description on the dataset page: https://huggingface.co/datasets/ShivRamSaud/mushroom-net.car_msrp_eda_netzer_v2
🚗 Car MSRP Analysis — Exploratory Data Analysis (EDA)
📘 Overview
This project performs an in-depth Exploratory Data Analysis (EDA) on a dataset of cars and their characteristics in order to understand which features most strongly influence the MSRP (Manufacturer Suggested Retail Price).
The work includes:
Data loading and cleaning
Handling missing values
Basic feature engineering
Outlier handling for visualization
Exploratory Data Analysis and visualizations… See the full description on the dataset page: https://huggingface.co/datasets/netzer97/car_msrp_eda_netzer_v2.hftesty
Dataset Card for Dataset Name
Collection of NC DOT Public Traffic Safety Cameras. Gathered to detect icy bridges for traffic safety.
This dataset card based upon this raw template.
Dataset Details
Dataset Description
North Carolina Traffic Safety Cameras. Gathered to detect icy bridges for traffic safety. Weather data from openweathermaps.
Curated by: John F. Davis davisjf@gmail.com
Funded by [optional]: John F. Davis
Shared by [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/netskink/hftesty.nettleTeleTables
TeleTables
A Benchmark for Large Language Models in Telecom Table Interpretation
Developed by the NetOp Team, Huawei Paris Research Center
📄 Read the Paper
🤗 Explore the Dataset
TeleTables is a benchmark designed to evaluate the ability of Large Language Models (LLMs) to interpret, reason over, and extract information from technical tables in the telecommunications domain.
A large portion of information in… See the full description on the dataset page: https://huggingface.co/datasets/netop/TeleTables.netflix-clone-imagesdream-network-prop-cards
Dream Network — Artifact & Prop Reference Cards (20)
20 fully-annotated weapon, artifact, and prop reference cards from the Dream Network universe. Each card is a self-contained 1:1 collectible item sheet featuring a 3D hero showcase render, inset material/angle details, and structured RPG stat annotations.
Companion to:
dream-network-player-cards (The Characters)
dream-network-environment-cards (The Locations)
unhinged-cast-20 (The Original Character Turnarounds)
Together… See the full description on the dataset page: https://huggingface.co/datasets/TheMindExpansionNetwork/dream-network-prop-cards.GUI-Net-NanoFor debugging purpose of training TongUI
Shoe-Net-10K
Shoe-Net-10K Dataset
The Shoe-Net-10K dataset is a curated collection of 10,000 shoe images annotated for multi-class image classification. This dataset is suitable for training deep learning models to recognize different types of shoes from images.
Dataset Details
Total Images: 10,000
Image Size: Varies (typical width range: 94 px to 519 px)
Format: Parquet
Split:
train: 10,000 images
Modality: Image
License: Apache 2.0
Labels
The dataset includes 5… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Shoe-Net-10K.pokemon-cards-image-and-annotationsADA-Net-dataset
Attention-Guided Domain Adaptation Network (ADA-Net)
This repository shares the data for ADA-Net: Attention-Guided Domain Adaptation Network with Contrastive Learning for Standing Dead Tree Segmentation Using Aerial Imagery and includes the annotated dataset for mapping standing dead trees. ADA-Nets are generic networks and they can be used in different domation adaptation and Image-to-Image translation problems. In this repository, we specifically focus on transforming… See the full description on the dataset page: https://huggingface.co/datasets/meteahishali/ADA-Net-dataset.
