datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KiTS23-sliced-2dmvsplat_dl3dv_2dMode_1Sub2D_FalseAlphaPreGen-NavierStokes-2D
PreGen Navier-Stokes 2D Dataset
Paper | Project Page | Code
Dataset Description
This dataset accompanies the research paper Pre-Generating Multi-Difficulty PDE Data For Few-Shot Neural PDE Solvers (under review at ICLR 2026). It contains systematically generated 2D incompressible Navier-Stokes fluid flow simulations designed to study difficulty transfer in neural PDE solvers.The key insight: by pre-generating many low and medium difficulty examples and including them… See the full description on the dataset page: https://huggingface.co/datasets/sage-lab/PreGen-NavierStokes-2D.transplat_dl3dv_2dMode_1Sub2D_FalseAlphadl3dv_2dMode_1Sub2D_TrueAlpha_fromRe10KPretrainedqwen35-2d-grounding-sfthipsc_2dSTS-2D-Tooth
STS-2D-Tooth
The 2D panoramic dental X-ray subset of the STS (Semi-supervised Teeth
Segmentation) multi-modal dataset, as released in
Wang et al., Scientific Data 12, 117 (2025)
and used in the MICCAI 2023 STS Challenge.
Composition
4,000 panoramic X-ray images (PNG, 640x320, 3-channel grayscale-as-RGB) split
across two demographic subsets:
Subset
Total
Labeled
Unlabeled
A-PXI (adult)
3,500
850
2,650
C-PXI (child)
500
50
450
Total
4,000
900
3,100… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/STS-2D-Tooth.kitti-2d-detection-lance
KITTI 2D Object Detection (Lance Format)
A Lance-formatted version of the KITTI 2D Object Detection benchmark, sourced from nateraw/kitti so no manual signup or download from cvlibs.net is required. Each row is a single driving frame with inline JPEG bytes, the full set of 2D and 3D object annotations stored as parallel per-object lists, plus a cosine-normalized OpenCLIP ViT-B-32 image embedding — all available directly from the Hub at… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/kitti-2d-detection-lance.2D_Video_Game_Cartoon_Character_Sprite-Sheets
Dataset Card for Dataset Name
Dataset Details
Experimental composition of 76 cartoon art-style video game character spritesheets. Resized to 512x512, mixed variation of animation styles.
Dataset Description
All images editted using Tiled image editting software as most assets are typically downloaded individually and not in sequence. I compiled each animation sequence into one img to display animations frame-by-frame evenly distributed across some common… See the full description on the dataset page: https://huggingface.co/datasets/mgane/2D_Video_Game_Cartoon_Character_Sprite-Sheets.cellmap-2d
CellMap 2D
This dataset contains all 2D slices from the EM volumes used in the CellMap segmentation challenge.
The dataset contains all x, y, z slices obtained from a total of 289 3D EM volume crops (the crops come from 22 different samples),
together with their corresponding labeled segmentation masks.
The slices were prepared and pushed to the HF datasets Hub with
this script.
You can load the dataset as follows (non-streaming mode):
ds =… See the full description on the dataset page: https://huggingface.co/datasets/eminorhan/cellmap-2d.pompomcork_YoloSegment-2D-to-3D-RebotARM
PompomCork RGB-D Tabletop Segmentation (2D-to-3D)
Instance-segmentation dataset of small tabletop objects (cork, pompom,
lighter) captured with an Intel RealSense camera. It is the data behind an
RGB-D perception and grasping pipeline: 2D YOLO instance segmentation, lifting
masks to 3D, and base-frame pose estimation for robot picking.
Code, training and the full pipeline (ROS 2 / MoveIt):… See the full description on the dataset page: https://huggingface.co/datasets/ddt1992/pompomcork_YoloSegment-2D-to-3D-RebotARM.transplat_dl3dv_2dMode_1Sub2D_TrueAlphaObjaverse_2d_renders
2D Image/Depth Rendering of Objaverse Dataset
In total, the rendered split contains 167,857 objects. The object ids are in the completed_renders.txt file. After unzipping, the image/depth renders are in the following folder strunture:
# e.g.,
000-000/000074a334c541878360457c672b6c2e
├── depth.zip
├── image.zip
├── metadata.json
└── transforms_train.json
Camera Intrinsics
72 views per-object, uniformly sampled on the upper hemisphere
Image dimensions: 400×400… See the full description on the dataset page: https://huggingface.co/datasets/ShapeSplats/Objaverse_2d_renders.openorganelle-2d
OpenOrganelle 2D
This dataset contains a large collection of 2D slices from the EM volumes on HHMI Janelia's OpenOrganelle data repository.
The dataset contains a total of ~2.67 million x, y, and z slices obtained from 79 different 3D EM volumes on OpenOrganelle.
Notes
The following volumes on OpenOrganelle are missing from the current repository because they were too large to process and store here:
jrc_fly-larva-1, jrc_fly-mb-z0419-20, jrc_mus-guard-hair-follicle… See the full description on the dataset page: https://huggingface.co/datasets/eminorhan/openorganelle-2d.ZOD-Mini-2D-Road-Scenes
ZOD-Mini-2D-Road-Scenes
The ZOD-Mini-2D-Road-Scenes dataset is derived from the Zenseact Open Dataset (ZOD), property of Zenseact AB (© 2022 Zenseact AB), and is licensed under the permissive CC BY-SA 4.0. Any public use, distribution, or display of this dataset must contain this entire notice:
For this dataset, Zenseact AB has taken all reasonable measures to remove all personally identifiable information, including faces and license plates. To the extent that you like to request… See the full description on the dataset page: https://huggingface.co/datasets/8bits-ai/ZOD-Mini-2D-Road-Scenes.cyp_p450_2d6_inhibition_veith_et_al-multimodalsam3-low-dice-2d-nnunet
SAM3 low-Dice 2D datasets for nnU-Net
Private research export of two small 2D datasets on which the balanced-finish
SAM3 LoRA validation Dice was below 0.5. The purpose is to test whether a
dataset-specific nnU-Net can fit these data and to distinguish data/training
limitations from inference bugs.
Dataset
SAM3 Dice
SAM3 IoU
Evaluated validation images
Actual SAM3 training images
DRIVE
0.212233
0.118717
2
14
RAVIR
0.224709
0.128455
2
16
The two-image validation… See the full description on the dataset page: https://huggingface.co/datasets/MedicalSAM3/sam3-low-dice-2d-nnunet.re10kDl3dv_2dMode_1Sub2D_TrueAlpha_finetune2D_Shape_Image_DatasetsFLARE-MLLM-2D
FLARE25 Medical Multimodal Dataset
This repository contains a multimodal medical imaging dataset for FLARE 2025 with question-answer pairs across various medical imaging modalities.
Dataset Structure
The dataset is organized into the following main directories:
training/: Training data
validation-public/: Public validation data
validation-hidden/: Hidden validation data (answer not released)
testing/: Hidden testing data (not released)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/FLARE-MedFM/FLARE-MLLM-2D.transforms_2d_base
Transforms-2D Base Dataset
This dataset contains foreground objects and background images used by the Transforms-2D dataset in the paper Understanding the Role of Invariance in Transfer Learning, published at TMLR 2024.
The code for the paper is available here, including the implementation of the Transforms-2D dataset.
The Transforms-2D dataset consists of transformed versions of image objects with transparency masks (from this base dataset), pasted onto background images (also from… See the full description on the dataset page: https://huggingface.co/datasets/tillspeicher/transforms_2d_base.2d-geometric-shapes-datasetpacman_2d_extreme_remap_imagined_rollout_5000_cot_v3
Pacman 2D Extreme Interleaved COT v3
Difficulty: extreme.
This dataset is a v3 training view derived from two sources:
5,000 imagined-rollout COT rows from the copied source trajectories.
5,000 disjoint final-stop COT rows from novastar112/pacman_extreme_v0 / pacman_2d_extreme_v0.
Main changes from v1:
The selected imagined-rollout COT conversation is truncated at the COT/action assistant turn.
The assistant-side current/context image after Now I observe... is removed.
Final-stop… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pacman_2d_extreme_remap_imagined_rollout_5000_cot_v3.Elf_encoded_flat_band_materials2D-icons-Dataset2D icons dataset for lora training
dl3dv_2dMode_1Sub2D_FalseAlpha2D-to-3D-groundingclean-kitti-2d-object-detection
Cleaned KITTI Dataset
Overview
This repository provides a cleaned version of the KITTI 2D Object Detection Benchmark. The dataset was curated as part of the research project:
Analyzing Training-Free Corruption Detection for Object Detection Datasets
The goal of the cleaning process was to identify potential annotation inconsistencies using a training-free feature-space based corruption detection approach.
The original KITTI annotation format and folder structure… See the full description on the dataset page: https://huggingface.co/datasets/Chris1095/clean-kitti-2d-object-detection.eval-pack-wjt-of-my-project-ir4dht-2dd66d78
Eval pack wjt of my Project ir4dht
Besoin d'un datapack pour entrainer mon bras robots à saisir des objets transparents
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless… See the full description on the dataset page: https://huggingface.co/datasets/physicl-test/eval-pack-wjt-of-my-project-ir4dht-2dd66d78.
