datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sam3d-flat-20260329-022951ff4d-sam3d-prep
ff4d-sam3d prepped training data (motion324 + syn4d)
STATUS: upload in progress (started 2026-09-07 ~00:20 CDT, ETA ~04:00 CDT).
Files appear a few GB at a time, stream by stream. pointmaps is the largest stream and lands
last for each root. Check the file list for what is already complete; a __NNN.tar that is
present is complete (each is committed only after a successful upload).
Prepped training samples for ff4d-sam3d — video-to-4D on the SAM-3D Objects backbone. Each sample… See the full description on the dataset page: https://huggingface.co/datasets/Testing12321111/ff4d-sam3d-prep.sam3-sgv4-experiment-artifacts
SAM3 SG-v4 experiment artifacts
Private backup of the experiment-specific artifacts under
sam3_sgv4_fullcoverage_20260731.
The original COCO/RefCOCO and ReasonSeg datasets remain in
xuzishan/sam3-datasets. This repository is for generated material that is
not recoverable from those original datasets:
RefCOCO, RefCOCO+, and RefCOCOg four-rollout training trajectories;
thinking histories, trajectory JSONL, predictions, masks, and GT snapshots;
ReasonSeg generated outputs;… See the full description on the dataset page: https://huggingface.co/datasets/xuzishan/sam3-sgv4-experiment-artifacts.libero_10_sam3_visual_prompts
LIBERO-10 SAM3 Visual Prompts
This dataset contains SAM3-generated visual prompts for successful episodes from
fracapuano/libero_10.
It does not contain a trained policy model. The files are per-episode visual
prompt annotations aligned to the original LeRobot episode frames.
Contents
libero_10/chunk-000/episode_*.npz: visual prompt annotations.
manifest_success.jsonl: one row per saved episode with task metadata and hit rates.
videos/episode_*_sam3_vp.mp4: rendered… See the full description on the dataset page: https://huggingface.co/datasets/TechieMoon/libero_10_sam3_visual_prompts.gpic-bcc-sam3-qwen38-27b
GPIC Bidirectional Concept Correspondence Data
This release was generated by ConCor Training Data Generation. Each training example connects a text mask—a set of caption character spans—to an image mask made from one or more segmented instances. Disjoint co-referring spans can therefore share the same correspondence.
The three training configs intentionally match the caption-row format used by UWGZQ/ConCor-1-Data. Our richer pipeline records and complete per-image dispositions… See the full description on the dataset page: https://huggingface.co/datasets/suryadv/gpic-bcc-sam3-qwen38-27b.ff4d-sam3d-prep
ff4d-sam3d prepped training data (motion324 + syn4d)
STATUS: upload complete (2026-09-07 23:01 UTC). Check COMPLETE.txt for the per-stream entry counts.
Files appear a few GB at a time, stream by stream. pointmaps is the largest stream and lands
last for each root. Check the file list for what is already complete; a __NNN.tar that is
present is complete (each is committed only after a successful upload).
Prepped training samples for ff4d-sam3d — video-to-4D on the SAM-3D… See the full description on the dataset page: https://huggingface.co/datasets/ncc2/ff4d-sam3d-prep.sam-3d-body-dataset
SAM-3D-Body Data
This repository provides the annotations used in SAM 3D Body.
Datasets
3DPW
AI Challenger
COCO
EgoExo4D
EgoHumans
Harmony4D
MPII
SA1B
Get Started
Please follow the instructions to download and preocess the annotations.
License
The SAM 3D Body data is licensed under SAM License.
Citing SAM 3D Body
If you use SAM 3D Body or the SAM 3D Body dataset in your research, please use the following BibTeX entry.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/sam-3d-body-dataset.SAM3-Joint-Surgical-Dataset
SAM3 Joint Surgical Dataset
Mehmet Kerem Turkcan, Soham Samal, Zoran Kostic
AIDL Lab, Columbia University,
IPAL Lab
The SAM3 Joint Surgical Dataset contains 41,372 physical images with 175,674 COCO annotations for surgical instruments, anatomy, and tissue. Its schema contains 29 categories and consolidates labels from CholecInstanceSeg, CholecSeg8k, the Dresden Surgical Anatomy Dataset, and Endoscapes. Positive annotations cover 25 categories. The test split… See the full description on the dataset page: https://huggingface.co/datasets/mehmetkeremturkcan/SAM3-Joint-Surgical-Dataset.metaworld_mt10_gen_sam3_masksThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 500,
"total_frames": 46327,
"total_tasks": 10,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 80,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Beegbrain/metaworld_mt10_gen_sam3_masks.basil-segmentation-sam3
maximilian-franz/basil-segmentation-sam3
Per-instance segmented basil crops produced by the sam3 backend. This is a Hugging Face ImageFolder dataset: file_name points to the black-background masked crop used for downstream image analysis and bbox_file_name points to the corresponding unmasked rectangular crop. Empty masks are omitted. Bounding boxes use native source-frame coordinates.
Load it with:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/maximilian-franz/basil-segmentation-sam3.det_qwen3vl_sam3
注意
请解压DIOR.zip和FAIR1Mv2.zip,将解压得到的文件夹放到./DIOR路径下和./FAIR1Mv2路径下,再将DIOR_incontext.jsonl与DIOR_incontext.parquet放到./DIOR/conv路径下。
pap_cube-sam3-maskspap_cube_v2-sam3-maskslow-alt-satellite-image-dataset-5k-sam3-segmented_jsonnastol-sam3d-output
vlordier/nastol-sam3d-output
Incremental SAM3D body keypoints & meshes. Auto-generated.
sam3
SAM3 Vision Scripts
Detect and segment objects in images using Meta's SAM3 (Segment Anything Model 3) with text prompts. Process HuggingFace datasets with zero-shot detection and segmentation using natural language descriptions.
Script
What it does
Output
detect-objects.py
Object detection with bounding boxes
objects column with bbox, category, score
segment-objects.py
Pixel-level segmentation masks
Segmentation maps or per-instance masks
Browse results… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/sam3.low-alt-satellite-image-dataset-5k-sam3-segmentedscenesmith-sam3d-objects
SceneSmith SAM3D Objects
This dataset contains 30,887 static, simulation-ready 3D object assets for Drake. The objects were generated with the SceneSmith SAM3D asset pipeline and are additional generated objects beyond those packaged in the SceneSmith Example Scenes dataset.
Each object is packaged as a self-contained SDFormat asset directory containing:
SDFormat model file (.sdf)
final visual mesh referenced by the SDF (.gltf plus referenced buffers and images)
convex… See the full description on the dataset page: https://huggingface.co/datasets/nepfaff/scenesmith-sam3d-objects.episodic_box_search_robot_sam3_v1
Episodic box search: robot demonstrations with SAM3 home points
45 bimanual ARX robot sessions, exported as 638 primitive episodes and 118,299 frames at 30 Hz. This is a derived LeRobot v3 dataset for a visual-goal-conditioned low-level policy. All retained sessions belong to the training split; no held-out evaluation performance is claimed.
The robot opens visually similar boxes, shows their contents to the overhead camera, returns boxes to their home positions, and empties… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/episodic_box_search_robot_sam3_v1.blackline-atlas-sam3-real-eval-v2
Blackline Atlas SAM3 Real-Image Eval Pack
This dataset packages the Blackline Atlas SAM3/SAM3.1 selected-site evidence eval pack
with real SimSat Sentinel image pairs.
It is an evaluation and integration dataset, not a training benchmark. The cases are exact
civilian lifeline sites with current/baseline satellite frames, text prompts, expected
visual evidence tags, expected triage action, and optional normalized bboxes.
Contents
Cases: 22
Images: 44 PNG files
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ChrisRPL/blackline-atlas-sam3-real-eval-v2.label_sam3sam3-low-dice-2d-nnunet
SAM3 low-Dice 2D datasets for nnU-Net
Private research export of two small 2D datasets on which the balanced-finish
SAM3 LoRA validation Dice was below 0.5. The purpose is to test whether a
dataset-specific nnU-Net can fit these data and to distinguish data/training
limitations from inference bugs.
Dataset
SAM3 Dice
SAM3 IoU
Evaluated validation images
Actual SAM3 training images
DRIVE
0.212233
0.118717
2
14
RAVIR
0.224709
0.128455
2
16
The two-image validation… See the full description on the dataset page: https://huggingface.co/datasets/MedicalSAM3/sam3-low-dice-2d-nnunet.traffi-sam3-resized-labelsperson-benchmark-sam3
person-benchmark-sam3
Element: manak0/Detect-Person (person)
Classes: 0: person
Layout: yolo — 260 images, 13081 boxes
Source: tool:build
split
images
labeled
boxes
train
260
260
13081
Steps:
build 2026-09-16T16:32:11Z
person-new-cctv-sam3
Person New CCTV — SAM3 Annotated
YOLO-format person boxes for synthetic CCTV images from jjjlimaus/person-new-cctv-synthetic.
Annotator: SAM3 (prepare_dataset_with_sam3.py --mode person)
Class: person only (nc: 1)
Augmentation: original images plus horizontal flips (_aug_hflip), each flip re-annotated with SAM3
Samples: 2060 (train 1648 / val 412 / test 0)
Split: 8:2:0 (seed 42); flips stay in the same split as the source image
Resolution: all images stretched to 1024×1024… See the full description on the dataset page: https://huggingface.co/datasets/jjjlimaus/person-new-cctv-sam3.traffi-sam3-inputsam3-segment-test-wildlife
Image Segmentation: Deer using SAM3
This dataset contains semantic segmentation maps for deer segmented in images from davanstrien/ena24-detection using Meta's SAM3.
Generated using: uv-scripts/sam3 segmentation script
Statistics
Objects Segmented: deer
Total Instances: 0
Images with Detections: 0 / 5 (0.0%)
Average Instances per Image: 0.00
Output Format: semantic-mask
Processing Details
Source Dataset: davanstrien/ena24-detection
Model: facebook/sam3… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/sam3-segment-test-wildlife.sam3_cylinder_ring_tracked_jun22This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so107",
"total_episodes": 1,
"total_frames": 173,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/thewisp/sam3_cylinder_ring_tracked_jun22.person-high-cctv-sam3
Person High CCTV SAM3
SAM3 --mode person YOLO annotations for jjjlimaus/person-high-cctv-synthetic.
Source: KIE nano-banana-2 bird's-eye / elevated person CCTV images (originals + hflip / l30 / r30)
Class: person (nc: 1)
Layout: Ultralytics {train,val,test}/{images,labels,visualizations}/ plus dataset.yaml
Labels: YOLO class cx cy w h (.txt) and SAM3 masks (.npz)
Split: random 80:12:8 from SAM3 stage-2 (_pick_split)
Counts: 824 paired samples (train 650, val 106, test 68).… See the full description on the dataset page: https://huggingface.co/datasets/jjjlimaus/person-high-cctv-sam3.traffi-sam3-resized-input
