datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pointcalib-corpus
PointCalib corpus — 1,051,756 frames, one format
Nine source datasets normalised into a single streamable WebDataset, plus the frozen
evaluation protocol and checkpoints behind our reported numbers. The point is not
to mirror upstream archives (TartanAir, Hypersim are already on the Hub) but to
remove the nine-decoder / nine-resolution / nine-depth-convention tax: one format,
one depth convention, streamable.
Contents
split prefix
frames
shards
source… See the full description on the dataset page: https://huggingface.co/datasets/ChenmingWu/pointcalib-corpus.pointarena-data
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
Long Cheng1∗,
Jiafei Duan1,2∗,
Yi Ru Wang1†,
Haoquan Fang1,2†,
Boyang Li1†,
Yushan Huang1,
Elvis Wang3,
Ainaz Eftekhar1,2,
Jason Lee1,2,
Wentao Yuan1,
Rose Hendrix2,
Noah A. Smith1,2,
Fei Xia1,
Dieter Fox1,
Ranjay Krishna1,2
1University of Washington,
2Allen Institute for Artificial Intelligence,
3Anderson Collegiate Vocational… See the full description on the dataset page: https://huggingface.co/datasets/PointArena/pointarena-data.ABotN-PointBench
ABotN-Bench
ABotN-Bench is a benchmark suite built on a high-fidelity 3D Gaussian Splatting (3DGS) reconstruction stack to advance the evaluation of closed-loop, social-rule-aware visual navigation in real-world indoor and outdoor environments. This benchmark was introduced in the paper ABot-N1: Toward a General Visual Language Navigation Foundation Model.
Dataset Summary
The benchmark consists of three complementary datasets:
Dataset
Task
Goal
Scenes… See the full description on the dataset page: https://huggingface.co/datasets/acvlab/ABotN-PointBench.CoSyn-point
CoSyn-point
CoSyn-point is a collection of diverse computer-generated images that are annotated with queries and answer points.
It can be used to train models to return points in the image in response to a user query.
The data was created by using the Claude large language model to generate code that can be executed to render an image,
The code used to generate this data is open source.
Synthetic question-answer data is also available in a seperate repo.
Quick links:
📃 CoSyn… See the full description on the dataset page: https://huggingface.co/datasets/allenai/CoSyn-point.pointdit-gallerypixmo-points
PixMo-Points
PixMo-Points is a dataset of images paired with referring expressions and points marking the locations the
referring expression refers to in the image. It was collected using human annotators and contains a diverse
range of points and expressions, with many high-frequency (10+) expressions.
PixMo-Points is a part of the PixMo dataset collection and was used to
provide the pointing capabilities of the Molmo family of models
Quick links:
📃 Paper
🎥 Blog with Videos… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-points.pointerbench
Pointerbench
Pointerbench is a small GUI grounding benchmark suite for computer-use models.
Each example has one screenshot, one instruction, target geometry in absolute
pixels, and a binary evaluation rule.
Links:
GitHub: https://github.com/warmwindOS/pointerbench
Blog post: https://about.warmwind.com/pointer-bench/
Add your model to the official benchmark leaderboard: https://warmwind.com/contact
The suite has three subsets:
Subset
Examples
What it tests… See the full description on the dataset page: https://huggingface.co/datasets/WarmwindOS/pointerbench.PointCloudPeople
Dataset Card for PointCloudPeople
The use of light detection and ranging (LiDAR) sensor technology for people detection offers a significant advantage in terms of data protection. However, to design
these systems cost- and energy-efficiently, the relationship between the measurement data and final object detection output with deep neural networks (DNNs) has to be
elaborated. Therefore, we present an automatically labeled LiDAR dataset for person detection, with different… See the full description on the dataset page: https://huggingface.co/datasets/LukasPro/PointCloudPeople.ABotN-POIBenchVis-Poison
Vis-Poison
This repository hosts the dataset accompanying
Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation.
The dataset contains 4,416 examples, each with a question, a correct
answer, a target incorrect answer, and a pair of original and
counterfactually edited images.
Paper: arXiv:2608.20756
GitHub repository: SWUFE-DB-Group/Vis-Poison
Source dataset: WebQA
Data source and construction
Vis-Poison is derived from WebQA… See the full description on the dataset page: https://huggingface.co/datasets/liangrujin/Vis-Poison.pixmo-point-count-concat_0-20pixmo-point-count-gen-undPhysical_poisoned_I_4image-pointing-1M-sft-swiftPointBenchro_sft_pixmo_points
Dataset Description
PixmoPoints is a dataset of images paired with referring expressions and points marking the locations the referring expression refers to in the image.
Here we provide the Romanian translation of the PixmoPoints dataset, translated with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_pixmo_points.SAM_PointPrompt_Dataset
Abstract
The remarkable capabilities of the Segment Anything Model (SAM) for tackling image segmentation tasks in an intuitive and interactive manner has sparked interest in the design of effective visual prompts. Such interest has led to the creation of automated point prompt selection strategies, typically motivated from a feature extraction perspective. However, there is still very little understanding of how appropriate these automated visual prompting strategies are… See the full description on the dataset page: https://huggingface.co/datasets/gOLIVES/SAM_PointPrompt_Dataset.xarm_dishwasher_points9_arrow_len0
xarm_dishwasher_points9_arrow_len0
xArm7 phone-teleop dataset for dishwasher ("pull the basket outside of the dishwasher and pick up the mug and put it into the basket") with the points9_arrow_len0
tactile overlay burned into the camera frames: 9 dots per gripper finger drawn at their
real 3D tactile-pad positions (forward kinematics, projected through camera intrinsics),
with force arrows rendered at zero length (arrow_length_scale=0).
This is the force-ablation control for the… See the full description on the dataset page: https://huggingface.co/datasets/EdwardoSunny/xarm_dishwasher_points9_arrow_len0.IISC_hand_all_final_point_object_tracksThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 26,
"total_frames": 6152,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:26"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hyzhang01/IISC_hand_all_final_point_object_tracks.pixmo-pointsYe
Open Schematics Dataset
A comprehensive dataset of electronic schematics from hardware projects. This dataset is designed for training AI models on circuit design, component recognition, and hardware engineering tasks.
Dataset Description
This dataset contains electronic schematic files along with their visual representations, component information, and metadata from various hardware projects.
Dataset Structure
Each record in the dataset contains:
schematic:… See the full description on the dataset page: https://huggingface.co/datasets/POISONX/Ye.xarm_charger_points9_arrow_len0
xarm_charger_points9_arrow_len0
xArm7 phone-teleop dataset for charger ("pick up the charger and plug it into the nearest plug") with the points9_arrow_len0
tactile overlay burned into the camera frames: 9 dots per gripper finger drawn at their
real 3D tactile-pad positions (forward kinematics, projected through camera intrinsics),
with force arrows rendered at zero length (arrow_length_scale=0).
This is the force-ablation control for the points9_arrow variant: identical dot… See the full description on the dataset page: https://huggingface.co/datasets/EdwardoSunny/xarm_charger_points9_arrow_len0.GPT_Pointillism_Style_Images
GPT Pointillism Style Images
Dataset description
This is a synthetic GPT-generated pointillism style image dataset. It contains 70 image-caption pairs with colorful dotted texture, painterly lighting, and pointillism-inspired compositions.
The images focus on colorful stippled brushwork, luminous dotted texture, painterly lighting, cozy subjects, gardens, landscapes, interiors, and decorative still-life scenes.
Contents
images/
metadata.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/GPT_Pointillism_Style_Images.ur5_3finger_pointThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5",
"total_episodes": 27,
"total_frames": 3880,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:27"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hyzhang01/ur5_3finger_point.lab_data_orange_cube_single_point_paired_25This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 50,
"total_frames": 12001,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ceilingfan456/lab_data_orange_cube_single_point_paired_25.GCA_robotiq_franka_point_object_tracksThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 32,
"total_frames": 10859,
"total_tasks": 2,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hyzhang01/GCA_robotiq_franka_point_object_tracks.Point-MAE-Zero
Procedural 3D Synthetic Shapes Dataset
Overview
This dataset contains 152,508 procedurally synthesized 3D shapes in order to help people better reproduce results for Semantic-Free Procedural 3D Shapes Are Surprisingly Good Teachers. The shapes are created using a procedural 3D program that combines primitive shapes (e.g., cubes, spheres, and cylinders) and applies various transformations and augmentations to enhance geometric diversity.
Our dataset is collected based on… See the full description on the dataset page: https://huggingface.co/datasets/uva-cv-lab/Point-MAE-Zero.pixmo-points-filtered-below10_imgContainedforeground_object_pointcloudslibero_10_poisoned_T_10This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 391,
"total_frames": 98435,
"total_tasks": 20,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:391"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LEE181204/libero_10_poisoned_T_10.
