datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pointodysseyFeb 21: updated to v1.2. Please see https://github.com/y-zheng18/point_odyssey/tree/main for release notes.
PointWorld-DROID
PointWorld-DROID
Dataset Description:
PointWorld-DROID is the packaged DROID-derived annotation release used for training and evaluating the 3D world model, PointWorld. It contains precomputed 3D annotations derived from the official DROID dataset, including episode-level 3D point trajectories, optimized camera metadata, optional downsampled depth, and the released expert confidence artifact used by the PointWorld DROID evaluation pipeline.
This dataset is for research… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PointWorld-DROID.PointCloudCorruptionconcerto_scannet_compressedpointcalib-corpus
PointCalib corpus — 1,051,756 frames, one format
Nine source datasets normalised into a single streamable WebDataset, plus the frozen
evaluation protocol and checkpoints behind our reported numbers. The point is not
to mirror upstream archives (TartanAir, Hypersim are already on the Hub) but to
remove the nine-decoder / nine-resolution / nine-depth-convention tax: one format,
one depth convention, streamable.
Contents
split prefix
frames
shards
source… See the full description on the dataset page: https://huggingface.co/datasets/ChenmingWu/pointcalib-corpus.Point-PRC
Datasets
We conduct experiments on three new 3D domain generalization (3DDG) benchmarks proposed by us, as introduced in the next section.
base-to-new class generalization (base2new)
cross-dataset generalization (xset)
few-shot generalization (fewshot)
The structure of these benchmarks should be organized as follows.
/path/to/Point-PRC
|----data # placed in the same level of `trainers`, `weights`, etc.
|----base2new
|----modelnet40… See the full description on the dataset page: https://huggingface.co/datasets/auniquesun/Point-PRC.pointarena-data
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
Long Cheng1∗,
Jiafei Duan1,2∗,
Yi Ru Wang1†,
Haoquan Fang1,2†,
Boyang Li1†,
Yushan Huang1,
Elvis Wang3,
Ainaz Eftekhar1,2,
Jason Lee1,2,
Wentao Yuan1,
Rose Hendrix2,
Noah A. Smith1,2,
Fei Xia1,
Dieter Fox1,
Ranjay Krishna1,2
1University of Washington,
2Allen Institute for Artificial Intelligence,
3Anderson Collegiate Vocational… See the full description on the dataset page: https://huggingface.co/datasets/PointArena/pointarena-data.droid-pointflowABotN-PointBench
ABotN-Bench
ABotN-Bench is a benchmark suite built on a high-fidelity 3D Gaussian Splatting (3DGS) reconstruction stack to advance the evaluation of closed-loop, social-rule-aware visual navigation in real-world indoor and outdoor environments. This benchmark was introduced in the paper ABot-N1: Toward a General Visual Language Navigation Foundation Model.
Dataset Summary
The benchmark consists of three complementary datasets:
Dataset
Task
Goal
Scenes… See the full description on the dataset page: https://huggingface.co/datasets/acvlab/ABotN-PointBench.CoSyn-point
CoSyn-point
CoSyn-point is a collection of diverse computer-generated images that are annotated with queries and answer points.
It can be used to train models to return points in the image in response to a user query.
The data was created by using the Claude large language model to generate code that can be executed to render an image,
The code used to generate this data is open source.
Synthetic question-answer data is also available in a seperate repo.
Quick links:
📃 CoSyn… See the full description on the dataset page: https://huggingface.co/datasets/allenai/CoSyn-point.PointWorld-BEHAVIOR
PointWorld-BEHAVIOR
Dataset Description:
PointWorld-BEHAVIOR is the packaged BEHAVIOR-derived annotation release used for training and evaluating the 3D world model PointWorld. It contains precomputed 3D annotations derived from BEHAVIOR simulation episodes, organized as episode-level HDF5 files that store robot state, camera parameters, initial RGB-D observations, and rigid-body scene geometry annotations.
This Hugging Face repository hosts the packaged release, not the… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PointWorld-BEHAVIOR.POINTS-Seeker-Eval
POINTS-Seeker-Eval
This repository serves as the evaluation hub for POINTS-Seeker. It contains benchmark datasets in .tsv format and comprehensive evaluation logs across different benchmarks.
pointdit-galleryPointLLMThe official dataset release of paper ECCV 2024: PointLLM: Empowering Large Language Models to Understand Point Clouds
vipevoice-checkpointswb_pointed_chair_pull_push_rgb
wb_pointed_chair_pull_push_rgb
Whole-body teleoperation data from a Unitree_G1_WholeBody_RGB, published in LeRobot v2.1 format.
Published in the v2.1 layout (one parquet and one video clip per episode) so it loads directly on older lerobot releases. On lerobot v3.0+ run the official upgrade first:
python -m lerobot.datasets.v30.convert_dataset_v21_to_v30 --repo-id=DaoyuanZhu/wb_pointed_chair_pull_push_rgb
Task — pull out the chair indicated by the human gesture, then push it… See the full description on the dataset page: https://huggingface.co/datasets/DaoyuanZhu/wb_pointed_chair_pull_push_rgb.PointMotionBench
PointMotionBench
This is the official benchmark for the paper MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction.
A benchmark for evaluating 3D point motion in video, covering egocentric and third-person scenes across three source datasets. Each sample pairs an RGB video clip with per-object 3D and 2D tracked surface points and a human-verified natural-language caption.
Overview
Dataset
Clips
Video format
Tracks
Scene type
DAVIS… See the full description on the dataset page: https://huggingface.co/datasets/allenai/PointMotionBench.crossed_arm_point_clouds
Crossed Arm Point Clouds Dataset
This dataset contains 3D point cloud data captured from a LiDAR scanner for crossed arm classification in the context of robot magic trick performance.
Overview
This dataset was collected for training and evaluating the Crossed Arm Voxel Network (CAVN) architecture, a deep learning model designed for 3D point cloud classification in human-robot interaction magic performances. The data supports classification of human arm positions during… See the full description on the dataset page: https://huggingface.co/datasets/ahanjaya/crossed_arm_point_clouds.common_crawl_pointers_by_collectionpixmo-points
PixMo-Points
PixMo-Points is a dataset of images paired with referring expressions and points marking the locations the
referring expression refers to in the image. It was collected using human annotators and contains a diverse
range of points and expressions, with many high-frequency (10+) expressions.
PixMo-Points is a part of the PixMo dataset collection and was used to
provide the pointing capabilities of the Molmo family of models
Quick links:
📃 Paper
🎥 Blog with Videos… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-points.ABotN-POIBenchpointerbench
Pointerbench
Pointerbench is a small GUI grounding benchmark suite for computer-use models.
Each example has one screenshot, one instruction, target geometry in absolute
pixels, and a binary evaluation rule.
Links:
GitHub: https://github.com/warmwindOS/pointerbench
Blog post: https://about.warmwind.com/pointer-bench/
Add your model to the official benchmark leaderboard: https://warmwind.com/contact
The suite has three subsets:
Subset
Examples
What it tests… See the full description on the dataset page: https://huggingface.co/datasets/WarmwindOS/pointerbench.puregen-poisonspointer-retrievalrelease-910k-new
SOMA UMR Release v260717 T2M 910k (realigned)
Realigned from release_v260710_t2m_910k per
code/soma-motion-tokenizer/doc/handoff_hml_split_realign.md.
split_seed: 20260717
HumanML3D: official tags.split
bones-seed: Kimodo test pinned; val+extra-test carved from Kimodo-train
others: per-subset 80:5:15 by source group
Old release_v260710_t2m_910k is preserved untouched.
Colored_Point_Clouds
Colored point-cloud completion — PoinTr-protocol occlusions
Colored ground truth and occluded partial inputs for colored point-cloud completion,
built from textured ShapeNetCore meshes. Three categories, 599 models, 1797 partials.
What makes this different from the usual completion sets: the ground truth carries
per-point color sampled from the mesh texture, and it is de-speckled before sampling,
so the color is actually correct rather than plausible-looking.… See the full description on the dataset page: https://huggingface.co/datasets/eylulpelinkilic/Colored_Point_Clouds.PointCloudPeople
Dataset Card for PointCloudPeople
The use of light detection and ranging (LiDAR) sensor technology for people detection offers a significant advantage in terms of data protection. However, to design
these systems cost- and energy-efficiently, the relationship between the measurement data and final object detection output with deep neural networks (DNNs) has to be
elaborated. Therefore, we present an automatically labeled LiDAR dataset for person detection, with different… See the full description on the dataset page: https://huggingface.co/datasets/LukasPro/PointCloudPeople.Vis-Poison
Vis-Poison
This repository hosts the dataset accompanying
Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation.
The dataset contains 4,416 examples, each with a question, a correct
answer, a target incorrect answer, and a pair of original and
counterfactually edited images.
Paper: arXiv:2608.20756
GitHub repository: SWUFE-DB-Group/Vis-Poison
Source dataset: WebQA
Data source and construction
Vis-Poison is derived from WebQA… See the full description on the dataset page: https://huggingface.co/datasets/liangrujin/Vis-Poison.common_crawl_pointer_indicesllm-graph-poisoning-data
Generation-Time Poisoning of LLM-Generated Social Networks
This dataset contains synthetic personas, LLM-generated social graphs, cached
text embeddings, and evaluation metrics for clean generation and three
generation-time attack families. All names and profiles are synthetic and do
not represent real people.
Dataset variants
Variant
Nodes
Generator
Graph seeds per condition
Attack rates
p50
50
Qwen3-Max
10
10%, 20%, 30%, 40%, 50%
p200
200… See the full description on the dataset page: https://huggingface.co/datasets/Kevynf/llm-graph-poisoning-data.
