datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
amara-spatial-10k
AmaraSpatial-10K
A Semantically Anchored, Metric-Scale 3D Dataset for Embodied AI and Spatial Computing
10,071 AI-generated 3D meshes across 10 top-level categories and 476 subcategories — from basilisks to bassoons, cottages to cosmic stations — curated by Zero One Creative to close the spatial alignment gap that makes most generative 3D repositories unusable for zero-shot deployment in game engines, robotics simulators, and AR/VR pipelines.
Every asset is… See the full description on the dataset page: https://huggingface.co/datasets/ZeroOneCreative/amara-spatial-10k.Awesome_Spatial_VQA_BenchmarksSpatialRGPT-Benchlibero_gen_spatial_combination_train_openpilibero_spatial_imageThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 432,
"total_frames": 52970,
"total_tasks": 10,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:432"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/libero_spatial_image.Spatial457Awesome_Spatial_VQA_Benchmarks_ViewSpatial-BenchSpatialBlock-15k
SpatialBlock-15k
This dataset accompanies the paper SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem. It contains 15,000 synthetic block-stacking problems for training large vision-language models (LVLMs) to improve spatial reasoning. The dataset includes three types of multiple-choice questions:
Q1: 3D-to-2D projection
Q2: viewpoint transformation
Q3: structural combination
The dataset is organized into a train split of 15,000… See the full description on the dataset page: https://huggingface.co/datasets/rsoohyun/SpatialBlock-15k.Q-Spatial-Bench
Dataset Card for Q-Spatial Bench
Q-Spatial Bench is a benchmark designed to measure the quantitative spatial reasoning 📏 in large vision-language models.
🔥The paper associated with Q-Spatial Bench is accepted by EMNLP 2024 main track!
Our paper: Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models [arXiv link]
Project website: [link]
Dataset Details
Q-Spatial Bench is a benchmark designed to measure the… See the full description on the dataset page: https://huggingface.co/datasets/andrewliao11/Q-Spatial-Bench.spatialvlm_qa
Synthetic Spatial Visual Language Question Answering Dataset
Automatically generated from the Blender Scene Dataset. Each example contains an image and a question-answer pair to probe metric (numeric) and relation (true/false) spatial reasoning skills.
Images: Rendered using Blender (1000 scenes, 5 random primitives each, random cameras and lighting).
Metadata: Object name, position, scale, color, material flags.
Questions: 10 per image, drawn from handcrafted templates… See the full description on the dataset page: https://huggingface.co/datasets/Litian2002/spatialvlm_qa.SpatialVQAlibero_spatial_with_depthThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 476,
"total_frames": 60429,
"total_tasks": 10,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 20,
"splits": {
"train": "0:476"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/WillMandil001/libero_spatial_with_depth.libero_spatial_imageThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 432,
"total_frames": 52970,
"total_tasks": 10,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:432"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/theayos/libero_spatial_image.SAT-Spatial-VQASpatialCLEVRSpatialTree-Benchspatial-visual-reasoning-66klibero_spatial_onlyThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 432,
"total_frames": 52970,
"total_tasks": 10,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:432"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Mirage415/libero_spatial_only.libero40_libero_spatial_v2_playThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 420,
"total_frames": 117600,
"total_tasks": 10,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:420"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/VarunGiridhar3/libero40_libero_spatial_v2_play.MER-BenchLerobot_Glue_299SpatialEval
🤔 About SpatialEval
SpatialEval is a comprehensive benchmark for evaluating spatial intelligence in LLMs and VLMs across four key dimensions:
Spatial relationships
Positional understanding
Object counting
Navigation
Benchmark Tasks
Spatial-Map: Understanding spatial relationships between objects in map-based scenarios
Maze-Nav: Testing navigation through complex environments
Spatial-Grid: Evaluating spatial reasoning within structured environments
Spatial-Real:… See the full description on the dataset page: https://huggingface.co/datasets/MilaWang/SpatialEval.SpatialWorld
SpatialWorld — Task Card Metadata (v1.0)
760 tasks from SpatialWorld.
Column
Notes
task_id
Task identifier
task_name
Short title (from task.json when available)
instruction
Natural-language instruction
backend
Simulator backend
scene
Scene / map identifier
category
Scenario category (Daily, Work, Travel, Entertain, Social, …)
task_type
Navigation, Interaction, or Hybrid
data_path
Path to task assets in data/
Golden trajectories: indoor/outdoor tasks… See the full description on the dataset page: https://huggingface.co/datasets/HongchengGao/SpatialWorld.libero_spatial_mask_lerobot_256x256This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 500,
"total_frames": 62250,
"total_tasks": 10,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/binhng/libero_spatial_mask_lerobot_256x256.vqasynth_sample_spatialai2thor_spatial_verification_val_v2spatialtunnel
SpatialTunnel
SpatialTunnel is a Blender-rendered diagnostic dataset for studying how vision-language models represent spatial relations internally. It was introduced in Why Far Looks Up: Probing Spatial Representation in Vision-Language Models (arXiv:2605.30161).
Resources
Project page
Contrastive-probing code
SpatialTunnel generation code
Dataset Configurations
Config
File
Rows
Description
phase_variation
phase_variation-*.parquet
12… See the full description on the dataset page: https://huggingface.co/datasets/cubec/spatialtunnel.libero_goal_object_spatialThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 1314,
"total_frames": 171996,
"total_tasks": 30,
"total_videos": 0,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:1314"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/siamaky1984/libero_goal_object_spatial.spatial-programmingColumns:
input_text — text prompt (may be null)
input_image — reference image (may be null); at least one of text / image is set
code — Python source that builds the asset
glb — the resulting GLB file, raw bytes
platform — e.g. blender
type — e.g. modeling
model — which model wrote the code (fable, fable 5.1, opus 5, astra, sol)
Rows are appended one parquet shard per upload under data/.
open-spatial-reasoning
Open Spatial Reasoning
A multiple-choice dataset of spatial reasoning questions and answers for evaluating 3D spatial reasoning from single driving images. Each image contains numbered bounding boxes referencing objects in the scene, and each question probes a model's ability to reconstruct the real 3D scene rather than rely on flat-image shortcuts (e.g. "lower in the frame = closer", "bigger box = nearer").
Dataset Description
Frontier vision-language models… See the full description on the dataset page: https://huggingface.co/datasets/ReasonCore/open-spatial-reasoning.
