datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MazeMAZEL16SitesMazePlanning-Testpiper_dataset20251222抓取物体数据
cranfield-synthetic-drone-detectionIf you use this, pleace cite Drone Detection using Deep Neural Networks Trained on Pure Synthetic Data (https://arxiv.org/abs/2411.09077)
ubr-maze-nav
ubr-maze-nav — vision + instrument waypoint planning for a small tracked robot
Synthetic navigation corpus for fine-tuning small vision-language models to
plan local waypoint paths for a 0.3 m tracked ground robot in corridor/maze
environments, plus the frozen evaluation suite used in our internal reports.
Each sample is one first-person RGB frame (640×480) from the robot's camera
in a procedurally generated MuJoCo scene, an instruction carrying the goal
(bearing/range) and a… See the full description on the dataset page: https://huggingface.co/datasets/ubr-physical-ai/ubr-maze-nav.Maze-ReasoningMAZEL8Sitesmaze-solving-for-gemma-4mazes-largeportuguese-ocr-datasettask_categories:
image-to-text
task_ids:
optical-character-recognition
text-recognition
Portuguese OCR Dataset
A comprehensive dataset for Portuguese OCR (Optical Character Recognition) generated from classic Portuguese literature with diverse fonts and visual styles.
Dataset Description
This dataset contains 20000 text images for OCR training, created from Portuguese books from Project Gutenberg. Each image contains a complete Portuguese sentence with proper… See the full description on the dataset page: https://huggingface.co/datasets/mazafard/portuguese-ocr-dataset.maze_15_AUGmaze_1_AUGsketchvlm-maze-navigation
SketchVLM: Maze Navigation
This dataset is associated with the paper: SketchVLM: Vision Language Models Can Annotate Images to Explain Thoughts and Guide Users.
SketchVLM is a training-free, model-agnostic framework that enables Vision-Language Models (VLMs) to produce non-destructive, editable SVG overlays on input images to visually explain their answers. The Maze Navigation dataset is one of the benchmarks introduced to evaluate a model's ability to trace a path from start to end… See the full description on the dataset page: https://huggingface.co/datasets/loganbolton/sketchvlm-maze-navigation.inkslop-mazes-hard
InkSlop Mazes Hard
Part of the InkSlop Benchmark a vibe-coded benchmark for spatial reasoning with digital ink.
Collection: InkSlop Benchmark
Task
Maze Solving: Given a maze image, find and draw the solution path as digital ink. This "hard" variant contains diverse synthetically generated maze families with varying visual styles.
Maze Families
bamboo - Bamboo forest style
cave_skeleton - Cave system layouts
floorplan_rooms - Architectural floor plans… See the full description on the dataset page: https://huggingface.co/datasets/amaksay/inkslop-mazes-hard.ball-maze-lerobot
Ball Maze Environment Dataset
This dataset contains episodes of a ball maze environment, converted to the LeRobot format for visualization and training.
Dataset Structure
Sagar18/ball-maze-lerobot/
├── data/ # Main dataset files
├── metadata/ # Dataset metadata
│ └── info.json # Configuration and version info
├── episode_data_index.safetensors # Episode indexing information
└── stats.safetensors # Dataset statistics
Features… See the full description on the dataset page: https://huggingface.co/datasets/Sagar18/ball-maze-lerobot.mazenav_vqaDisclaimer: This is slightly modified version of SpatialEval.
Maze-Bench-v0.2Maze-Bench-v0.1javatari_sft_8_AUG_shooting_sports_maze_actionMazeBench
MazeBench
The evaluation set from the paper From Pixels to BFS: High Maze Accuracy Does Not Imply Visual Planning.
Paper: arXiv:2603.26839
Code: github.com/alrod97/LLMs_mazes
Overview
110 procedurally generated maze images spanning 8 structural families and grid sizes from 5x5 to 20x20, with ground-truth shortest-path annotations. These are the mazes used to evaluate multimodal LLMs in the paper.
Maze Families
sanity — open corridors, minimal… See the full description on the dataset page: https://huggingface.co/datasets/albertoRodriguez97/MazeBench.mazes-smallmazes-extralargeaigcAsian photography dataset
win3000: about 18k asian celebrity photo.
jiepaigou: streetsnap and celebrity
cybesx: about 13k street photography
figuritas-de-mazapanmazes-mediumcranfield-synthetic-drone-classificationPart of the "Drone Model Classification Using Convolutional Neural Network Trained on Synthetic Data" paper (https://www.mdpi.com/2313-433X/8/8/218).
Contains four classes:
DJI Mavic
DJI Phantom
DJI Inspire
No Drone
If you find this useful, please consider citing:
@Article{jimaging8080218,
AUTHOR = {Wisniewski, Mariusz and Rana, Zeeshan A. and Petrunin, Ivan},
TITLE = {Drone Model Classification Using Convolutional Neural Network Trained on Synthetic Data},
JOURNAL = {Journal of Imaging}… See the full description on the dataset page: https://huggingface.co/datasets/mazqtpopx/cranfield-synthetic-drone-classification.license-detection-paligemmatwitter-werwervov-2026.02.26-2026904290804298165-pPDFeh9_MAZse21Q-part1
