datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openimages-bboxImages and nounding box annotations from the OpenImages dataset.
droid_120_stsg_bboxesTextVQA_GT_bbox
TextVQA validation set with grounding truth bounding box
The dataset used in the paper MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs for studying MLLMs' attention patterns.
The dataset is sourced from TextVQA and annotated manually with ground-truth bounding boxes.
We consider questions with a single area of interest in the image so that 4370 out of 5000 samples are kept.
Citation
If you find our paper and code useful… See the full description on the dataset page: https://huggingface.co/datasets/jrzhang/TextVQA_GT_bbox.chris_stsg_bboxes_v1_part0persian-ocr-bench-submitted10-bbox-crops
Persian OCR benchmark — selected submitted bbox crops
This dataset contains the non-empty OCR bboxes from the ten explicitly selected
submitted pages in persian_ocr_bench_bbox_review.
Each row is one PNG crop. gold_text is the current editable OCR content from
the live Argilla bbox field (content_text). Geometry is stored both as source
page pixels and as percentages of the source page. The original record ID,
external ID, bbox ID, source URL, and SHA-256 hashes are included for… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-ocr-bench-submitted10-bbox-crops.dhivehi-image-bbox-prompt
Dhivehi Image Bounding Box Prompt Dataset
This dataset, alakxender/dhivehi-image-bbox-prompt, contains 58,738 images annotated with COCO-style bounding boxes and Dhivehi (Thaana script) text, along with layout categories such as Text, Title, Picture, Caption, and Columns. It is designed for OCR, document layout analysis, and multimodal vision–language research focused on Dhivehi.
Dataset
Each row includes:
image — the RGB image (preserved original dimensions)
width… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-image-bbox-prompt.qwen3-resize-easyr1-110k-bbox0p05-remove-pixmo-uground-seeclickchris_stsg_bboxes_v1_part1snakeaid-yolov12-5000-bbox
SnakeAid YOLOv12 5000 BBox
Dataset Summary
This repository contains a YOLO-format SnakeAid object-detection dataset for snake detection experiments. It is organized as image/label pairs across train, valid, test splits and is intended for training or evaluating YOLO-family detectors, including the related SnakeAid Detect YOLOv12 checkpoints linked below.
Safety note: snake detection can be safety-critical in real-world use. Treat model outputs trained on this data as… See the full description on the dataset page: https://huggingface.co/datasets/the-khiem7/snakeaid-yolov12-5000-bbox.chris_stsg_bboxes_v1_part2dhivehi-image-bbox-ds-fmt
Dhivehi Image Bounding Box Dataset - DeepSeek Format
Dataset Description
This dataset is a transformed version of alakxender/dhivehi-image-bbox-prompt, specifically formatted to align with DeepSeek OCR model requirements for training vision-language models with grounding capabilities.
The original dataset contained Dhivehi (Thaana script) text with bounding box annotations. This version restructures the annotations into DeepSeek's grounding token format, enabling the… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-image-bbox-ds-fmt.sat-bbox-metadata-sft-v1
Dataset Summary
NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills).
The goal is to create high-signal, production-shaped supervision for multimodal chat models:
Captioning for satellite chips
Grounding (bounding boxes in normalized coordinates) for land-cover regions
Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-bbox-metadata-sft-v1.droid_120_stsg_bboxes_v4easyr1-110k-bbox0p05-minus-stage3-rl0p2-noise-add-rest-yt-4MP
easyr1-110k-bbox0p05-minus-stage3-rl0p2-noise-add-rest-yt-4MP
Merged dataset composed of the following sources:
datasets/easyr1-103k-bbox0p05-minus-stage3-rl0p2-noise (100155 samples in split train)
mlfoundations-cua-dev/66-yt-app-ui-claude-instructions-no-filter-4MP-gta1-correct-qwen7b-not-correct (7573 samples in split train)
Summary
Generated on: 2025-09-23 15:18:54 UTC
Split: train
Column strategy: intersection
Samples after merge: 107728
Usage
from… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-110k-bbox0p05-minus-stage3-rl0p2-noise-add-rest-yt-4MP.LightOnOCR-bbox-mix-0126
LightOnOCR-bbox-mix-0126
LightOnOCR-bbox-mix-0126 is a large-scale OCR training dataset including layout information built via distillation: a strong vision–language model is prompted to produce naturally ordered full-page transcriptions (Markdown with LaTeX math spans and HTML tables) from rendered document pages. The dataset is designed as supervision for end-to-end OCR / document-understanding models that aim to output clean, human-readable text in a consistent format.
This… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/LightOnOCR-bbox-mix-0126.indoor-3d-bbox-indices
Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding
This repository contains the ScanNet dataset (3D scene data and 2D frame data) and refined annotations used for the paper Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding.
TAB is a dynamic agentic framework designed for zero-shot 3D Visual Grounding (3D-VG). By operating directly on raw RGB-D streams, TAB reformulates 3D… See the full description on the dataset page: https://huggingface.co/datasets/AntonioJun/indoor-3d-bbox-indices.publaynet_bboxclevr-bboxchris_stsg_bboxes_v1_part5easyr1-110k-bbox0p05-minus-stage3-rl0p2-noise-add-rest-yt-4MP-remove-pixmo-uground-seeclickham10000_bbox
HAM10000 with Spatial Annotations and Bounding Box Coordinates
Enhanced version of HAM10000 dataset with bounding box coordinates and spatial descriptions for skin lesion localization.
Dataset Description
This dataset extends the original HAM10000 dermatology dataset with:
Bounding box coordinates for lesion localization
Spatial descriptions (e.g., "located in center-center region")
Area coverage statistics
Mask availability flags
Features
image: RGB skin… See the full description on the dataset page: https://huggingface.co/datasets/abaryan/ham10000_bbox.chris_stsg_bboxes_v1_part4invoice-annotated-bboxManually annotated invoice page images exported from AnnotateEverything, with axis-aligned bounding boxes for 8 document-layout regions. Built for training object detectors (YOLO, DETR, etc.) on invoice macro-structure.
Dataset summary
Property
Value
Pages
76
Documents
1
Source PDF
train_images.pdf
Total annotations
771
Avg boxes / page
10.14
Image width range
425 – 2853 px
Image height range
570 – 4096 px
Export date
2026-06-22T19:27:38.375Z… See the full description on the dataset page: https://huggingface.co/datasets/AvoCahDoe/invoice-annotated-bbox.object_detection_bbox_merged_coco_foodqwen3-resize-easyr1-110k-bbox0p05-remove-pixmo-uground-seeclick-normalizedvlm-project-with-images-with-bbox-images-with-tree-of-thoughts-v3obstaclenet-voc2007-bbox
obstaclenot-voc2007-grid
Binary segmentation masks derived from PASCAL VOC 2007, built for training
the seg_net of the ObstacleNet
robot-vision pipeline.
What is in each row?
Each row contains the original JPEG image and a 28×28 binary mask where
255 = obstacle (any object bounding box covers this pixel) and 0 = clear.
The mask is 28×28 to match the SegNet output: 32×32 input → two 3×3 convolutions
without padding → 28×28 output.
Superclass mapping… See the full description on the dataset page: https://huggingface.co/datasets/robro/obstaclenet-voc2007-bbox.ubuntu_traj_rl_bbox_max_500_samples_max_turns_10droid_120_stsg_bboxes_v2ref_val_bbox
