datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openimages-bboxImages and nounding box annotations from the OpenImages dataset.
TextVQA_GT_bbox
TextVQA validation set with grounding truth bounding box
The dataset used in the paper MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs for studying MLLMs' attention patterns.
The dataset is sourced from TextVQA and annotated manually with ground-truth bounding boxes.
We consider questions with a single area of interest in the image so that 4370 out of 5000 samples are kept.
Citation
If you find our paper and code useful… See the full description on the dataset page: https://huggingface.co/datasets/jrzhang/TextVQA_GT_bbox.dhivehi-image-bbox-prompt
Dhivehi Image Bounding Box Prompt Dataset
This dataset, alakxender/dhivehi-image-bbox-prompt, contains 58,738 images annotated with COCO-style bounding boxes and Dhivehi (Thaana script) text, along with layout categories such as Text, Title, Picture, Caption, and Columns. It is designed for OCR, document layout analysis, and multimodal vision–language research focused on Dhivehi.
Dataset
Each row includes:
image — the RGB image (preserved original dimensions)
width… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-image-bbox-prompt.dhivehi-image-bbox-ds-fmt
Dhivehi Image Bounding Box Dataset - DeepSeek Format
Dataset Description
This dataset is a transformed version of alakxender/dhivehi-image-bbox-prompt, specifically formatted to align with DeepSeek OCR model requirements for training vision-language models with grounding capabilities.
The original dataset contained Dhivehi (Thaana script) text with bounding box annotations. This version restructures the annotations into DeepSeek's grounding token format, enabling the… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-image-bbox-ds-fmt.easyr1-110k-bbox0p05-minus-stage3-rl0p2-noise-add-rest-yt-4MP
easyr1-110k-bbox0p05-minus-stage3-rl0p2-noise-add-rest-yt-4MP
Merged dataset composed of the following sources:
datasets/easyr1-103k-bbox0p05-minus-stage3-rl0p2-noise (100155 samples in split train)
mlfoundations-cua-dev/66-yt-app-ui-claude-instructions-no-filter-4MP-gta1-correct-qwen7b-not-correct (7573 samples in split train)
Summary
Generated on: 2025-09-23 15:18:54 UTC
Split: train
Column strategy: intersection
Samples after merge: 107728
Usage
from… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-110k-bbox0p05-minus-stage3-rl0p2-noise-add-rest-yt-4MP.clevr-bboxpublaynet_bboxeasyr1-110k-bbox0p05-minus-stage3-rl0p2-noise-add-rest-yt-4MP-remove-pixmo-uground-seeclickham10000_bbox
HAM10000 with Spatial Annotations and Bounding Box Coordinates
Enhanced version of HAM10000 dataset with bounding box coordinates and spatial descriptions for skin lesion localization.
Dataset Description
This dataset extends the original HAM10000 dermatology dataset with:
Bounding box coordinates for lesion localization
Spatial descriptions (e.g., "located in center-center region")
Area coverage statistics
Mask availability flags
Features
image: RGB skin… See the full description on the dataset page: https://huggingface.co/datasets/abaryan/ham10000_bbox.object_detection_bbox_merged_coco_foodinvoice-annotated-bboxManually annotated invoice page images exported from AnnotateEverything, with axis-aligned bounding boxes for 8 document-layout regions. Built for training object detectors (YOLO, DETR, etc.) on invoice macro-structure.
Dataset summary
Property
Value
Pages
76
Documents
1
Source PDF
train_images.pdf
Total annotations
771
Avg boxes / page
10.14
Image width range
425 – 2853 px
Image height range
570 – 4096 px
Export date
2026-06-22T19:27:38.375Z… See the full description on the dataset page: https://huggingface.co/datasets/AvoCahDoe/invoice-annotated-bbox.vlm-project-with-images-with-bbox-images-with-tree-of-thoughts-v3qwen3-resize-easyr1-110k-bbox0p05-remove-pixmo-uground-seeclick-normalizedref_val_bboxvlm-project-with-images-with-bbox-images-with-tree-of-thoughtsvlm-project-with-images-with-bbox-images-v5so101_car_pick_and_place-bboxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 98,
"total_frames": 67279,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:98"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jonathm126/so101_car_pick_and_place-bbox.vlm-project-with-images-with-bbox-images-official-q3-updatevlm-project-with-images-with-bbox-images-with-tree-of-thoughts-v2ViRFT_COCO_bbox2dvlm-project-with-images-with-bbox-images-with-tree-of-thoughts-with-originalbrain_tumor_bboxvlm-project-with-images-with-bbox-images-with-tree-of-thoughts-RLHF-v6vlm-project-with-images-with-bbox-images-with-tree-of-thoughts-original-onlyvlm-project-with-images-with-bbox-images-v6audio_bbox_balancedso101_pick_pen-bboxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 51,
"total_frames": 35904,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jonathm126/so101_pick_pen-bbox.vlm-project-with-images-with-bbox-images-officialvabench-point-bboxvstar-bench-with-bbox
vstar-bench with Bounding Box Annotations
Dataset Description
This dataset extends lmms-lab/vstar-bench by adding bounding box annotations for target objects. The bounding box information was extracted from craigwu/vstar_bench and mapped to the lmms-lab version.
Key Features
Visual Spatial Reasoning: Tests understanding of spatial relationships in images
Bounding Box Annotations: Each sample includes target object bounding boxes
Multiple Choice QA: 4-way… See the full description on the dataset page: https://huggingface.co/datasets/jae-minkim/vstar-bench-with-bbox.
