datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
od-syn-page-annotations-com
📦 Dhivehi Synthetic Document Layout + Textline Dataset
This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis, visual document understanding, OCR fine-tuning, and related tasks specifically for Dhivehi script.
Note: this version image are compressed.
Raw version 📁 Repository: Hugging Face Datasets
📋 Dataset Summary
Total Examples: ~58… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations-com.od-syn-page-annotations
📦 Dhivehi Synthetic Document Layout + Textline Dataset
This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis , visual document understanding , OCR fine-tuning, and related tasks specifically for Dhivehi script.
📋 Dataset Summary
Total Examples: ~58,738
Image Content: Synthetic Dhivehi documents generated to simulate real-world layouts… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations.forest-fire-annotations
Forest Fire Detection Dataset — Auto-Annotated
Bounding-box annotated version of touati-kamel/forest-fire-dataset,
built for training forest-fire / smoke / fog object detection models.
Overview
This dataset contains video frames auto-labeled with bounding boxes for fire and
smoke-related visual phenomena, using a zero-shot open-vocabulary object detector
(Grounding DINO). It is derived from the original touati-kamel/forest-fire-dataset image
classification dataset… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/forest-fire-annotations.real-resumes-section-detection-annotationsbower-waste-annotations
Dataset Card for waste annotations made by the recycling solution Bower
The data offered by Bower (Sugi Group AB) in collaboration with Google.org
Dataset Summary
The bower-waste-annotations dataset consists of 1440 images of waste and various consumer items taken by consumer phone cameras. The images are annotated with Material type and Object type classes, listed below.
The images and annotations has been manually reviewed to ensure correctness. It is assumed… See the full description on the dataset page: https://huggingface.co/datasets/BowerApp/bower-waste-annotations.amz-image-annotationsImageIn_annotations_resized_images
Dataset Card for ImageIn_annotations_resized_images
More Information needed
annotations_only_10pct_gpt5_miniImageIn_annotationsInitial annotated dataset derived from ImageIN/IA_unlabelled
TriConflict-hallucination-annotationscc3m-grounded-annotations
CC3M grounded annotations
Region-level grounding for Conceptual Captions 3M: bounding boxes, the noun
phrase each box grounds, and the span of the caption that phrase came from, for
3,016,640 of CC3M's 3,318,333 rows.
No images here. This is metadata only, joinable onto a CC3M copy you already
have. That is the point of it: the grounding is 354 MB, the pixels are 125 GB.
Files
file
rows
size
annotations-0000..0482.parquet
3,016,640
199 MB… See the full description on the dataset page: https://huggingface.co/datasets/freek23/cc3m-grounded-annotations.new_year-50-annotationsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 23,
"total_frames": 42588,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:23"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jadechoghari/new_year-50-annotations.leicester_loaded_annotations_binary
Dataset Card for "leicester_loaded_annotations_binary"
More Information needed
multilingual-image-annotations
Multilingual Image Annotations
Image annotations across 7 languages (en, es, fr, hi, zh, ar, pt) generated by google/gemma-4-31B-it via the Hugging Face Router. Each row pairs an image with an English description, multilingual descriptions, 21 VQA pairs (3 per language), and conditional object detections with normalized bounding boxes. When detections are present, a derivative image with rectangles drawn is included as boxed_image.
Stats
Images: 464… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/multilingual-image-annotations.ImageIn_annotationstask2-1_original-annotationsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 45,
"total_frames": 37489,
"total_tasks": 3,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:45"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jasontchan/task2-1_original-annotations.leicester_loaded_annotations
Dataset Card for "leicester_loaded_annotations"
More Information needed
Multimodal-Hallucination-Annotations-reducedTriConflict-annotationsbrick-annotations-v2ads-annotationsMulti-Hallucination-Annotations-exampleMultimodal-Hallucination-Annotationshindi-annotationsCMultimodal_Hallucination_Annotationsreasoning_annotations_002TriConflict-hallucination-annotations-80_20GPTMultimodal_Hallucination_Annotationsgenerated-resumes-section-detection-annotationsTriConflict-hallucination-annotations-80_20
