datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ProcVQA-20M-annotations
ProcVQA-20M Annotations
Project Page |
arXiv |
Code |
Model |
Media
This repository contains the text annotations for the ProcVQA-20M dataset. The full image files are hosted separately on ProcVQA-20M-media.
Overview
This dataset is constructed from over 26 embodied datasets, comprising:
20M QA pairs for training
330K original trajectories
50M annotated frames from ~5,000 hours of manipulation data
200+ different tasks
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/ce-amtic/ProcVQA-20M-annotations.AmharicCLIP-annotation
AmharicCLIP Annotation Dataset
69,629 images organized by category for Amharic caption annotation.
Structure
images/
animals10/ 18,644 images — 10 animal classes
cat/
dog/
horse/ ...
intel/ 11,998 images — 6 scene classes
forest/
mountain/ ...
fruits360/ 38,987 images — 131 fruit classes
apple/
banana/ ...
Image URL Format… See the full description on the dataset page: https://huggingface.co/datasets/CLIPAMharic/AmharicCLIP-annotation.Drone-Orthomosaic-Vehicles-Yolo-annotation
Dataset Tailings Mining Vehicles & Instruments (High-Res Drone Imagery)
Dataset Summary
This dataset contains high-resolution aerial imagery focused on vehicle detection and geotechnical monitoring instruments within active mining environments (tailings dams). The data was acquired using a DJI Zenmuse P1 sensor at 120m altitude.
Photogrammetric Context
The images originate from large-scale georeferenced orthomosaics generated from bi-daily… See the full description on the dataset page: https://huggingface.co/datasets/titoruizh/Drone-Orthomosaic-Vehicles-Yolo-annotation.od-syn-page-annotations-com
📦 Dhivehi Synthetic Document Layout + Textline Dataset
This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis, visual document understanding, OCR fine-tuning, and related tasks specifically for Dhivehi script.
Note: this version image are compressed.
Raw version 📁 Repository: Hugging Face Datasets
📋 Dataset Summary
Total Examples: ~58… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations-com.od-syn-page-annotations
📦 Dhivehi Synthetic Document Layout + Textline Dataset
This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis , visual document understanding , OCR fine-tuning, and related tasks specifically for Dhivehi script.
📋 Dataset Summary
Total Examples: ~58,738
Image Content: Synthetic Dhivehi documents generated to simulate real-world layouts… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations.bower-waste-annotations
Dataset Card for waste annotations made by the recycling solution Bower
The data offered by Bower (Sugi Group AB) in collaboration with Google.org
Dataset Summary
The bower-waste-annotations dataset consists of 1440 images of waste and various consumer items taken by consumer phone cameras. The images are annotated with Material type and Object type classes, listed below.
The images and annotations has been manually reviewed to ensure correctness. It is assumed… See the full description on the dataset page: https://huggingface.co/datasets/BowerApp/bower-waste-annotations.map-annotation-tool-data
Map Annotation Tool Media
This dataset contains Docling-extracted figure crops grouped by source PDF. It is the
media input for map-annotation-tool; human annotations are maintained separately in
the application repository and submitted through pull requests.
Contents
media/figures/brgm-v1/: 1,512 figures from the earlier BRGM parsing batch,
of which 156 had legacy annotations at migration time.
media/figures/brgm-v2/: 6,046 figures from the newer BRGM parsing… See the full description on the dataset page: https://huggingface.co/datasets/marijanic/map-annotation-tool-data.malaysia-trash-annotation
Notice:
The foundational manuscript detailing the methodology, taxonomy (Build 20260418), and benchmark performance of the MTA dataset is currently Under Review.
If you are utilizing this dataset for research, please bookmark this repository. The official BibTeX citation will be published here upon the paper's acceptance.
malaysian-trash-annotation > Build-20260418-x1
-x1 refer to dataset without any augmentation applied to.
Check out the Roboflow Universe… See the full description on the dataset page: https://huggingface.co/datasets/larmkaixian/malaysia-trash-annotation.human_annotation_web_rm_version_1{
"total_stats": {
"total_action_data": 19016,
"total_website": 50,
"action_type": {
"bid": 852,
"coord": 848
},
"level": {
"easy": 486,
"medium": 856,
"hard": 358
},
"viewport_type": {
"full": 707,
"laptop": 698,
"mobile": 295
},
"judge_type": {
"string_match": 858,
"url_match": 836… See the full description on the dataset page: https://huggingface.co/datasets/WPRM/human_annotation_web_rm_version_1.forest-fire-annotations
Forest Fire Detection Dataset — Auto-Annotated
Bounding-box annotated version of touati-kamel/forest-fire-dataset,
built for training forest-fire / smoke / fog object detection models.
Overview
This dataset contains video frames auto-labeled with bounding boxes for fire and
smoke-related visual phenomena, using a zero-shot open-vocabulary object detector
(Grounding DINO). It is derived from the original touati-kamel/forest-fire-dataset image
classification dataset… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/forest-fire-annotations.ToothXpert.MM-OPG-Annotationsreal-resumes-section-detection-annotationsHyperNeRF-Annotation
This is the language query annotations for the HyperNeRF dataset, which are used in 4DLangSplat For original dataset, please visit https://github.com/google/hypernerf. For the usage of annotations, please visit https://github.com/zrporz/4DLangSplat
license: cc-by-nc-4.0
toy-car-annotation-YOLOHey everyone,
In my final year project, I created Smart Traffic Management System.The project was to manage traffic lights' delays based on the number of vehicles on road.I made everything worked using Raspberry Pi and pre-recorded videos but it was a "final year project", it was needed to be tested by changing videos frequently which was a kind of hustle. Collecting tons of videos and loading them in Pi was not too hard but it would have cost time, by every time changing names of videos in… See the full description on the dataset page: https://huggingface.co/datasets/tubasid/toy-car-annotation-YOLO.amz-image-annotationsGCA_suction_franka_annotationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 40,
"total_frames": 3964,
"total_tasks": 2,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hyzhang01/GCA_suction_franka_annotation.4k-video-annotations
4K Video Annotations — Shot Segmentation and Camera Motion
This dataset contains 12 frame-accurate shot clips segmented from five short cinematic video sequences. Every clip is paired with a detailed, manually reviewed annotation covering visible content, subject actions, shot scale, camera angle, camera movement, movement direction, stabilization, composition, lighting, color, pacing, transitions, timecodes, and technical properties.
The footage depicts a tense nighttime… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/4k-video-annotations.The-Oxford-IIIT-Pet-Dataset-With-Annotationsusd-side-coco-annotations
USD Side Detection Dataset (Front/Back)
A refined COCO-format dataset for detecting US Dollar currency and classifying whether the front or back side is visible.
Dataset Summary
Total Images: 3,618
Total Annotations: 3,746
Format: COCO + HuggingFace JSONL
Classes: 24 (denominations × front/back × authentic/counterfeit)
Classification Accuracy: 100% (all Front/Back classified)
Split
Images
Annotations
Train
2,671
2,738
Valid
597
627
Test
350
381… See the full description on the dataset page: https://huggingface.co/datasets/ebowwa/usd-side-coco-annotations.Neu3D-AnnotationThis is the language query annotations for the HyperNeRF dataset, which are used in 4DLangSplat For original dataset, please visit https://github.com/facebookresearch/Neural_3D_Video For the usage of annotations, please visit https://github.com/zrporz/4DLangSplat
license: cc-by-nc-4.0
ImageIn_annotations_resized_images
Dataset Card for ImageIn_annotations_resized_images
More Information needed
focus-frame-point-annotation
FOCUS — Foreign-Object Point Annotation
What you are doing
For every foreign object in a surgical frame, drop one point on it and pick its class.
You are not counting. We already have the counts. What the model is missing is
where the objects are — it knows "this frame contains 5" but not which 5 pixels,
which is exactly why it cannot learn to count.
Getting started (1 minute)
unzip shardN.zip -d shardN
cd shardN
python3 -m http.server 8000
Open… See the full description on the dataset page: https://huggingface.co/datasets/Potestates/focus-frame-point-annotation.White_Blood_Cells_with_annotationImageIn_annotationsInitial annotated dataset derived from ImageIN/IA_unlabelled
fruit-dataset-annotation
FruitDet – Object Detection Dataset
A community-driven fruit image dataset annotated for object detection, original repo.The images are real-world fruit photos collected from various environments. All images have been resized to 920×1080 pixels, and each fruit instance is labeled with a bounding box in YOLO format.
Features
Real-world images from diverse sources
19 fruit/vegetable categories, each containing 30 images
All images resized to 920×1080 pixels… See the full description on the dataset page: https://huggingface.co/datasets/sirunchained/fruit-dataset-annotation.OCR-Tibetan_line_segmentation_esukhia_coordinate_annotation
Data Splits
Split: without_superscript_subscript
Total Rows: 766,245
method
Type: categorical
Data Type: string
Unique Values: 1
Value
Count
Percentage
Transkribus
766,245
100.00%
Split: with_superscript_subscript
Total Rows: 263,346
method
Type: categorical
Data Type: string
Unique Values: 1
Value
Count
Percentage
Transkribus
263,346
100.00%
cc3m-grounded-annotations
CC3M grounded annotations
Region-level grounding for Conceptual Captions 3M: bounding boxes, the noun
phrase each box grounds, and the span of the caption that phrase came from, for
3,016,640 of CC3M's 3,318,333 rows.
No images here. This is metadata only, joinable onto a CC3M copy you already
have. That is the point of it: the grounding is 354 MB, the pixels are 125 GB.
Files
file
rows
size
annotations-0000..0482.parquet
3,016,640
199 MB… See the full description on the dataset page: https://huggingface.co/datasets/freek23/cc3m-grounded-annotations.OCR-Tibetan_line_segmentation_coordinate_annotation
Tibetan OCR Line Segementation Coordinates Annotation
This dataset contains coordinated annotation data for lines. It includes features related to text lines, image details, and processing methods used for data annotation.
Features
line_id: Text Line Information
line_coordinates: Coordinates of the text lines
source_image: Image filename or identifier
image_size: Size of the images in pixel
format: Image Format
bdrc_work_id: Identifier for BDRC work
image_url:… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/OCR-Tibetan_line_segmentation_coordinate_annotation.TriConflict-hallucination-annotationsstorage-for-data-annotation-ovariancaner
