datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PittImageVideoAdsDataset
Dataset Card for PittImageVideoAdsDataset
Dataset Summary
PittImageVideoAdsDataset is the image and video advertisement dataset released with Automatic Understanding of Image and Video Advertisements. The paper reports 64,832 image advertisements and 3,477 YouTube advertisement videos, with human annotations for topics, sentiments, slogans, persuasive strategies, symbolic references, and action/reason Q/A. This Hugging Face version exposes the public annotation… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PittImageVideoAdsDataset.devanagari_ocr_graphemes
Devanagari OCR Grapheme Dataset
This repository hosts a grapheme‑level OCR dataset for the Devanagari script. Each data point consists of an image of a single grapheme and its corresponding Unicode text.
Dataset Structure
.
├── data/ # Directory containing all PNG images (e.g., 00000.png, 00001.png, ...)
├── devanagari_ocr_graphemes.json # ShareGPT‑formatted JSON file
└── README.md # This file
data/ – Images are stored in… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/devanagari_ocr_graphemes.graph_dataset_generated_v2
Graph Dataset - Image & LabelMe & OBB Annotation (Train/Val Split)
Dataset Overview
Comprehensive graph/chart detection dataset with ground truth LabelMe polygon annotations and OBB (Oriented Bounding Box) data, split into training and validation sets.
Total examples: 35561 image-annotation pairs
Train: 28448 (80.0%)
Validation: 7113 (20.0%)
Total size: 2134.30 MB
Language: Khmer (km)
Document types: Graph/Chart documents
Ground truth: LabelMe polygon annotations… See the full description on the dataset page: https://huggingface.co/datasets/vichetkao/graph_dataset_generated_v2.
