datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
od-syn-page-annotations-com
📦 Dhivehi Synthetic Document Layout + Textline Dataset
This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis, visual document understanding, OCR fine-tuning, and related tasks specifically for Dhivehi script.
Note: this version image are compressed.
Raw version 📁 Repository: Hugging Face Datasets
📋 Dataset Summary
Total Examples: ~58… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations-com.ImageIn_annotations_resized_images
Dataset Card for ImageIn_annotations_resized_images
More Information needed
ImageIn_annotationsInitial annotated dataset derived from ImageIN/IA_unlabelled
scent-engine-annotations
ScentEngine — Universal Training Annotations
Two-stage Gemma 4 labels (E4B primary, 31B verifier for conf ∈ [0.4, 0.7])
over public game screenshots. Each row maps an image to 6 PWM values
(0–100) in locked channel order: Flora · Aqua · Earth · Pyric · Ozone · Civic.
Files
File
Purpose
public_games_final.jsonl
All accepted labels (quarantine excluded)
universal_train.jsonl
90% split for training the universal head
universal_val.jsonl
10% split for… See the full description on the dataset page: https://huggingface.co/datasets/Dennishuang85/scent-engine-annotations.
