datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
typed_digital_signatures
Typed Digital Signatures Dataset
This comprehensive dataset contains synthetic digital signatures rendered across 30 different Google Fonts, specifically selected for their handwriting and signature-style characteristics. Each font contributes unique stylistic elements, making this dataset ideal for robust signature analysis and font recognition tasks.
Dataset Overview
Total Fonts: 30 different Google Fonts
Images per Font: 3,000 signatures
Total Dataset Size:… See the full description on the dataset page: https://huggingface.co/datasets/Benjy/typed_digital_signatures.typeface_dataset
Typeface Dataset
Designed by COLIGNUM/webXOS (x.com/colignum)
Under Development.
github.com/webxos for more info.
Dataset Information
Name: colignum_typeface_dataset_512px_2026-01-22T01-21-36-081Z
Total Characters: 155
Resolution: 512×512 pixels
Format: PNG + CSV/Parquet
Generated: 2026-01-22
Character Sets Included
A-Z Uppercase (26 characters)
a-z Lowercase (26 characters)
0-9 Numbers (10 characters)
Symbols (32 characters):… See the full description on the dataset page: https://huggingface.co/datasets/webxos/typeface_dataset.Logo_typecav3_t-type_calcium_channels_butkiewicz-multimodalMMMU_img_typetypes-of-film-shots
What a Shot!
2,919 film frames labeled with shot scale (8 classes), from film-grab.com.
annotator
count
how labeled
human
863
hand-labeled (gold)
ai
2,056
DINOv2 classifier; low-confidence frames re-judged by Claude Opus (active learning)
Columns
image, label — ambiguous, closeUp, detail, extremeLongShot, fullShot, longShot, mediumCloseUp, mediumShot
annotator — human or ai
source — human / v2_highconf / opus_review / opus_resolved_amb… See the full description on the dataset page: https://huggingface.co/datasets/szymonrucinski/types-of-film-shots.TYPERAhair_type_comparison_resultsmyanmar_typeset_dictionary_OCR
Myanmar Typeset Dictionary OCR Dataset
This is a synthetically generated, realistically formatted dataset modeling a Myanmar-Myanmar dictionary. It is designed for training and validating OCR models, Document Layout Analysis (DLA) pipelines, and structural key-value extraction models.
The dataset contains a highly diverse set of pages containing multiple column flows, tabular glossaries, running headers/footers, realistic backgrounds, and dynamic typography (four fonts paired… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_typeset_dictionary_OCR.south-africa-crop-type-clouds
Dataset Card for South Africa Crop Type Clouds
This dataset contains the cloud masks generated and used for the paper KAN You See It? KANs and Sentinel for Effective and Explainable Crop Field Segmentation.
Curated by: Daniele Rege Cambrin
License: OpenRAIL
Uses
The dataset will provide a quality assessment for Sentinel-2 images of the South Africa Crop Type dataset.
Since MSI is ineffective through clouds, it was used to exclude samples that contain a large portion… See the full description on the dataset page: https://huggingface.co/datasets/DarthReca/south-africa-crop-type-clouds.eo-crop-type-belgium
Crop Type Segmentation Earth Observation Dataset of Belgium
This dataset is a ML ready dataset with Earth Observation data from ESA Sentinel-2 with a GSD of 10m. The segmentation labels are given for the crop-type in each 10m pixel. The spectral bands provided are those suggest in ESA WorldCereal documentation (Bands: 2,3,4,5,6,7,8,11,12). The dataset covers the majority of Belgium. Each image is 256x256 pixels.
Author of this dataset: Robert Cowlishaw (0x365)
Input data -… See the full description on the dataset page: https://huggingface.co/datasets/0x365/eo-crop-type-belgium.Document-Type-Detection
Document-Type-Detection
Dataset Summary
The Document-Type-Detection dataset is a large-scale image classification dataset consisting of scanned or photographed document images. Each image is categorized into one of nine document types. This dataset is ideal for training document classification models in finance, administration, OCR, and automation workflows.
Supported Tasks
Multiclass Document Classification
Classify an input document image into one of the… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Document-Type-Detection.Minerals_type_images_labelehri-danish-typewritten-lines
EHRI Danish Typewritten Lines
This dataset contains 1,007 cropped line images from 36 Danish typewritten pages in the EHRI Dataset. The source material consists of Danish diplomatic reports from the Second World War held by the Danish National Archives.
It is published as a ready-to-use line recognition evaluation dataset. It is typewritten material, not book or newspaper typesetting.
Dataset structure
The dataset has one test split because the upstream dataset… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/ehri-danish-typewritten-lines.GreenLabel-Waste-TypesOriginal Datasets: https://www.kaggle.com/datasets/techsash/waste-classification-data?select=DATASET
Sanskrit-OCR-Typed-Dataset
Sanskrit OCR Dataset
This dataset contains Sanskrit text images paired with their corresponding text labels, designed for OCR (Optical Character Recognition) tasks.
Dataset Structure
The dataset is split into training and validation sets:
Training set: Contains unique Sanskrit text images
Validation set: Contains separate unique Sanskrit text images
Features
image: The image containing Sanskrit text
label: The corresponding Sanskrit text label
filename:… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Sanskrit-OCR-Typed-Dataset.Minerals_type_images_6114TypeCoffee_16x16qode-vehicle-typescloud-types
Dataset Card for cloud-types
** The original COCO dataset is stored at dataset.tar.gz**
Dataset Summary
cloud-types
Supported Tasks and Leaderboards
object-detection: The dataset can be used to train a model for Object Detection.
Languages
English
Dataset Structure
Data Instances
A data point comprises an image and its object annotations.
{
'image_id': 15,
'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB… See the full description on the dataset page: https://huggingface.co/datasets/Francesco/cloud-types.document-type-detectionANIMATION_TYPES_3.05-Flower-Types-Classification-Datasetlux-typed-docs
Yale LUX dots.ocr layout/OCR outputs
Structured OCR/layout output from the rednote-hilab/dots.ocr model for Yale LUX document images. Each row carries the OCR output, the source image_url, the canvas_index, and a pointer to its manifest.
Rows: 626,586.
Built from the Yale LUX manifest processing database.
Tooth-Agenesis-6_Types
Tooth-Agenesis-6_Types
This dataset contains labeled dental images intended for training machine learning models on the classification of six different oral health conditions. It is designed for image classification tasks, specifically targeting types of tooth agenesis and other related dental conditions.
Dataset Summary
Total Images: 12,320
Format: Parquet
Modality: Image
Split: train only
Size: 183 MB
Language: English
License: Apache 2.0
Labels
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/strangerguardhf/Tooth-Agenesis-6_Types.pokemon-type-captions
Dataset Card for Pokémon type captions
Contains official artwork and type-specific caption for Pokémon #1-898 (Bulbasaur-Calyrex).
Each Pokémon is represented once by the default form from PokéAPI
Each row contains image and text keys:
image is a 475x475 PIL jpg of the Pokémon's official artwork.
text is a label describing the Pokémon by its type(s)
Attributions
Images and typing information pulled from PokéAPI
Based on the Lambda Labs Pokémon Blip Captions Dataset
Death_SE_type_datasettyped_final_chart_to_table
Dataset Card for "typed_final_chart_to_table"
More Information needed
pokemon-typesType2_dataset_235
