datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-characters
Synthetic Characters Dataset
A synthetic image dataset generated with Flux Schnell featuring structured character prompts designed for training character generation, fashion understanding, and portrait synthesis models.
Recommended Filters:
Age
People Count
Hair color
Camera angle
Anime/Realistic
Nudity/Clothed
I'll likely recaption everything with a list of classifications attached to them for easy filtering later.
Dataset Description
This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/synthetic-characters.kuzushiji-dataset-characters
Dataset Card for Kuzushiji Character Dataset
Dataset Summary
The Kuzushiji Character Dataset contains individual character crops generated from full
page images and character-coordinate annotations. Crops retain their original pixel size;
they are not resized or padded. Each record keeps the source book, page, character ID,
Unicode label, original page box, and the clamped box actually used for cropping.
This card describes the generated repository… See the full description on the dataset page: https://huggingface.co/datasets/Kotomiya07/kuzushiji-dataset-characters.malayalam_characters
Malayalam Character Dataset
This dataset contains 136,344 images of Malayalam characters, covering 662 distinct classes.
Dataset Structure
The dataset is organized as a single split (train) with the following columns:
image_path: The character image (128x128 grayscale).
text: The Malayalam character/word represented in the image.
label_id: Integer label ID (0-661).
filename: Original filename.
Classes
The dataset includes:
Basic consonants and vowels… See the full description on the dataset page: https://huggingface.co/datasets/mangalathkedar/malayalam_characters.
