datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Devanagari-Characters-Image
Devanagari Characters Image Dataset
Dataset Summary
The Devanagari Characters Image Dataset is a high-resolution dataset designed to support research and experimentation in generative modeling, specifically for the Hindi script. It includes images for:
Vowels (स्वर)
Consonants (व्यंजन)
Matra combinations (e.g., का, कि, की, कु)
Hindi numerals (०-९)
The dataset was created to address the limitations of existing Devanagari datasets, which often suffer from low resolution… See the full description on the dataset page: https://huggingface.co/datasets/Mayank022/Devanagari-Characters-Image.Devanagari-Characters-Image
Devanagari Characters Image Dataset
Dataset Summary
The Devanagari Characters Image Dataset is a high-resolution dataset designed to support research and experimentation in generative modeling, specifically for the Hindi script. It includes images for:
Vowels (स्वर)
Consonants (व्यंजन)
Matra combinations (e.g., का, कि, की, कु)
Hindi numerals (०-९)
The dataset was created to address the limitations of existing Devanagari datasets, which often suffer from low resolution… See the full description on the dataset page: https://huggingface.co/datasets/rhythmjain30/Devanagari-Characters-Image.Devanagari-Characters-Image
Devanagari Characters Image Dataset
Dataset Summary
The Devanagari Characters Image Dataset is a high-resolution dataset designed to support research and experimentation in generative modeling, specifically for the Hindi script. It includes images for:
Vowels (स्वर)
Consonants (व्यंजन)
Matra combinations (e.g., का, कि, की, कु)
Hindi numerals (०-९)
The dataset was created to address the limitations of existing Devanagari datasets, which often suffer from low resolution… See the full description on the dataset page: https://huggingface.co/datasets/Yash141414/Devanagari-Characters-Image.synthetic-characters
Synthetic Characters Dataset
A synthetic image dataset generated with Flux Schnell featuring structured character prompts designed for training character generation, fashion understanding, and portrait synthesis models.
Recommended Filters:
Age
People Count
Hair color
Camera angle
Anime/Realistic
Nudity/Clothed
I'll likely recaption everything with a list of classifications attached to them for easy filtering later.
Dataset Description
This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/synthetic-characters.kuzushiji-dataset-characters
Dataset Card for Kuzushiji Character Dataset
Dataset Summary
The Kuzushiji Character Dataset contains individual character crops generated from full
page images and character-coordinate annotations. Crops retain their original pixel size;
they are not resized or padded. Each record keeps the source book, page, character ID,
Unicode label, original page box, and the clamped box actually used for cropping.
This card describes the generated repository… See the full description on the dataset page: https://huggingface.co/datasets/Kotomiya07/kuzushiji-dataset-characters.Gujarati-Handwritten-Characters-Dataset
Gujarati-Handwritten-Characters-Dataset
Images: Contains scanned images of data collection sheets.
output_boxes: contains character wise separated images.
preprocessed_images: contains character wise separated images after preprocessing.
Note: In the files archive files(.zip) are also provided for individual downloads.
Github Repository: https://github.com/MKPTechnicals/Gujarati-Handwritten-Characters-Dataset
Kaggle Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Meet1265/Gujarati-Handwritten-Characters-Dataset.geez-characters
Amharic (Geʽez) Handwritten Character Dataset (32×32)
Dataset Details
Description
This dataset contains handwritten images of Amharic (Geʽez script) characters intended for character-level Optical Character Recognition (OCR) and handwriting recognition research.
Property
Value
Total Images
13,000+
Classes
287 distinct characters
Image Size
32 × 32 pixels
Format
Grayscale
Distribution
Balanced across all classes
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Yaredoffice/geez-characters.malayalam_characters
Malayalam Character Dataset
This dataset contains 136,344 images of Malayalam characters, covering 662 distinct classes.
Dataset Structure
The dataset is organized as a single split (train) with the following columns:
image_path: The character image (128x128 grayscale).
text: The Malayalam character/word represented in the image.
label_id: Integer label ID (0-661).
filename: Original filename.
Classes
The dataset includes:
Basic consonants and vowels… See the full description on the dataset page: https://huggingface.co/datasets/mangalathkedar/malayalam_characters.anime-with-you-our-love-will-make-it-through-characters
With You Our Love Will Make It Through Anime Full Character Dataset
This dataset contains images of all characters from the anime "With you, Our Love will Make it Through" sorted into different folders.
Original forum post:
https://diffused.to/Thread-Image-Anime-With-You-Our-Love-Will-Make-It-Through-Full-Character-Dataset
Dataset collection date
Dec 2025
Total Characters:
3
Total Images:
109
Dataset structure:
├── 📂 256x256/
│… See the full description on the dataset page: https://huggingface.co/datasets/wallstoneai/anime-with-you-our-love-will-make-it-through-characters.
