datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
game_character_skins
Game Character Skins Dataset
Summary
This comprehensive dataset contains game character skins and artwork from multiple popular mobile and PC games, providing a rich collection of character visual assets for computer vision research and game development applications. The dataset spans eight major game titles including Arknights, Azur Lane, Blue Archive, Fate/Grand Order, Genshin Impact, Girls' Frontline, Neural Cloud, Nikke, Path to Nowhere, and Honkai: Star Rail… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/game_character_skins.Devanagari-Characters-Image
Devanagari Characters Image Dataset
Dataset Summary
The Devanagari Characters Image Dataset is a high-resolution dataset designed to support research and experimentation in generative modeling, specifically for the Hindi script. It includes images for:
Vowels (स्वर)
Consonants (व्यंजन)
Matra combinations (e.g., का, कि, की, कु)
Hindi numerals (०-९)
The dataset was created to address the limitations of existing Devanagari datasets, which often suffer from low resolution… See the full description on the dataset page: https://huggingface.co/datasets/Mayank022/Devanagari-Characters-Image.generic_character_skins
Generic Character Skins Dataset
Summary
This comprehensive dataset provides an extensive collection of character images sourced from Zerochan across multiple popular anime, game, and manga franchises. The dataset contains meticulously organized character artwork spanning diverse genres including gacha games, idol franchises, fantasy series, and action RPGs. With over 2,000 character folders and thousands of high-quality images, this repository serves as a valuable… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/generic_character_skins.Devanagari-Characters-Image
Devanagari Characters Image Dataset
Dataset Summary
The Devanagari Characters Image Dataset is a high-resolution dataset designed to support research and experimentation in generative modeling, specifically for the Hindi script. It includes images for:
Vowels (स्वर)
Consonants (व्यंजन)
Matra combinations (e.g., का, कि, की, कु)
Hindi numerals (०-९)
The dataset was created to address the limitations of existing Devanagari datasets, which often suffer from low resolution… See the full description on the dataset page: https://huggingface.co/datasets/rhythmjain30/Devanagari-Characters-Image.synthetic-characters
Synthetic Characters Dataset
A synthetic image dataset generated with Flux Schnell featuring structured character prompts designed for training character generation, fashion understanding, and portrait synthesis models.
Recommended Filters:
Age
People Count
Hair color
Camera angle
Anime/Realistic
Nudity/Clothed
I'll likely recaption everything with a list of classifications attached to them for easy filtering later.
Dataset Description
This dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/synthetic-characters.kuzushiji-dataset-characters
Dataset Card for Kuzushiji Character Dataset
Dataset Summary
The Kuzushiji Character Dataset contains individual character crops generated from full
page images and character-coordinate annotations. Crops retain their original pixel size;
they are not resized or padded. Each record keeps the source book, page, character ID,
Unicode label, original page box, and the clamped box actually used for cropping.
This card describes the generated repository… See the full description on the dataset page: https://huggingface.co/datasets/Kotomiya07/kuzushiji-dataset-characters.2D_Video_Game_Cartoon_Character_Sprite-Sheets
Dataset Card for Dataset Name
Dataset Details
Experimental composition of 76 cartoon art-style video game character spritesheets. Resized to 512x512, mixed variation of animation styles.
Dataset Description
All images editted using Tiled image editting software as most assets are typically downloaded individually and not in sequence. I compiled each animation sequence into one img to display animations frame-by-frame evenly distributed across some common… See the full description on the dataset page: https://huggingface.co/datasets/mgane/2D_Video_Game_Cartoon_Character_Sprite-Sheets.Devanagari-Characters-Image
Devanagari Characters Image Dataset
Dataset Summary
The Devanagari Characters Image Dataset is a high-resolution dataset designed to support research and experimentation in generative modeling, specifically for the Hindi script. It includes images for:
Vowels (स्वर)
Consonants (व्यंजन)
Matra combinations (e.g., का, कि, की, कु)
Hindi numerals (०-९)
The dataset was created to address the limitations of existing Devanagari datasets, which often suffer from low resolution… See the full description on the dataset page: https://huggingface.co/datasets/Yash141414/Devanagari-Characters-Image.game_character_skins
Game Character Skins Dataset
Summary
This comprehensive dataset contains game character skins and artwork from multiple popular mobile and PC games, providing a rich collection of character visual assets for computer vision research and game development applications. The dataset spans eight major game titles including Arknights, Azur Lane, Blue Archive, Fate/Grand Order, Genshin Impact, Girls' Frontline, Neural Cloud, Nikke, Path to Nowhere, and Honkai: Star Rail… See the full description on the dataset page: https://huggingface.co/datasets/JosephPaul2244/game_character_skins.anime-demon-slayer-all-character-images
Demon Slayer Anime Full Character Images Dataset
This dataset contains images of all characters from the anime "Demon Slayer" sorted into different folders.
Original forum post:
https://diffused.to/Thread-Image-Anime-Demon-Slayer-Full-Character-Images-Dataset
Dataset collection date
Dec 2025
Total Characters:
19
Total Images:
554
Dataset structure:
├── 📂 256x256/
│ ├── 📂 001/
│ ├── 00001.jpg
│ ├── 00002.jpg
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/wallstoneai/anime-demon-slayer-all-character-images.Gujarati-Handwritten-Characters-Dataset
Gujarati-Handwritten-Characters-Dataset
Images: Contains scanned images of data collection sheets.
output_boxes: contains character wise separated images.
preprocessed_images: contains character wise separated images after preprocessing.
Note: In the files archive files(.zip) are also provided for individual downloads.
Github Repository: https://github.com/MKPTechnicals/Gujarati-Handwritten-Characters-Dataset
Kaggle Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Meet1265/Gujarati-Handwritten-Characters-Dataset.geez-characters
Amharic (Geʽez) Handwritten Character Dataset (32×32)
Dataset Details
Description
This dataset contains handwritten images of Amharic (Geʽez script) characters intended for character-level Optical Character Recognition (OCR) and handwriting recognition research.
Property
Value
Total Images
13,000+
Classes
287 distinct characters
Image Size
32 × 32 pixels
Format
Grayscale
Distribution
Balanced across all classes
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Yaredoffice/geez-characters.burmese-character-dataset
Burmese Character Dataset (Combined)
A combined image-classification dataset of handwritten and air-written Burmese
(Myanmar) characters and digits, merged from two publicly available Kaggle
datasets into a single imagefolder-compatible layout (one subfolder per
character class).
Classes: 54 (digits 0–9 plus Burmese consonant/vowel characters,
e.g. Ah, Aha, O, ba_kone, ga_nge, nga, za_kwear, ...)
Images: 13,269 JPG files
Size: ~106 MB
Source datasets & attribution… See the full description on the dataset page: https://huggingface.co/datasets/aungthuhein-dev/burmese-character-dataset.popular_anime_characterkidney-3dct-lesion-characterization
Kidney 3D-CT Lesion Characterization — runnable data
Companion data for the code repository
kidney-3dct-lesion-characterization
(paper: Multi-Granularity 3D Kidney Lesion Characterization from CT Volumes).
The private UF Health dataset used in the paper cannot be released (patient
privacy). This repository instead provides two small, fully runnable datasets so
the published code can be executed end-to-end:
Subset
What it is
Covers
License
synthetic/
Fake volumes +… See the full description on the dataset page: https://huggingface.co/datasets/LiangRenjie/kidney-3dct-lesion-characterization.burmese-pyu-character-recognition
Burmese Pyu Character Recognition Dataset
English | မြန်မာဘာသာ
Overview
This Dataset is an image collection created for the purpose of computer recognition (Character Recognition) and research of the "Pyu" alphabet, an ancient script of Myanmar.
Brief Historical Background
The Pyu people were one of the earliest major ethnic groups to inhabit Myanmar, settling in the region since the early AD periods. The Pyu culture is of great importance when studying the… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/burmese-pyu-character-recognition.malayalam_characters
Malayalam Character Dataset
This dataset contains 136,344 images of Malayalam characters, covering 662 distinct classes.
Dataset Structure
The dataset is organized as a single split (train) with the following columns:
image_path: The character image (128x128 grayscale).
text: The Malayalam character/word represented in the image.
label_id: Integer label ID (0-661).
filename: Original filename.
Classes
The dataset includes:
Basic consonants and vowels… See the full description on the dataset page: https://huggingface.co/datasets/mangalathkedar/malayalam_characters.db-sfw-512px-character-filter
Danbooru SFW 512px Character Filter
This dataset is meant to be used for training a simple binary classifier that can filter the
Danbooru SFW 2021 dataset. It is similar to db-sfw-512-general-filter-dataset
but it has different class criteria. Just like the general dataset there are two
classes: "accepted" and "rejected", with "accepted" representing samples that should pass
through the filter and "rejected" representing samples that should not.
To be accepted, a sample should… See the full description on the dataset page: https://huggingface.co/datasets/hayden-donnelly/db-sfw-512px-character-filter.anime-with-you-our-love-will-make-it-through-characters
With You Our Love Will Make It Through Anime Full Character Dataset
This dataset contains images of all characters from the anime "With you, Our Love will Make it Through" sorted into different folders.
Original forum post:
https://diffused.to/Thread-Image-Anime-With-You-Our-Love-Will-Make-It-Through-Full-Character-Dataset
Dataset collection date
Dec 2025
Total Characters:
3
Total Images:
109
Dataset structure:
├── 📂 256x256/
│… See the full description on the dataset page: https://huggingface.co/datasets/wallstoneai/anime-with-you-our-love-will-make-it-through-characters.zerochan-character-w640-ws-m50-full
Zerochan Single-Character Webdataset Full Dataset
This is the webdataset subset dataset for animetimm/zerochan-character-w640.
All the images here are guaranteed to be non-monochrome, single-person, single-headed, single-faced and have one primary character.
Can be used for single-label anime character classification training.
Images here are resized to min(width, height) <= 640.
Bounding boxes of faces, heads and persons are provided in the json metadata, in the form of (x0, y0… See the full description on the dataset page: https://huggingface.co/datasets/animetimm/zerochan-character-w640-ws-m50-full.zerochan-character-w640-ws-m200-100k
Zerochan Single-Character Webdataset 100k Sub-Dataset
This is the webdataset subset dataset for animetimm/zerochan-character-w640.
All the images here are guaranteed to be non-monochrome, single-person, single-headed, single-faced and have one primary character.
Can be used for single-label anime character classification training.
Images here are resized to min(width, height) <= 640.
Bounding boxes of faces, heads and persons are provided in the json metadata, in the form of (x0… See the full description on the dataset page: https://huggingface.co/datasets/animetimm/zerochan-character-w640-ws-m200-100k.
