datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anime-art-curated
anime-art-curated
A curated WebDataset of anime / digital illustration images with Danbooru-style
text tags, intended as a clean training corpus for text-to-image and
illustration-style model fine-tuning.
Provenance
This dataset is a filtered re-publication of an earlier dataset
(advokat/artist390k, now deleted) which was itself algorithmically curated
from broader booru sources by selecting popular images. The original source
inadvertently contained material that the… See the full description on the dataset page: https://huggingface.co/datasets/advokat/anime-art-curated.anime-bgAnime background and wallpaper images based on skytnt/anime-segmentation.
Archived with indexed tar files, you can easily download any of these images with hfutils or dghs-imgutils library.
For example:
from imgutils.resource import get_bg_image_file, random_image
# get and download this background image file
# return value should be the local path of given file
get_bg_image_file('000001.jpg')
# random select one background image from deepghs/anime-bg
random_image()
See Documentation of… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/anime-bg.pexels-tagger-v0-w640-ws-full
Pexels Tagger V0 Webdataset Full Dataset
This is the webdataset dataset for animetimm/pexels-wdtagger-w640.
Images here are resized to min(width, height) <= 640.
How to Use It
from datasets import load_dataset
dataset = load_dataset('animetimm/pexels-tagger-v0-w640-ws-full')
print(dataset["train"][0])
Images
3122908 images in total.
Split
Image Count
Total Size
train
2810634
184 GB
test
156409
10.2 GB
val
155865
10.2 GB
Tags… See the full description on the dataset page: https://huggingface.co/datasets/animetimm/pexels-tagger-v0-w640-ws-full.Anime-2026animer_data_backupanimetimm-Danbooru-VLManime_oldFigures-Plushies-Animeanime_dump_redditanimevox-wdsanime-collectionanimepics
AnimePics
This dataset is a pure image dataset in WebDataset format, designed for pre-training anime-style models.
It can also be used to evaluate the effect of additional pre-training on backbones trained primarily on real-world images when applied to datasets of a completely different nature.
The dataset is designed to be continuously updated, leveraging the features of WebDataset.
Dataset Details
Dataset Description
The animepics dataset is a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/zenless-fab/animepics.anime-classification-v1.5
Anime Image Classification Dataset (v1.5)
This is the webdataset dataset, containing 323060 images in total.
Images here are resized to min(width, height) <= 640.
How to Use It
from datasets import load_dataset
dataset = load_dataset('just-a-try/anime-classification-v1.5')
print(dataset["train"][0])
Images
323060 images in total.
Split
Image Count
Total Size
train
257996
14.5 GB
test
32506
1.83 GB
val
32558
1.84 GB
Class
Image Count… See the full description on the dataset page: https://huggingface.co/datasets/just-a-try/anime-classification-v1.5.animevox-vb-wdsanime_1080p_cliped
