datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PixelArt_Multiview
Multiview PixelArt
Dataset Summary
Contains sets of images representing a full 360° turnaround of characters, animals and objects in pixel art.
Each row contains 9 images from all angles.
Camera Data can be downloaded
Examples
Input (f1)
f2
f3
f4
f5
f6
f7
f8
f9
Input (f1)
f2
f3
f4
f5
f6
f7
f8
f9
Input (f1)
f2
f3
f4
f5
f6
f7
f8
f9
Input (f1)
f2
f3
f4
f5
f6
f7
f8
f9… See the full description on the dataset page: https://huggingface.co/datasets/Scaryplasmon96/PixelArt_Multiview.pixelart-308k
PixelArt-308K
A large-scale synthetic pixel art image-text dataset containing 308,765 256×256 RGB images, each paired with a text caption describing the image.
The dataset was created as a two-stage generation pipeline:
Caption generation — prompts were generated using Google's Gemma models through LM Studio.
Image generation — the generated prompts were used with FLUX.2-klein-4B to create the corresponding pixel art images.
The goal of this dataset is to provide a large… See the full description on the dataset page: https://huggingface.co/datasets/Chan-Y/pixelart-308k.test-HunyuanVideo-pixelart-videos
trojblue/test-HunyuanVideo-pixelart-images
👋 Heads up—this repository is just a PARTIAL dataset. For the full pixelart-images dataset, make sure to grab both parts:
Images Part
Video Part (this repo)
What's in the Dataset?
This dataset is all about anime-styled pixel art images that have been carefully selected to make your models shine. Here’s what makes these images special:
Rich in detail: Pixelated, yes—but still full of life and not overly simplified.… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/test-HunyuanVideo-pixelart-videos.lpc-4view-pixel-art-diffusion
LPC 4-View Pixel Art Diffusion Dataset
This dataset provides LPC-style pixel-art character sprites for training
diffusion models for unconditional and text-to-image character generation.
It is developed as part of an undergraduate Third Year Project and is intended
for research and educational use.
The accompanying training code, preprocessing scripts, and experiments are
available in the project GitHub repository:
PIXEL-T2I.
Dataset Overview
The dataset consists… See the full description on the dataset page: https://huggingface.co/datasets/carlosuperb/lpc-4view-pixel-art-diffusion.diffusiondb-pixelartDiffusionDB is the first large-scale text-to-image prompt dataset. It contains 2
million images generated by Stable Diffusion using prompts and hyperparameters
specified by real users. The unprecedented scale and diversity of this
human-actuated dataset provide exciting research opportunities in understanding
the interplay between prompts and generative models, detecting deepfakes, and
designing human-AI interaction tools to help users more easily use these models.pixel-art-synthetic-10k
Pixel Art Dataset
A synthetic dataset of 9,852 pixel art images with text prompts.
Sample Preview
Image
Prompt
concept art of a silent hill monster. painted by edward hopper.
64×64 pixels, nearest-neighbor upscaled 4× for display
Dataset Details
Samples: 9,852
Grid size: 64×64 pixels
Palette: 256 colors (shared palette stored in palette.npy)
Base model: FLUX.1-dev with Retro-Pixel LoRA
Prompt style: Various art styles and subjects… See the full description on the dataset page: https://huggingface.co/datasets/Sc077y/pixel-art-synthetic-10k.pixel-art-character
Pixel Art Character Dataset
⚠️ CONTENT WARNING: This dataset contains partially NSFW content. Some images may include suggestive themes, violence, or mature content. Viewer discretion advised.
A dataset of 500 pixel art character sprites for training LoRA models.
License
Derived License: Apache 2.0
This dataset is provided under the Apache 2.0 License, inherited from the base models used for generation.
Copyright 2026 Limbicnation
Licensed under the Apache License… See the full description on the dataset page: https://huggingface.co/datasets/Limbicnation/pixel-art-character.free-to-use-pixelart
Free-to-use Pixel Art
Dataset Details
This dataset was collected on 25th May, 2024.
It's a small subset of the free-to-use images on PixilArt.
At the time of publication, this dataset was covered by permissive terms that allow commercial use.
Dataset Description
This dataset is unique in that it contains the pixel group size for each collected sample, which might assist in experiments on microconditioning inputs on an adapter to control this value of the unit… See the full description on the dataset page: https://huggingface.co/datasets/bghira/free-to-use-pixelart.Pixel_Art
Pixel Art
This dataset contains free-licensed images, downloaded from unsplash. Curated and created by:
Vadim Bogulov
Fujiphilm
SIMON LEE
pixel-art-nouns
Dataset Card for "pixel-art-nouns"
More Information needed
test-HunyuanVideo-pixelart-videos
Drive From trojblue/test-HunyuanVideo-pixelart-videos
Reorganized version of Wild-Heart/Disney-VideoGeneration-Dataset. This is needed for Mochi-1 fine-tuning.
draw_pixel_artThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 50,
"total_frames": 26066,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Beegbrain/draw_pixel_art.pick-style-pixel-art
Margin-aware Preference Optimization for Aligning Diffusion Models without Reference
We propose MaPO, a reference-free, sample-efficient, memory-friendly alignment technique for text-to-image diffusion models. For more details on the technique, please refer to our paper here.
Developed by
Jiwoo Hong* (KAIST AI)
Sayak Paul* (Hugging Face)
Noah Lee (KAIST AI)
Kashif Rasul (Hugging Face)
James Thorne (KAIST AI)
Jongheon Jeong (Korea University)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/mapo-t2i/pick-style-pixel-art.pokemon-pixel-artThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.lpc-action-pixel-art-diffusion
LPC Action Pixel Art Diffusion Dataset
This dataset provides LPC-style action spritesheets for training
image-conditional diffusion models, where a 4-view character image is used
as the conditioning input and an action spritesheet is generated as the output.
It is developed as part of an undergraduate Third Year Project and is intended
for research and educational use.
The accompanying training code, preprocessing scripts, and experiments are
available in the project GitHub… See the full description on the dataset page: https://huggingface.co/datasets/carlosuperb/lpc-action-pixel-art-diffusion.anime-pixel-art-animefacesv2
Anime Faces Pixelized
Pixel-art conversion of images generated from data/pixelized/bob80333-animefacesv2.
This dataset was prepared with anime-pixel-dataset and the upstream
WuZongWei6/Pixelization implementation. Review the source dataset license
and Pixelization's non-commercial scientific research license before publishing
or using the uploaded result.
Files
Split: train
Images: 92538
Parquet shards: 36
Each row contains only:
image: decoded by Hugging Face Datasets… See the full description on the dataset page: https://huggingface.co/datasets/nullHawk/anime-pixel-art-animefacesv2.pixel-art-bench-v1
Pixel Art Benchmark Dataset (Source)
The Pixel Art Benchmark Dataset is a structured collection of pixel-art outputs generated by large language models (LLMs). Each sample consists of a discrete color palette and a grid-based representation of pixel art, along with generation metadata such as token usage, cost, and model provenance.
Each row in the dataset represents a single generated pixel-art sample.
Encoding Details
Each string in grid represents one row of pixels.… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/pixel-art-bench-v1.trojblue-pixelart-imagesA copy of https://huggingface.co/datasets/trojblue/test-HunyuanVideo-pixelart-images to make it compatible for training with https://github.com/a-r-r-o-w/finetrainers (until more dataset formats are supported).
trojblue-pixelart-videosA copy of https://huggingface.co/datasets/trojblue/test-HunyuanVideo-pixelart-videos to make it compatible for training with https://github.com/a-r-r-o-w/finetrainers (until more dataset formats are supported).
160-LegoBot-lego_pixelart_stage_1_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 16,
"total_frames": 4747,
"total_tasks": 1,
"total_videos": 32,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:16"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_1_v2.diffusiondb-pixelart-v2
DiffusionDB-Pixelart
Dataset Summary
This is a subset of the DiffusionDB 2M dataset which has been turned into pixel-style art.
DiffusionDB is the first large-scale text-to-image prompt dataset. It contains 14 million images generated by Stable Diffusion using prompts and hyperparameters specified by real users.
DiffusionDB is publicly available at 🤗 Hugging Face Dataset.
Supported Tasks and Leaderboards
The unprecedented scale and diversity of this… See the full description on the dataset page: https://huggingface.co/datasets/Clawffice/diffusiondb-pixelart-v2.lego_pixelart_stage_1_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 28,
"total_frames": 10945,
"total_tasks": 1,
"total_videos": 56,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:28"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/lego_pixelart_stage_1_v3.160-LegoBot-lego_pixelart_stage_2_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower_bimanual",
"total_episodes": 8,
"total_frames": 3989,
"total_tasks": 1,
"total_videos": 32,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:8"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_2_v2.anime-pixel-art-2
Anime Faces Pixelized
Pixel-art conversion of images generated from aipracticecafe/anime-faces-256px.
This dataset was prepared with WuZongWei6/Pixelization implementation.
Files
Split: train
Images: 48167
Parquet shards: 7
Each row contains only:
image: decoded by Hugging Face Datasets as an image feature.
tags: source text tags/caption for that image.
pixelart-98kpixel-art-bench-lite
🎨 Pixel Art Bench Lite
Pixel Art Bench Lite is a structured-output benchmark designed to evaluate small language models on their ability to generate valid, interpretable, and semantically meaningful JSON outputs under strict constraints.
The benchmark is based on Pixel Art Bench focuses on pixel art generation over a fixed 24×24 grid, requiring models to produce outputs that are syntactically correct but also visually coherent.
While many benchmarks evaluate free-form text… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/pixel-art-bench-lite.captioned-pixelart-palette160-LegoBot-lego_pixelart_stage_2_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower_bimanual",
"total_episodes": 31,
"total_frames": 18204,
"total_tasks": 1,
"total_videos": 93,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:31"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_2_v1.pixel-art-nouns-2k
Dataset Card for "pixel-art-nouns-2k"
More Information needed
160-LegoBot-lego_pixelart_stage_2_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower_bimanual",
"total_episodes": 52,
"total_frames": 28727,
"total_tasks": 1,
"total_videos": 208,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:52"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_2_v3.
