datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PixelArt_Multiview
Multiview PixelArt
Dataset Summary
Contains sets of images representing a full 360° turnaround of characters, animals and objects in pixel art.
Each row contains 9 images from all angles.
Camera Data can be downloaded
Examples
Input (f1)
f2
f3
f4
f5
f6
f7
f8
f9
Input (f1)
f2
f3
f4
f5
f6
f7
f8
f9
Input (f1)
f2
f3
f4
f5
f6
f7
f8
f9
Input (f1)
f2
f3
f4
f5
f6
f7
f8
f9… See the full description on the dataset page: https://huggingface.co/datasets/Scaryplasmon96/PixelArt_Multiview.pixelart-308k
PixelArt-308K
A large-scale synthetic pixel art image-text dataset containing 308,765 256×256 RGB images, each paired with a text caption describing the image.
The dataset was created as a two-stage generation pipeline:
Caption generation — prompts were generated using Google's Gemma models through LM Studio.
Image generation — the generated prompts were used with FLUX.2-klein-4B to create the corresponding pixel art images.
The goal of this dataset is to provide a large… See the full description on the dataset page: https://huggingface.co/datasets/Chan-Y/pixelart-308k.test-HunyuanVideo-pixelart-videos
trojblue/test-HunyuanVideo-pixelart-images
👋 Heads up—this repository is just a PARTIAL dataset. For the full pixelart-images dataset, make sure to grab both parts:
Images Part
Video Part (this repo)
What's in the Dataset?
This dataset is all about anime-styled pixel art images that have been carefully selected to make your models shine. Here’s what makes these images special:
Rich in detail: Pixelated, yes—but still full of life and not overly simplified.… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/test-HunyuanVideo-pixelart-videos.diffusiondb-pixelartDiffusionDB is the first large-scale text-to-image prompt dataset. It contains 2
million images generated by Stable Diffusion using prompts and hyperparameters
specified by real users. The unprecedented scale and diversity of this
human-actuated dataset provide exciting research opportunities in understanding
the interplay between prompts and generative models, detecting deepfakes, and
designing human-AI interaction tools to help users more easily use these models.pixel-art-synthetic-10k
Pixel Art Dataset
A synthetic dataset of 9,852 pixel art images with text prompts.
Sample Preview
Image
Prompt
concept art of a silent hill monster. painted by edward hopper.
64×64 pixels, nearest-neighbor upscaled 4× for display
Dataset Details
Samples: 9,852
Grid size: 64×64 pixels
Palette: 256 colors (shared palette stored in palette.npy)
Base model: FLUX.1-dev with Retro-Pixel LoRA
Prompt style: Various art styles and subjects… See the full description on the dataset page: https://huggingface.co/datasets/Sc077y/pixel-art-synthetic-10k.pixel-art-character
Pixel Art Character Dataset
⚠️ CONTENT WARNING: This dataset contains partially NSFW content. Some images may include suggestive themes, violence, or mature content. Viewer discretion advised.
A dataset of 500 pixel art character sprites for training LoRA models.
License
Derived License: Apache 2.0
This dataset is provided under the Apache 2.0 License, inherited from the base models used for generation.
Copyright 2026 Limbicnation
Licensed under the Apache License… See the full description on the dataset page: https://huggingface.co/datasets/Limbicnation/pixel-art-character.free-to-use-pixelart
Free-to-use Pixel Art
Dataset Details
This dataset was collected on 25th May, 2024.
It's a small subset of the free-to-use images on PixilArt.
At the time of publication, this dataset was covered by permissive terms that allow commercial use.
Dataset Description
This dataset is unique in that it contains the pixel group size for each collected sample, which might assist in experiments on microconditioning inputs on an adapter to control this value of the unit… See the full description on the dataset page: https://huggingface.co/datasets/bghira/free-to-use-pixelart.Pixel_Art
Pixel Art
This dataset contains free-licensed images, downloaded from unsplash. Curated and created by:
Vadim Bogulov
Fujiphilm
SIMON LEE
pixel-art-nouns
Dataset Card for "pixel-art-nouns"
More Information needed
test-HunyuanVideo-pixelart-videos
Drive From trojblue/test-HunyuanVideo-pixelart-videos
Reorganized version of Wild-Heart/Disney-VideoGeneration-Dataset. This is needed for Mochi-1 fine-tuning.
pixel-art-bench-v1
Pixel Art Benchmark Dataset (Source)
The Pixel Art Benchmark Dataset is a structured collection of pixel-art outputs generated by large language models (LLMs). Each sample consists of a discrete color palette and a grid-based representation of pixel art, along with generation metadata such as token usage, cost, and model provenance.
Each row in the dataset represents a single generated pixel-art sample.
Encoding Details
Each string in grid represents one row of pixels.… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/pixel-art-bench-v1.trojblue-pixelart-imagesA copy of https://huggingface.co/datasets/trojblue/test-HunyuanVideo-pixelart-images to make it compatible for training with https://github.com/a-r-r-o-w/finetrainers (until more dataset formats are supported).
trojblue-pixelart-videosA copy of https://huggingface.co/datasets/trojblue/test-HunyuanVideo-pixelart-videos to make it compatible for training with https://github.com/a-r-r-o-w/finetrainers (until more dataset formats are supported).
160-LegoBot-lego_pixelart_stage_1_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 16,
"total_frames": 4747,
"total_tasks": 1,
"total_videos": 32,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:16"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_1_v2.diffusiondb-pixelart-v2
DiffusionDB-Pixelart
Dataset Summary
This is a subset of the DiffusionDB 2M dataset which has been turned into pixel-style art.
DiffusionDB is the first large-scale text-to-image prompt dataset. It contains 14 million images generated by Stable Diffusion using prompts and hyperparameters specified by real users.
DiffusionDB is publicly available at 🤗 Hugging Face Dataset.
Supported Tasks and Leaderboards
The unprecedented scale and diversity of this… See the full description on the dataset page: https://huggingface.co/datasets/Clawffice/diffusiondb-pixelart-v2.lego_pixelart_stage_1_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 28,
"total_frames": 10945,
"total_tasks": 1,
"total_videos": 56,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:28"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/lego_pixelart_stage_1_v3.160-LegoBot-lego_pixelart_stage_2_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower_bimanual",
"total_episodes": 8,
"total_frames": 3989,
"total_tasks": 1,
"total_videos": 32,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:8"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_2_v2.pixelart-98kpixel-art-bench-lite
🎨 Pixel Art Bench Lite
Pixel Art Bench Lite is a structured-output benchmark designed to evaluate small language models on their ability to generate valid, interpretable, and semantically meaningful JSON outputs under strict constraints.
The benchmark is based on Pixel Art Bench focuses on pixel art generation over a fixed 24×24 grid, requiring models to produce outputs that are syntactically correct but also visually coherent.
While many benchmarks evaluate free-form text… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/pixel-art-bench-lite.captioned-pixelart-palette160-LegoBot-lego_pixelart_stage_2_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower_bimanual",
"total_episodes": 31,
"total_frames": 18204,
"total_tasks": 1,
"total_videos": 93,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:31"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_2_v1.pixel-art-nouns-2k
Dataset Card for "pixel-art-nouns-2k"
More Information needed
160-LegoBot-lego_pixelart_stage_2_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower_bimanual",
"total_episodes": 52,
"total_frames": 28727,
"total_tasks": 1,
"total_videos": 208,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:52"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_2_v3.160-LegoBot-lego_pixelart_stage_1_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 21,
"total_frames": 8064,
"total_tasks": 1,
"total_videos": 42,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_1_v1.HunyuanVideo-pixelart-videos-sample
trojblue/test-HunyuanVideo-pixelart-images
👋 Heads up—this repository is just a PARTIAL dataset. For the full pixelart-images dataset, make sure to grab both parts:
Images Part
Video Part (this repo)
What's in the Dataset?
This dataset is all about anime-styled pixel art images that have been carefully selected to make your models shine. Here’s what makes these images special:
Rich in detail: Pixelated, yes—but still full of life and not overly simplified.… See the full description on the dataset page: https://huggingface.co/datasets/inlineresearch/HunyuanVideo-pixelart-videos-sample.160-LegoBot-lego_pixelart_stage_1_point_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 54,
"total_frames": 23460,
"total_tasks": 1,
"total_videos": 108,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:54"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_1_point_v1.test-HunyuanVideo-pixelart-images
trojblue/test-HunyuanVideo-pixelart-images
Hey there! 👋 Heads up—this repository is just a PARTIAL dataset. For the full pixelart-images dataset, make sure to grab both parts:
Images Part (this repo)
Video Part
This dataset is a collection of anime-style pixel art images and is perfect for debugging general anime text-to-image (T2I) training or testing Hunyuan Video models. 🎨
What's in the Dataset?
This dataset is all about anime-styled pixel art images that have… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/test-HunyuanVideo-pixelart-images.captioned-pixelart-cannytopdown-medieval-pixelart
Dataset Card for Top-Down Medieval Pixel Art
Quick Facts
Images
100
Resolution
1024 × 1024 px
Format
PNG, white background
Asset families
structure 97 / vegetation 2 / terrain 1
License
Non-commercial, with Open RAIL-M usage restrictions (see §License)
Dataset Description
Top-Down Medieval Pixel Art is a synthetic dataset of 100 top-down
(overhead) medieval fantasy pixel-art game assets, each paired with a
descriptive… See the full description on the dataset page: https://huggingface.co/datasets/stixxert/topdown-medieval-pixelart.160-LegoBot-lego_pixelart_stage_2_v1
160_legobot_lego_pixelart_stage_2_v1
This dataset was converted from
LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_2_v1
to Apache TsFile format.
Dataset Description
The original dataset is a LeRobot-format robot dataset for a LegoBot pixel-art
stage task.
Modalities: Time-series. The original dataset also includes video streams.
Robot type: so100_follower_bimanual
Codebase version: LeRobot v2.1
Recording frequency: 30 Hz
Split: train (0:53)
Scale:… See the full description on the dataset page: https://huggingface.co/datasets/THULab/160-LegoBot-lego_pixelart_stage_2_v1.
