datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
better_daily_dialogpixelprose
From Pixels to Prose: A Large Dataset of Dense Image Captions
[ arXiv paper ] | [ 🌮 image tars ]
PixelProse is a comprehensive dataset of over 16M (million) synthetically generated captions,
leveraging cutting-edge vision-language models (Gemini 1.0 Pro Vision) for detailed and accurate descriptions.
1. Details
Total number of image-caption pairs: 16,896,214 (16.9M)
6,538,898 (6.5M) pairs in the split of CommonPool
9,066,455 (9.1M) pairs in the split of CC12M
1,290… See the full description on the dataset page: https://huggingface.co/datasets/tomg-group-umd/pixelprose.test-HunyuanVideo-pixelart-videos
trojblue/test-HunyuanVideo-pixelart-images
👋 Heads up—this repository is just a PARTIAL dataset. For the full pixelart-images dataset, make sure to grab both parts:
Images Part
Video Part (this repo)
What's in the Dataset?
This dataset is all about anime-styled pixel art images that have been carefully selected to make your models shine. Here’s what makes these images special:
Rich in detail: Pixelated, yes—but still full of life and not overly simplified.… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/test-HunyuanVideo-pixelart-videos.pixelprose
From Pixels to Prose: A Large Dataset of Dense Image Captions
[[ arXiv paper ]]
PixelProse is a comprehensive dataset of over 16M (million) synthetically generated captions,
leveraging cutting-edge vision-language models (Gemini 1.0 Pro Vision) for detailed and accurate descriptions.
@article{pixelprose24,
title = {{From Pixels to Prose: A Large Dataset of Dense Image Captions}},
author = {Vasu Singla and Kaiyu Yue and Sukriti Paul and Reza Shirkavand and Mayuka Jayawardhana… See the full description on the dataset page: https://huggingface.co/datasets/lodestones/pixelprose.cobe-dmr-pixelized-differential-data
COBE DMR four-year pixelized differential data
The preview bins the 31A source rows by their ordered PIX_PLUS and
PIX_MINU identities; colour records the number of source rows in each 24 by
24 pixel bin.
Regenerate it from the published Parquet data with
python tools/render_hub_preview.py cobe-dmr-pixelized-differential-data datasets/cobe-dmr-pixelized-differential-data/preview.png.
This dataset contains the COBE Differential Microwave Radiometer four-year
Pixelized… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/cobe-dmr-pixelized-differential-data.free-to-use-pixelart
Free-to-use Pixel Art
Dataset Details
This dataset was collected on 25th May, 2024.
It's a small subset of the free-to-use images on PixilArt.
At the time of publication, this dataset was covered by permissive terms that allow commercial use.
Dataset Description
This dataset is unique in that it contains the pixel group size for each collected sample, which might assist in experiments on microconditioning inputs on an adapter to control this value of the unit… See the full description on the dataset page: https://huggingface.co/datasets/bghira/free-to-use-pixelart.draw_pixel_artThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 50,
"total_frames": 26066,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Beegbrain/draw_pixel_art.swe-explore-fix-trajectories
swe-explore-fix trajectories
Agent trajectories on swe-explore-fix. One row per task = its folded agent loop.
dual_cam_test_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
4
],
"names": [
"joint_basis_arm1.pos",
"joint_arm1_arm2.pos",
"joint_arm2_arm3.pos",
"joint_arm3_greifer.pos"
]
}… See the full description on the dataset page: https://huggingface.co/datasets/pixelouie/dual_cam_test_3.javanese-Komodo-pixelgpt
Javanese PixelGPT Dataset
This dataset contains preprocessed Javanese text data for training PixelGPT models.
Dataset Statistics
Language: Javanese (jawa)
Total samples: 401,542
Train samples: 400,726
Test samples: 816
Tokenizers
[TO BE EDITED]
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation of aksara text
<tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/javanese-Komodo-pixelgpt.llm-smartrouter-benchmark
LLM SmartRouter & Agent Highway Latency & Cost Benchmark (v1.4.0)
Empirical performance benchmark dataset comparing direct model endpoints (OpenAI, Anthropic Claude, Google Gemini) against the PixelRouter / BLUN SmartRouter proxy layer and Autonomous Agent Web Highway (https://api.pixeloffice.eu/v1).
v1.4.0 Benchmark Highlights
Anthropic Claude Messages API: Sub-35ms proxy routing for native /v1/messages payloads with 94%+ cost savings.
Machine Web Highway… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/llm-smartrouter-benchmark.Rebuttal-javanese-pixelgpt
Javanese PixelGPT Tokenizer Ablation Dataset
Optimized with Font Size 6 and Dynamic Trimming.
Tokenizer Schema
tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer)
tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT)
tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1)
tok_mt5: Google Multilingual Unigram (google/mt5-small)
balinese-Komodo-pixelgpt
Balinese PixelGPT Dataset
This dataset contains preprocessed Balinese text data for training PixelGPT models.
Dataset Statistics
Language: Balinese (bali)
Total samples: 54,467
Train samples: 54,017
Test samples: 450
Tokenizers
[TO BE EDITED]
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation of aksara text
<tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/balinese-Komodo-pixelgpt.sundanese-pixelgpt
Sundanese PixelGPT Dataset
This dataset contains preprocessed Sundanese text data for training PixelGPT models.
Dataset Statistics
Language: Sundanese (sunda)
Total samples: 294,756
Train samples: 293,933
Test samples: 823
Tokenizers
Grapheme tokenizer: izzako/sunda-llama-tokenizer
LLaMA tokenizer: ernie-research/DualGPT
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation… See the full description on the dataset page: https://huggingface.co/datasets/izzako/sundanese-pixelgpt.dual_cam_test_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
4
],
"names": [
"joint_basis_arm1.pos",
"joint_arm1_arm2.pos",
"joint_arm2_arm3.pos",
"joint_arm3_greifer.pos"
]
}… See the full description on the dataset page: https://huggingface.co/datasets/pixelouie/dual_cam_test_2.pixel-art-bench-v1
Pixel Art Benchmark Dataset (Source)
The Pixel Art Benchmark Dataset is a structured collection of pixel-art outputs generated by large language models (LLMs). Each sample consists of a discrete color palette and a grid-based representation of pixel art, along with generation metadata such as token usage, cost, and model provenance.
Each row in the dataset represents a single generated pixel-art sample.
Encoding Details
Each string in grid represents one row of pixels.… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/pixel-art-bench-v1.balinese-pixelgpt
Balinese PixelGPT Dataset
This dataset contains preprocessed Balinese text data for training PixelGPT models.
Dataset Statistics
Language: Balinese (bali)
Total samples: 54,467
Train samples: 54,017
Test samples: 450
Tokenizers
Grapheme tokenizer: izzako/javanese-llama-tokenizer
LLaMA tokenizer: ernie-research/DualGPT
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel… See the full description on the dataset page: https://huggingface.co/datasets/izzako/balinese-pixelgpt.trojblue-pixelart-videosA copy of https://huggingface.co/datasets/trojblue/test-HunyuanVideo-pixelart-videos to make it compatible for training with https://github.com/a-r-r-o-w/finetrainers (until more dataset formats are supported).
HunyuanVideo-pixelart-videos-sample
trojblue/test-HunyuanVideo-pixelart-images
👋 Heads up—this repository is just a PARTIAL dataset. For the full pixelart-images dataset, make sure to grab both parts:
Images Part
Video Part (this repo)
What's in the Dataset?
This dataset is all about anime-styled pixel art images that have been carefully selected to make your models shine. Here’s what makes these images special:
Rich in detail: Pixelated, yes—but still full of life and not overly simplified.… See the full description on the dataset page: https://huggingface.co/datasets/inlineresearch/HunyuanVideo-pixelart-videos-sample.160-LegoBot-lego_pixelart_stage_1_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 16,
"total_frames": 4747,
"total_tasks": 1,
"total_videos": 32,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:16"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_1_v2.pinchbench-sclean-trajectories
PinchBench Lean Trajectories
Full agent trajectories and scores for small models run through a stripped down,
single model opencode backend on the 116 task PinchBench suite. One row per
(model, task, run).
Harness
The backend is a lean, three agent opencode setup, not the stock product stack:
default agent: the coordinator. Talks to the task, does simple work with
its own tools (read, edit, bash, todowrite), and delegates the rest.
websearch-agent: the only path… See the full description on the dataset page: https://huggingface.co/datasets/pixelxiong/pinchbench-sclean-trajectories.Rebuttal-sundanese-pixelgpt
Sundanese PixelGPT Tokenizer Ablation Dataset
Optimized with Font Size 6 and Dynamic Trimming.
Tokenizer Schema
tok_grapheme: Language-specific Grapheme BPE (izzako/sunda-llama-tokenizer)
tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT)
tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1)
tok_mt5: Google Multilingual Unigram (google/mt5-small)
brand-hallucination-and-ai-citation-benchmark
🛡️ Global Brand Hallucination & LLM Citation Benchmark Dataset
Official open dataset by Pixel Office EU tracking empirical brand hallucination rates, stale pricing quotes, and competitor deflection vectors across leading LLMs (ChatGPT GPT-4o, Claude 3.5 Sonnet, Perplexity AI, Google Gemini 2.5 Flash, and DeepSeek V3).
📊 Dataset Summary
Target Problem: Autonomous AI purchasing agents and AI search engines frequently cite outdated pricing tiers, non-existent… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/brand-hallucination-and-ai-citation-benchmark.lego_pixelart_stage_1_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 28,
"total_frames": 10945,
"total_tasks": 1,
"total_videos": 56,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:28"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/lego_pixelart_stage_1_v3.160-LegoBot-lego_pixelart_stage_2_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower_bimanual",
"total_episodes": 8,
"total_frames": 3989,
"total_tasks": 1,
"total_videos": 32,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:8"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_2_v2.160-LegoBot-lego_pixelart_stage_2_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower_bimanual",
"total_episodes": 31,
"total_frames": 18204,
"total_tasks": 1,
"total_videos": 93,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:31"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_2_v1.160-LegoBot-lego_pixelart_stage_2_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower_bimanual",
"total_episodes": 52,
"total_frames": 28727,
"total_tasks": 1,
"total_videos": 208,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:52"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_2_v3.160-LegoBot-lego_pixelart_stage_1_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 21,
"total_frames": 8064,
"total_tasks": 1,
"total_videos": 42,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_1_v1.160-LegoBot-lego_pixelart_stage_1_point_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 54,
"total_frames": 23460,
"total_tasks": 1,
"total_videos": 108,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:54"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_1_point_v1.calm-coach-81ac5b
calm-coach-81ac5b
Synthetic sensors test data: 33 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Pixel-John655/calm-coach-81ac5b.
