datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
better_daily_dialogpixelprose
From Pixels to Prose: A Large Dataset of Dense Image Captions
[ arXiv paper ] | [ 🌮 image tars ]
PixelProse is a comprehensive dataset of over 16M (million) synthetically generated captions,
leveraging cutting-edge vision-language models (Gemini 1.0 Pro Vision) for detailed and accurate descriptions.
1. Details
Total number of image-caption pairs: 16,896,214 (16.9M)
6,538,898 (6.5M) pairs in the split of CommonPool
9,066,455 (9.1M) pairs in the split of CC12M
1,290… See the full description on the dataset page: https://huggingface.co/datasets/tomg-group-umd/pixelprose.pixelprose
From Pixels to Prose: A Large Dataset of Dense Image Captions
[[ arXiv paper ]]
PixelProse is a comprehensive dataset of over 16M (million) synthetically generated captions,
leveraging cutting-edge vision-language models (Gemini 1.0 Pro Vision) for detailed and accurate descriptions.
@article{pixelprose24,
title = {{From Pixels to Prose: A Large Dataset of Dense Image Captions}},
author = {Vasu Singla and Kaiyu Yue and Sukriti Paul and Reza Shirkavand and Mayuka Jayawardhana… See the full description on the dataset page: https://huggingface.co/datasets/lodestones/pixelprose.cobe-dmr-pixelized-differential-data
COBE DMR four-year pixelized differential data
The preview bins the 31A source rows by their ordered PIX_PLUS and
PIX_MINU identities; colour records the number of source rows in each 24 by
24 pixel bin.
Regenerate it from the published Parquet data with
python tools/render_hub_preview.py cobe-dmr-pixelized-differential-data datasets/cobe-dmr-pixelized-differential-data/preview.png.
This dataset contains the COBE Differential Microwave Radiometer four-year
Pixelized… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/cobe-dmr-pixelized-differential-data.free-to-use-pixelart
Free-to-use Pixel Art
Dataset Details
This dataset was collected on 25th May, 2024.
It's a small subset of the free-to-use images on PixilArt.
At the time of publication, this dataset was covered by permissive terms that allow commercial use.
Dataset Description
This dataset is unique in that it contains the pixel group size for each collected sample, which might assist in experiments on microconditioning inputs on an adapter to control this value of the unit… See the full description on the dataset page: https://huggingface.co/datasets/bghira/free-to-use-pixelart.draw_pixel_artThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 50,
"total_frames": 26066,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Beegbrain/draw_pixel_art.dual_cam_test_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
4
],
"names": [
"joint_basis_arm1.pos",
"joint_arm1_arm2.pos",
"joint_arm2_arm3.pos",
"joint_arm3_greifer.pos"
]
}… See the full description on the dataset page: https://huggingface.co/datasets/pixelouie/dual_cam_test_3.swe-explore-fix-trajectories
swe-explore-fix trajectories
Agent trajectories on swe-explore-fix. One row per task = its folded agent loop.
javanese-Komodo-pixelgpt
Javanese PixelGPT Dataset
This dataset contains preprocessed Javanese text data for training PixelGPT models.
Dataset Statistics
Language: Javanese (jawa)
Total samples: 401,542
Train samples: 400,726
Test samples: 816
Tokenizers
[TO BE EDITED]
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation of aksara text
<tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/javanese-Komodo-pixelgpt.Rebuttal-javanese-pixelgpt
Javanese PixelGPT Tokenizer Ablation Dataset
Optimized with Font Size 6 and Dynamic Trimming.
Tokenizer Schema
tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer)
tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT)
tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1)
tok_mt5: Google Multilingual Unigram (google/mt5-small)
pixel-art-bench-v1
Pixel Art Benchmark Dataset (Source)
The Pixel Art Benchmark Dataset is a structured collection of pixel-art outputs generated by large language models (LLMs). Each sample consists of a discrete color palette and a grid-based representation of pixel art, along with generation metadata such as token usage, cost, and model provenance.
Each row in the dataset represents a single generated pixel-art sample.
Encoding Details
Each string in grid represents one row of pixels.… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/pixel-art-bench-v1.dual_cam_test_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
4
],
"names": [
"joint_basis_arm1.pos",
"joint_arm1_arm2.pos",
"joint_arm2_arm3.pos",
"joint_arm3_greifer.pos"
]
}… See the full description on the dataset page: https://huggingface.co/datasets/pixelouie/dual_cam_test_2.balinese-Komodo-pixelgpt
Balinese PixelGPT Dataset
This dataset contains preprocessed Balinese text data for training PixelGPT models.
Dataset Statistics
Language: Balinese (bali)
Total samples: 54,467
Train samples: 54,017
Test samples: 450
Tokenizers
[TO BE EDITED]
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation of aksara text
<tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/balinese-Komodo-pixelgpt.sundanese-pixelgpt
Sundanese PixelGPT Dataset
This dataset contains preprocessed Sundanese text data for training PixelGPT models.
Dataset Statistics
Language: Sundanese (sunda)
Total samples: 294,756
Train samples: 293,933
Test samples: 823
Tokenizers
Grapheme tokenizer: izzako/sunda-llama-tokenizer
LLaMA tokenizer: ernie-research/DualGPT
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation… See the full description on the dataset page: https://huggingface.co/datasets/izzako/sundanese-pixelgpt.160-LegoBot-lego_pixelart_stage_1_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 16,
"total_frames": 4747,
"total_tasks": 1,
"total_videos": 32,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:16"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_1_v2.lego_pixelart_stage_1_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 28,
"total_frames": 10945,
"total_tasks": 1,
"total_videos": 56,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:28"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/lego_pixelart_stage_1_v3.balinese-pixelgpt
Balinese PixelGPT Dataset
This dataset contains preprocessed Balinese text data for training PixelGPT models.
Dataset Statistics
Language: Balinese (bali)
Total samples: 54,467
Train samples: 54,017
Test samples: 450
Tokenizers
Grapheme tokenizer: izzako/javanese-llama-tokenizer
LLaMA tokenizer: ernie-research/DualGPT
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel… See the full description on the dataset page: https://huggingface.co/datasets/izzako/balinese-pixelgpt.160-LegoBot-lego_pixelart_stage_2_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower_bimanual",
"total_episodes": 8,
"total_frames": 3989,
"total_tasks": 1,
"total_videos": 32,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:8"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_2_v2.160-LegoBot-lego_pixelart_stage_2_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower_bimanual",
"total_episodes": 31,
"total_frames": 18204,
"total_tasks": 1,
"total_videos": 93,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:31"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_2_v1.160-LegoBot-lego_pixelart_stage_1_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 21,
"total_frames": 8064,
"total_tasks": 1,
"total_videos": 42,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_1_v1.160-LegoBot-lego_pixelart_stage_2_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower_bimanual",
"total_episodes": 52,
"total_frames": 28727,
"total_tasks": 1,
"total_videos": 208,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:52"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_2_v3.160-LegoBot-lego_pixelart_stage_1_point_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 54,
"total_frames": 23460,
"total_tasks": 1,
"total_videos": 108,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:54"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/LeRobot-worldwide-hackathon/160-LegoBot-lego_pixelart_stage_1_point_v1.Rebuttal-sundanese-pixelgpt
Sundanese PixelGPT Tokenizer Ablation Dataset
Optimized with Font Size 6 and Dynamic Trimming.
Tokenizer Schema
tok_grapheme: Language-specific Grapheme BPE (izzako/sunda-llama-tokenizer)
tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT)
tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1)
tok_mt5: Google Multilingual Unigram (google/mt5-small)
pinchbench-sclean-trajectories
PinchBench Lean Trajectories
Full agent trajectories and scores for small models run through a stripped down,
single model opencode backend on the 116 task PinchBench suite. One row per
(model, task, run).
Harness
The backend is a lean, three agent opencode setup, not the stock product stack:
default agent: the coordinator. Talks to the task, does simple work with
its own tools (read, edit, bash, todowrite), and delegates the rest.
websearch-agent: the only path… See the full description on the dataset page: https://huggingface.co/datasets/pixelxiong/pinchbench-sclean-trajectories.pixelvision-441k-recap-qwen2.5-vl-3B
Caption Dataset
This dataset contains 441,053 captioned items exported from CaptionFlow.
Dataset Structure
Data Fields
job_id: object
dataset: object
shard: object
chunk_id: object
item_key: object
item_index: int64
url: object
caption_count: int64
contributor_id: object
timestamp: datetime64[ns]
image_width: int64
image_height: int64
image_format: object
processing_time_ms: float64
metadata: object
captions: List of captions/outputs… See the full description on the dataset page: https://huggingface.co/datasets/RareConcepts/pixelvision-441k-recap-qwen2.5-vl-3B.dual_cam_test_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"shape": [
7
],
"names": [
"joint_basis_arm1.pos",
"joint_arm1_arm2.pos",
"joint_arm2_arm3.pos",
"joint_arm3_greifer.pos",
"joint_greifer_finger1.pos"… See the full description on the dataset page: https://huggingface.co/datasets/pixelouie/dual_cam_test_4.pixelprose-flowersvizdoom-inference-pixel-datasetAs for the details of this dataset, please refer to https://github.com/Masao-Taketani/GameNGen.
nema_pick_place_test_20260605_164708This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"joint_basis_arm1.pos",
"joint_arm1_arm2.pos",
"joint_arm2_arm3.pos",
"joint_arm3_greifer.pos"
],
"shape": [
4
]
}… See the full description on the dataset page: https://huggingface.co/datasets/pixelouie/nema_pick_place_test_20260605_164708.nema_test_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
4
],
"names": [
"joint_basis_arm1.pos",
"joint_arm1_arm2.pos",
"joint_arm2_arm3.pos",
"joint_arm3_greifer.pos"
]
}… See the full description on the dataset page: https://huggingface.co/datasets/pixelouie/nema_test_v3.
