datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
franka_real-tennis_bucket_uprightThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 100,
"total_frames": 31442,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/qingpowuwu/franka_real-tennis_bucket_upright.photo-concept-bucket
Photo Concept Bucket
The purpose of this dataset was to distribute a high quality, free-to-use dataset containing samples that require no attribution and have an open license.
All of the images were captioned in a cluster containing:
38x 3090 24G
6x 4090 24G
8x A5000 24G
2x A100 80G
A couple volunteers running a 3090 or 4090.
The model was running in fp8 precision using 🤗Transformers and 🤗Accelerate for easy multi-GPU captioning.
The captioning was spread across 10 different… See the full description on the dataset page: https://huggingface.co/datasets/bghira/photo-concept-bucket.img-bucketimg-bucket
Work In Progress Demo Images
PCA Grid
Skin Luminance x Chroma
cc12_imagenet21k_recap_hq_bucketed
cc12_imagenet21k_recap_hq_bucketed
Title: cc12_imagenet21k_recap_hq_bucketed
Description: This ~18M rows dataset is a re upload of https://huggingface.co/datasets/gmongaras/CC12M_and_Imagenet21K_Recap_Highqual where the images have
been pre bucketed into SDXL style aspect ratio buckets for target training at ~512^2 and ~256^2 pixels, and where about 7M rows were recaptioned with either Gemini or Ministral.
To avoid re encoding the images they have been left untouched so cropping… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/cc12_imagenet21k_recap_hq_bucketed.product-database-bucketimagenet_22k_512_bucketable
ImageNet-22k 512-Bucketable Captioned Subset
This dataset is a pre-bucketed, captioned subset of timm/imagenet-22k-wds.
It is intended for text-to-image training and similar workflows that want images already grouped into aspect-ratio buckets near a 512-base training resolution. Images were kept only if they could fit one of the target buckets without upsampling after deterministic resize and crop.
Summary
Source: timm/imagenet-22k-wds (fall11 ImageNet-22k WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/imagenet_22k_512_bucketable.photo-concept-bucket-wdsLAION_Aesthetics_1024_bucketed_512
LAION Aesthetics 1024 Bucketed 512 Captioned
This is a captioned bucketed-shards export of images from limingcv/LAION_Aesthetics_1024.
Images were filtered and resized/cropped into SDXL-style aspect-ratio buckets at a 512 base resolution, without upsampling. The export contains 382,144 images across 397 uncompressed WebDataset-style tar shards.
The .txt files now contain model-generated captions, not the original LAION web-scrape alt text or surrounding page text. Captions were… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/LAION_Aesthetics_1024_bucketed_512.pentester_bucketbg_photo_concepts_bucketed_512
bg_photo_concepts_bucketed_512
Title: bg_photo_concepts_bucketed_512
Description: A recaptioned, self contained, bucketed and ready to train with version of https://huggingface.co/datasets/bghira/photo-concept-bucket, exported at 512^2 ish resolution buckets.
I lost the tracking data of which version of Gemini this was captioned with, likely 2.0 flash or 2.5 flash. The captions are on the long and datailed side and sometimes slightly redundant, but overall high quality.… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/bg_photo_concepts_bucketed_512.recipe-synthetic-images-10k-bucketclassical_paintings_bucketed_1024
Classical Paintings Captioned
A curated dataset of 7,131 classical paintings by 42 artists spanning the Baroque period through the 19th century, each with a descriptive plain-language caption (100--150 words). Intended for fine-tuning text-to-image models.
Artists (42)
Aelbert Cuyp, Albert Bierstadt, Anders Zorn, Anthony van Dyck, Artemisia Gentileschi, Caravaggio, Diego Velazquez, Frans Hals, Frederic Edwin Church, Georges de La Tour, Gerard ter Borch, Gerrit Dou… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/classical_paintings_bucketed_1024.anyline-photo-concept-bucketnuextract3-nls-bucket-smoke
NuExtract3 on NationalLibraryOfScotland/nls-index-cards-object-detection
This dataset contains outputs from NationalLibraryOfScotland/nls-index-cards-object-detection processed with NuExtract3, a 4B vision-language model for document understanding.
Processing Details
Source Dataset: NationalLibraryOfScotland/nls-index-cards-object-detection
Model: numind/NuExtract3
Mode: structured-extractionNumber of Samples: 5
Processing Time: 4.4 min
Processing Date: 2026-05-20… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/nuextract3-nls-bucket-smoke.image-bucket-d1a2108f51f6bucketimnet1k_bucket_pailimage-bucket-78b10c30f996ice-cal-bucketgemma-4-E2B-it-bucketimage-bucket-62405f2679ddimage-bucket-89cb76ffb30bimage-bucket-0c50c6c003a9image-bucket-c78f60c3cda1image-bucket-917d14335de7image-bucket-328399026bd3image-bucket-7bc9390bc44bimage-bucket-cd0adf62f914image-bucket-d86c121683d8
