data-archetype/LAION_Aesthetics_1024_bucketed_1024
LAION Aesthetics 1024 Bucketed 1024 Captioned Bucketed-shards export of images from limingcv/LAION_Aesthetics_1024, with a small high-resolution addon from limingcv/LAION_Aesthetics_512. Images: 382,412 Shards: 398 uncompressed WebDataset-style TAR files Format: bucketed_shards_v1 Base resolutions: [1024, 512] Manifest: manifest.json Images are organized under buckets/<bucket_id>/<shard>.tar. Each sample contains: <key>.jpg <key>.txt <key>.json The .txt files contain… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/LAION_Aesthetics_1024_bucketed_1024.
LAION Aesthetics 1024 Bucketed 1024 Captioned
Bucketed-shards export of images from limingcv/LAION_Aesthetics_1024, with a small high-resolution addon from limingcv/LAION_Aesthetics_512.
- Images:
382,412 - Shards:
398uncompressed WebDataset-style TAR files - Format:
bucketed_shards_v1 - Base resolutions:
[1024, 512] - Manifest:
manifest.json
Images are organized under buckets/<bucket_id>/<shard>.tar. Each sample contains:
<key>.jpg<key>.txt<key>.json
The .txt files contain model-generated captions. The original LAION scrape text is not used as the training caption.
Caption priority:
google/gemini-2.5-flash-lite:32,941google/gemini-2.0-flash-001:348,861mistralai/mistral-medium-3.1:147openai/gpt-5-mini:463
Most images were resized/cropped into SDXL-style 1024 buckets without upsampling. The 463-image addon is stored as passthrough JPEG bytes and assigned to matching 1024 buckets; loaders should use the bucket metadata in each sample JSON/manifest for target dimensions.
See manifest.json for bucket counts, shard checksums, caption source IDs, and prompt hashes. Prompt text is intentionally not included.
