CoolFace
Datasetpublic

data-archetype/cc12_imagenet21k_recap_hq_bucketed

cc12_imagenet21k_recap_hq_bucketed Title: cc12_imagenet21k_recap_hq_bucketed Description: This ~18M rows dataset is a re upload of https://huggingface.co/datasets/gmongaras/CC12M_and_Imagenet21K_Recap_Highqual where the images have been pre bucketed into SDXL style aspect ratio buckets for target training at ~512^2 and ~256^2 pixels, and where about 7M rows were recaptioned with either Gemini or Ministral. To avoid re encoding the images they have been left untouched so… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/cc12_imagenet21k_recap_hq_bucketed.

sourceHugging Faceotherupdated 9mo agoView on Hugging Face
0likes58downloads
Dataset Card

cc12imagenet21krecaphqbucketed

  • —Title: cc12imagenet21krecaphqbucketed
  • —Description: This ~18M rows dataset is a re upload of https://huggingface.co/datasets/gmongaras/CC12MandImagenet21KRecapHighqual where the images have been pre bucketed into SDXL style aspect ratio buckets for target training at ~512^2 and ~256^2 pixels, and where about 7M rows were recaptioned with either Gemini or Ministral. To avoid re encoding the images they have been left untouched so cropping and resizing must be done at loading time

Technical details

This repository contains a bucketed-shards export (uncompressed TAR shards).

Format

  • —Format: bucketed_shards_v2
  • —Created: 2026-01-10T15:53:34.486914+00:00
  • —Export ID: export-2026-01-10T15:53:34.486914+00:00
  • —Manifest: manifest.json
  • —Image mode: passthrough_jpeg

Directory layout:

  • —manifest.json (global metadata + per-bucket shard listing)
  • —buckets/<bucket_id>/shard-*.tar

Each TAR shard contains 3 files per sample:

  • —<key>.jpg (JPEG bytes; either re-encoded RGB JPEG or source JPEG passthrough depending on image_mode)
  • —<key>.txt (caption text, UTF-8, newline-terminated)
  • —<key>.json (per-sample metadata: w, h, jpeg, image_mode, caption_variant, caption_selector_index, caption_source_id)

Image preprocessing

Unlike other datasets available in this repo, the images have been left unprocesed and are only pre bucketed into target aspect ratio buckets.

Resize/crop intended usage details at load time:

  • —Cover scale is scale = max(target_w / src_w, target_h / src_h); if scale > 1, the sample is skipped.
  • —After resize, a crop box is chosen deterministically from the sample key (sha256 of image_id).
  • —Corner strategy chooses a corner from allowed_corners where 0=TL, 1=TR, 2=BL, 3=BR (optional small jitter for corner_jitter).

Buckets / resolutions

  • —Buckets follow SDXL-style proto buckets defined at a 1024×1024 base.
  • —Base resolution(s): [512, 256]
  • —In single-res exports, bucket_id is the proto (1024-base) bucket, e.g. p1024x1024.
  • —In multi-res exports, buckets are namespaced by base resolution: r<base>_<proto>, e.g. r512_p1024x1024.
  • —The actual target resolution for each bucket (scaled by the per-bucket base resolution and divisible=32) is stored in:
  • —manifest.json → buckets[<bucket_id>].scaled.w/h (and base_resolution)
  • —each sample’s <key>.json → w/h

Bucket IDs (preview): r256_p1024x1024, r256_p1088x896, r256_p1152x896, r256_p1216x832, r256_p1344x704, r256_p1344x768, r256_p1472x704, r256_p1600x640, r256_p1728x576, r256_p1856x512, r256_p1984x512, r256_p2048x512, r256_p512x1920, r256_p512x2048, r256_p576x1664, r256_p576x1792, r256_p640x1536, r256_p704x1408, r256_p768x1280, r256_p832x1152, … (+42 more)

Bucket distribution:

bucket_idtarget_w×haspectcount
r256_p1152x896288×2241.2862,917,754
r256_p1216x832288×1921.5002,194,582
r512_p1216x832608×4161.4622,103,325
r512_p1152x832576×4161.3851,670,669
r512_p1024x1024512×5121.0001,371,409
r256_p1024x1024256×2561.0001,076,075
r256_p896x1152224×2880.778973,834
r256_p832x1152192×2880.667852,564
r512_p832x1216416×6080.684776,575
r512_p832x1152416×5760.722600,705
r512_p1344x768672×3841.750598,503
r512_p1152x896576×4481.286347,970
r256_p1088x896256×2241.143333,188
r512_p1280x768640×3841.667327,133
r512_p896x1152448×5760.778310,259
r256_p960x1024224×2560.875237,259
r256_p1344x768320×1921.667210,921
r512_p1088x896544×4481.214158,051
r512_p896x1088448×5440.824151,487
r512_p960x1024480×5120.938151,345
r512_p768x1280384×6400.600110,424
r512_p1344x704672×3521.909107,761
r512_p1088x960544×4801.133103,963
r512_p1024x960512×4801.067101,368
r512_p960x1088480×5440.88293,788
r256_p768x1280192×3200.60088,633
r256_p1344x704320×1602.00084,077
r512_p768x1344384×6720.57171,153
r512_p1408x704704×3522.00067,854
r512_p1472x704736×3522.09141,942
r512_p1536x640768×3202.40032,080
r256_p704x1408160×3520.45529,786
r256_p1600x640384×1602.40028,130
r512_p704x1408352×7040.50026,929
r256_p1472x704352×1602.20024,696
r256_p1728x576416×1283.25014,366
r512_p704x1472352×7360.47813,622
r256_p640x1536160×3840.41711,144
r512_p640x1536320×7680.4178,459
r512_p1600x640800×3202.5007,885
r256_p576x1664128×4160.3084,034
r256_p2048x512512×1284.0003,650
r512_p640x1600320×8000.4003,129
r256_p1856x512448×1283.5002,886
r256_p1984x512480×1283.7502,369
r512_p1664x576832×2882.8891,293
r512_p576x1664288×8320.346913
r512_p1792x576896×2883.111732
r256_p512x2048128×5120.250730
r256_p576x1792128×4480.286729
r256_p512x1920128×4800.267565
r512_p1856x512928×2563.625522
r512_p576x1792288×8960.321518
r512_p512x1856256×9280.276476
r512_p1728x576864×2883.000450
r512_p576x1728288×8640.333313
r512_p1920x512960×2563.750136
r512_p512x1920256×9600.267114
r512_p1984x512992×2563.87588
r512_p512x1984256×9920.25874
r512_p512x2048256×10240.25059
r512_p2048x5121024×2564.00056

Caption selection

Available caption variants

selectedvariantimages_with_ok_caption
✓caption_original18,655,051
✓captionministral14b_25124,000,718
✓caption_gemini3,299,328

Missing caption policy: drop

Export summary

  • —images_seen: 18,655,051
  • —images_exported: 18,455,504
  • —skippednocaption: 0
  • —skippedtoosmall: 199,547
  • —decode_errors: 0
  • —encode_errors: 0

Efficient loading

Recommended

Treat this as a webdataset-style collection of tar shards:

  • —Prefer sequential reads of tar files for throughput.
  • —Shuffle at the shard level (and optionally within-shard) for good randomness without expensive random I/O.
  • —Use manifest.json to list buckets and shards.
Python (webdataset)
python
import webdataset as wds

urls = "buckets/*/shard-*.tar"  # glob; adjust if you want a single bucket only
ds = (
    wds.WebDataset(urls)
    .decode("pil")            # decodes .jpg to PIL.Image
    .to_tuple("jpg", "txt", "json")
)
for jpg, caption, meta in ds:
    ...
Python (tarfile, no extra deps)
python
import io, json, tarfile
from pathlib import Path

tar_path = next(Path("buckets").rglob("shard-*.tar"))
with tarfile.open(tar_path, "r") as tf:
    members = tf.getmembers()
    for m in members:
        if not m.name.endswith(".txt"):
            continue
        key = m.name[:-4]
        caption = tf.extractfile(m).read().decode("utf-8").strip()
        meta = json.loads(tf.extractfile(tf.getmember(key + ".json")).read().decode("utf-8"))
        jpg_bytes = tf.extractfile(tf.getmember(key + ".jpg")).read()
        ...