CoolFace
Datasetpublic

sirus/megalith-10m-5.5k-claude-opus-5-recaptioned

Megalith-10M 5.5K — Claude Opus 5 Recaptioned This is a 5,511-image derivative subset of madebyollin/megalith-10m, selected through the megalith10m portion of zlab-princeton/i1-captions. The bytes were retrieved from the drawthingsai/megalith-10m image archive. It is not the complete Megalith-10M collection. Every image has one newly generated, detailed English caption. The recaptioning was performed with Claude Opus 5 via Claude Code on August 2, 2026. The image was the primary… See the full description on the dataset page: https://huggingface.co/datasets/sirus/megalith-10m-5.5k-claude-opus-5-recaptioned.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes42downloads
Dataset Card

Megalith-10M 5.5K — Claude Opus 5 Recaptioned

This is a 5,511-image derivative subset of `madebyollin/megalith-10m`, selected through the `megalith10m` portion of `zlab-princeton/i1-captions`. The bytes were retrieved from the `drawthingsai/megalith-10m` image archive. It is not the complete Megalith-10M collection.

Every image has one newly generated, detailed English caption. The recaptioning was performed with Claude Opus 5 via Claude Code on August 2, 2026. The image was the primary evidence; the pre-existing i1 caption was supplied as a lower-priority hint. The output was required to consolidate the useful visual facts into a natural master caption covering subject, action, spatial relationships, setting, composition, viewpoint, lighting, color, material, texture, photographic style, and legible text when relevant.

Derivative and source chain

RoleSourceExact identification
Original collection and Flickr links`madebyollin/megalith-10m`Megalith-10M source card and records
Image-byte archive`drawthingsai/megalith-10m`Revision `829c008`
Selection and caption hints`zlab-princeton/i1-captions`Revision `bb8c4a4`, megalith10m subset
New annotationClaude Opus 5 through Claude CodeOne synthetic master caption per image

The image bytes were not edited. Each row includes its original Flickr source_url, Megalith source key, exact archive and caption-source revisions, and a SHA-256 hash for byte-level verification.

Fields

  • —image, caption: embedded source image and the new Claude Opus 5 caption.
  • —source_*: original Megalith collection, image archive, i1 selection/caption-hint dataset, exact revisions, source key, and original Flickr URL.
  • —image_sha256, width, height: integrity and dimensions.
  • —caption_model, caption_interface, captioned_at, caption_policy, is_derivative: caption provenance.

License and attribution

The Megalith-10M source card describes the source links as Flickr Commons/no-known-restrictions, U.S. government works, CC0, or Public Domain Mark material, while also warning users to perform their own copyright analysis. This subset retains the original Flickr URL on every row, but the local extraction did not retain a more specific public-domain category per image. For that reason this repository uses license: other rather than asserting one blanket license. Check the source record and Flickr URL before use; captions do not change the status of the images.

Quality, intended use, and limitations

The captions average 1,122.8 characters (median 1,124; range 678–1,560). The dataset is intended for image-captioning research, text-to-image training experiments, and inspection of richer synthetic descriptions.

Captions are model-generated and may contain hallucinated details, missed small objects, inaccurate named entities, or imperfect text recognition. The subset was selected for another training project and is not statistically representative of all Megalith-10M images. The upstream source card estimates that a small fraction of the full collection may contain copyright-constrained, edited, non-photographic, or non-wholesome material; this subset has not received a new exhaustive legal or content audit.

Loading

python
from datasets import load_dataset

dataset = load_dataset(
    "sirus/megalith-10m-5.5k-claude-opus-5-recaptioned",
    split="train",
)

Citation

Please cite the i1 work associated with the selection and caption-hint dataset:

bibtex
@article{zeng2026i1,
  title={i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models},
  author={Zeng, Boya and Luo, Tianze and Pu, Shu and Shen, Jucheng and Lu, Taiming and Sarch, Gabriel and Liu, Zhuang},
  journal={arXiv preprint arXiv:2606.11289},
  year={2026}
}

Also retain the original Megalith-10M source link and any attribution required by the corresponding Flickr record.