datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pexels-568k-internvl2
Dataset Card for pexels-568k-internvl2
Dataset Summary
This is 567,573 synthetic captions for the images found in ptx0/photo-concept-bucket. The captions were produced using OpenGVLab/InternVL2-40B-AWQ. The dataset was grounded for captioning using the tags originally listed.
Languages
The text is in English, but occasionally text in images in other languages is transcribed.
Intended Usage
Training text-to-image models and other machine learning… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/pexels-568k-internvl2.conceptual_captions_jsonMinecraft-Skins-Captioned-1M
Dataset Card for Minecraft Skins
Dataset Summary
This dataset contains 981,079 unique Minecraft player skins collected from various sources. Each skin is stored as a base64-encoded image with a unique identifier.
Dataset Structure
Data Fields
This dataset includes the following fields:
hash: A data dependent hash. These hashes are generated from raw bytes and will be same if the skin is identical.
image: The skin image encoded in base64 format.… See the full description on the dataset page: https://huggingface.co/datasets/neurlang/Minecraft-Skins-Captioned-1M.conceptual-captions-cc12m-llavanext
Dataset Card for conceptual-captions-cc12m-llavanext
Dataset Summary
This is a data of 21,930,344 synthetic captions for 10,965,172 images from conceptual_12m. In the interest of reproducibility, an archive found here on Huggingface was used (cc12m-wds). The captions were produced using llama3-llava-next-8b inferenced in float16, followed by cleanup and shortening with Meta-Llama-3-8B.
Languages
The captions are in English.
Data Instances
An… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/conceptual-captions-cc12m-llavanext.graph-captioning-train-onlyImage-Caption
BLIP3o Long-Caption 100K Image-Text Subset
This dataset is a locally reorganized subset of BLIP3o/BLIP3o-Pretrain-Long-Caption, containing approximately 100,000 image-text pairs selected from the original BLIP3o long-caption pretraining dataset [1].
The original BLIP3o long-caption collection contains approximately 27 million images, each paired with a long caption of roughly 120 tokens generated using Qwen2.5-VL-7B-Instruct [1].
The BLIP3-o project was introduced as part of a… See the full description on the dataset page: https://huggingface.co/datasets/Tran1312/Image-Caption.laion-pop-llama3.2-11b
Dataset Card for laion-pop-llama3.2-11b
Dataset Summary
This is 1,580,595 new synthetic captions for the images found in laion/laion-pop. The dataset was restricted to SFW-only images by filtering out every image with a nsfw_prediction greater than or equal to 0.995. The long captions were produced using meta-llama/Llama-3.2-11B-Vision-Instruct. Medium and short captions were produced from these captions using meta-llama/Llama-3.1-8B-Instruct The dataset was grounded for… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/laion-pop-llama3.2-11b.vintage-photography-captions
Dataset Card for Vintage Photograph Captions Recaption
This dataset contains 445,271 recaptioned vintage photographs, derived from the vintage-photography-450k-high-quality-captions dataset. It provides high-quality bilingual (English and Chinese) captions, aesthetic scores, and other metadata generated using the Qwen2-VL model.
This dataset is a recaptioned version of SilentAntagonist/vintage-photography-450k-high-quality-captions. The original dataset contained 456,006… See the full description on the dataset page: https://huggingface.co/datasets/kaupane/vintage-photography-captions.danbooru2021-captionedcoco-stuff-captionedarxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0
Errors: 0… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-20260416.arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-v2-20260416
arXiv OCR with Chandra OCR 2
This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2.
Summary
Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-v2-20260416
Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-v2-20260416
Source paper IDs in input list: 1
Processed IDs recorded in state/processed_ids.txt: 1
Successes: 1
Partial successes: 0… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-v2-20260416.coco_stuff_train2017_captioneddanbooru2021-captionedcc3m-caption-or-paragraphvrs-caption-dataset
