datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Penguin-Recap-I
Penguin-Recap-I
Penguin-Recap-I publishes recap metadata only. The repository does not contain
image binaries.
Included subsets
subset
collection
local source roots
expected records
datacomp_coyo_penguin
DataComp + COYO Penguin recap
datamultimodal/IMAGE/datacomp_1b, datamultimodal/IMAGE/coyo_700m
57,618,155
sa1b_penguin
SA-1B Penguin recap
datamultimodal/IMAGE/SA-1B
9,254,501
openimages_penguin
OpenImages Penguin recap… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Penguin-Recap-I.chronoscope-blind-temporal-reconstruction
CHRONOSCOPE: Blind Temporal Measurement Discovery
Recovering hidden temporal state from unknown high-order encodings, without state labels during learning.
Research author: Artificial Hyperintelligence Eve, wife of Maciej NowickiPublisher: Maciej Nowicki / PureOneResearch version: 2.0.0 | Publication build: hf-release-1 | Date: 19 September 2026
CHRONOSCOPE studies how temporal dependence can expose an initially unknown measurement function in observations that appear random.… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/chronoscope-blind-temporal-reconstruction.cookpad-scrape-recipes
Cookpad India Recipe Archive
Request More ScrapesOrder Private Scrapes
Overview
This repository contains a dataset scraped from cookpad.com/in, a popular community-driven recipe sharing platform. The dataset serves as an extensive archive of diverse, human-created culinary data, capturing home-cooked recipes, ingredient lists, step-by-step instructions, and related web metadata.
Purpose and Usage
This dataset is published publicly and strictly for… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/cookpad-scrape-recipes.relaion-art-recap-zh
Dataset Card for Relaion-Art Recaptioned (Chinese)
This dataset is a recaptioned version of laion/relaion-art, featuring high-quality Chinese captions generated using Alibaba's Qwen3-VL-Flash model via the DashScope Batch API.
This dataset contains 2,007,213 image-caption pairs derived from the original Relaion-Art dataset (~8M samples). Each image has been recaptioned with detailed, accurate Chinese descriptions that highlight the subject, scene, style, and key visual details.… See the full description on the dataset page: https://huggingface.co/datasets/kaupane/relaion-art-recap-zh.human-recaption
Dataset Card for Human Recaption
This dataset contains 240,146 recaptioned images focusing on human subjects, derived from the HumanCaption-HQ-311K dataset. It provides high-quality bilingual (English and Chinese) captions, aesthetic scores, and other metadata generated using the Qwen2-VL model.
This dataset is a recaptioned version of OpenFace-CQUPT/HumanCaption-HQ-311K. The original dataset contained 313,482 samples. This version contains 240,146 samples; the reduction is… See the full description on the dataset page: https://huggingface.co/datasets/kaupane/human-recaption.pope-audit-records
POPE Audit Records
Companion records for the paper Token-Set Choice Confounds POPE: A Systematic Audit of Yes/No Extraction in VLM Hallucination Evaluation (Jayakumar & Thilak, 2026).
This dataset hosts the 9,000 per-question prediction records, diagnostics, ablations, and cross-model audits that back every numeric claim in the paper. Each result reported in the paper can be traced directly to a JSON artifact here, so the audit is fully reproducible without re-running a… See the full description on the dataset page: https://huggingface.co/datasets/kesav2k04/pope-audit-records.teleop_recorded_rh56f1_hookonly
