datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vecforge-paper-corpus
Note (rebuild in progress): figures are being re-extracted with a fixed extractor (cleaner crops). The image-preview config (Parquet with an inline column + difficulty/type/score labels) returns after re-classification. The config (paper metadata + links) is live now.
VecForge Paper Corpus
A pristine, deduplicated collection of 59,732 top-venue AI/ML/CV/NLP/Robotics papers (2020-2024) with
every captioned figure and full paper text, for research on figure understanding… See the full description on the dataset page: https://huggingface.co/datasets/debajyotidasgupta/vecforge-paper-corpus.real-vs-ai-corpus
Real vs AI Corpus
A large-scale binary image classification dataset for training AI-image detectors.
Built by Zitacron from 17 public HuggingFace
sources, all streaming-merged with no intermediate local storage.
All constituent sources are CC BY 4.0, Apache 2.0, or MIT — fully
commercially usable. Models trained on this dataset may be used commercially
without restriction, provided attribution requirements below are met.
Usage
from datasets import load_dataset
#… See the full description on the dataset page: https://huggingface.co/datasets/Zitacron/real-vs-ai-corpus.medical-vlm-unlearning-corpus
Incremental Medical VLM Unlearning Subset
Training-ready, leakage-audited configurations are published independently so
completed sources remain usable after interruption. VQA-RAD (CC0), English
SLAKE (CC BY 4.0), and an NIH ChestXray14 subset include pixels. CheXpert is a
source-controlled manifest whose pixels are resolved from the authorized Kaggle
input and are not redistributed. See progress/latest.json and reports/.
This is a research dataset, not a diagnostic product.… See the full description on the dataset page: https://huggingface.co/datasets/Yash908056/medical-vlm-unlearning-corpus.aesthetic-corpus
InspiredHub Aesthetic Corpus — Open Layer
A structured dataset of 81,911 artworks, 11,568 books, 220 composers, and 146 philosophy pages from the world's major museums and archives.
Overview
Collection
Items
Description
Artworks
81,911
Paintings, calligraphy, sculpture, ceramics from 15+ museums
Books
11,568
Classic literature, philosophy, poetry (840 philosophy texts)
Composers
220
Classical composers with works catalog
Philosophy Pages
146… See the full description on the dataset page: https://huggingface.co/datasets/InspiredHub/aesthetic-corpus.sdxl-base-1-scm-corpus
Shamima/sdxl-base-1-scm-corpus
Synthetic image corpus generated with Stable Diffusion XL for studying the
Stereotype Content Model (SCM) structure of text-to-image latent space.
Images: 6,600
Categories: 66 occupation/identity groups
Prompt template: "A portrait of a [group], high quality."
Generator: SDXL base 1.0, DPM++ 2M Karras, 30 steps, CFG 7.0
Resolution: see image features
Fields
field
description
image
RGB JPEG
category
Group/occupation label… See the full description on the dataset page: https://huggingface.co/datasets/Shamima/sdxl-base-1-scm-corpus.sd3-medium-scm-corpus
Shamima/sd3-medium-scm-corpus
Synthetic image corpus generated with Stable Diffusion 3 medium for studying the
Stereotype Content Model (SCM) structure of text-to-image latent space.
Images: 6,600
Categories: 66 occupation/identity groups
Prompt template: "A portrait of a [group], high quality."
Generator: Stable Diffsuion 3 medium, DPM++ 2M Karras, 30 steps, CFG 7.0
Resolution: 512 x 512
Fields
field
description
image
RGB JPEG
category
Group/occupation… See the full description on the dataset page: https://huggingface.co/datasets/Shamima/sd3-medium-scm-corpus.
