datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dragon
Dataset Card for DRAGON
🧾 ArXiv Preprint
DRAGON is a large-scale Dataset of Realistic imAges Generated by diffusiON models.
The dataset includes a total of 2.5 million training images and 100,000 test images generated using 25 diffusion models, spanning both recent advancements and older, well-established architectures.
Dataset Details
Dataset Description
The remarkable ease of use of diffusion models for image generation has led to a proliferation of… See the full description on the dataset page: https://huggingface.co/datasets/lesc-unifi/dragon.casia-char-1
CASIA Character Sample Dataset
This dataset is adapted from CASIA Online and Offline Chinese Handwriting Databases,
but this only contains character level sample data (from the offline database). The first column is the ground truth label (single character from
GB2312 charset) and the second one is byte sequences of the decoded PNG files from the original .gnt files.
Conditions of Academic Use
Please refer to the official page for more information.
All samples in the… See the full description on the dataset page: https://huggingface.co/datasets/UndefinedCpp/casia-char-1.the-un-laion-templeAll files uploaded. Enjoy!
Dataset Card for The Unlaion Temple
Dataset Details
Dataset Description
Laion-5B is still not public, so we decided to create our own dataset.
The Unlaion Temple is a raw dataset of CommonCrawl images (Estimated to be a total of 2 Billion urls). We haven't verified whether the links in this dataset are functional.
You are responsible for handling the data.
We've made some improvements to the dataset based on user feedback:
All… See the full description on the dataset page: https://huggingface.co/datasets/DSULT-Chiharu/the-un-laion-temple.UnsafeBench
Dataset Card for Dataset Name
[Update]: we added the caption/prompt information (if there is one) in case other researchers need it. It is not used in our study though.
The dataset consists of 10K safe/unsafe images of 11 different types of unsafe content and two sources (real-world VS AI-generated).
Dataset Details
Source
# Safe
# Unsafe
# All
LAION-5B (real-world)
3,228
1,832
5,060
Lexica (AI-generated)
2,870
2,216
5,086
All
6,098
4,048
10,146… See the full description on the dataset page: https://huggingface.co/datasets/yiting/UnsafeBench.ungenerated
Ungenerated
Dataset Details
This dataset contains reconstructed PNG artworks and associated metadata collected from ungenerated.io.
Art Styles
Examples include user-labeled styles such as digital art, traditional art, 3D art, pixel art, photography, and other categories present in the source platform.
Schema
Each example contains:
Column
Type
Description
image
Image
Reconstructed PNG
id
string
Unique artwork ID
title
string
Title given… See the full description on the dataset page: https://huggingface.co/datasets/Thereallo/ungenerated.NFT-70M_transactions
Dataset Card for "NFT-70M_transactions"
Dataset summary
The NFT-70M_transactions dataset is the largest and most up-to-date collection of Non-Fungible Tokens (NFT) transactions between 2021 and 2023 sourced from OpenSea, the leading trading platform in the Web3 ecosystem.
With more than 70M transactions enriched with metadata, this dataset is conceived to support a wide range of tasks, ranging from sequential and transactional data processing/analysis to graph-based… See the full description on the dataset page: https://huggingface.co/datasets/MLNTeam-Unical/NFT-70M_transactions.dermatology-dataset-acne-redness-and-bags-under-the-eyes
Skin Defects Dataset
The dataset contains images of individuals with various skin conditions: acne, skin redness, and bags under the eyes. Each person is represented by 3 images showcasing their specific skin issue. The dataset encompasses diverse demographics, age, ethnicities, and genders.
The dataset is created on the basis of Facial Skin Condition Dataset
Types of defects in the dataset: acne, skin redness & bags under the eyes
Acne photos: display different… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/dermatology-dataset-acne-redness-and-bags-under-the-eyes.NFT-70M_text
Dataset Card for "NFT-70M_text"
Dataset summary
The NFT-70M_text dataset is a companion for our released NFT-70M_transactions dataset,
which is the largest and most up-to-date collection of Non-Fungible Tokens (NFT) transactions between 2021 and 2023 sourced from OpenSea.
As we also reported in the "Data anonymization" section of the dataset card of NFT-70M_transactions,
the textual contents associated with the NFT data were replaced by identifiers to numerical… See the full description on the dataset page: https://huggingface.co/datasets/MLNTeam-Unical/NFT-70M_text.medical-vlm-unlearning-corpus
Incremental Medical VLM Unlearning Subset
Training-ready, leakage-audited configurations are published independently so
completed sources remain usable after interruption. VQA-RAD (CC0), English
SLAKE (CC BY 4.0), and an NIH ChestXray14 subset include pixels. CheXpert is a
source-controlled manifest whose pixels are resolved from the authorized Kaggle
input and are not redistributed. See progress/latest.json and reports/.
This is a research dataset, not a diagnostic product.… See the full description on the dataset page: https://huggingface.co/datasets/Yash908056/medical-vlm-unlearning-corpus.unlabelled_IA_with_snorkel_labels
Historic book pages illustration weak annotations
openm3chest-labels-v2
OpenM3Chest Labels v2
Stratified label subset of OpenM3Chest for LoRA fine-tuning of MedGemma 1.5-4B-IT.
CT scan images (NPY): UngLong/openm3chest-npy-v2
Dataset Summary
Agent
Tasks
Train samples
Test samples
Radiology
chest_abn_54–61, nodule_presence/location/attenuation/margin/size
~350/task
full
Cardiology
CVD_diagnosis, CVD_mortality
1,500/task
3,759/task
Oncology
lung_cancer_risk
2,000
10,308
Training subset: stratified sampling — binary tasks… See the full description on the dataset page: https://huggingface.co/datasets/UngLong/openm3chest-labels-v2.openm3chest-labels
OpenM3Chest Labels (OM3C)
JSON label files and Series UIDs from the OpenM3Chest dataset, prepared for fine-tuning medical vision-language models such as MedGemma.
Raw imaging data (DICOM) can be downloaded from IDC (Imaging Data Commons) using the Series Instance UIDs provided in unique_keys.txt.
Dataset Summary
OpenM3Chest is a medical multimodal multitask dataset for diagnosing chest abnormalities with a focus on lung cancer screening. The original raw data comes… See the full description on the dataset page: https://huggingface.co/datasets/UngLong/openm3chest-labels.mirabest-radio-astronomy-unofficial
MiraBest Radio Astronomy Dataset (Unofficial)
⚠️ IMPORTANT: This is an unofficial repository containing a processed version of the MiraBest dataset formatted for stable diffusion fine-tuning. This repository is not affiliated with the original authors.
Unofficial processing of the MiraBest radio astronomy dataset with original classification labels and natural language captions for diffusion fine-tuning. Original dataset by Porter & Scaife (2023).
Original Dataset
The… See the full description on the dataset page: https://huggingface.co/datasets/kwazzi-jack/mirabest-radio-astronomy-unofficial.pitvqa-unified-vlm
PitVQA Unified VLM Classification Dataset
Surgical workflow classification dataset for training vision-language models on pituitary surgery phase detection, step recognition, and instrument identification.
🔗 GitHub: https://github.com/matheus-rech/pit_project
🤖 Trained Model: mmrech/pitvqa-qwen2vl-unified
📄 Original Dataset: UCL Research Data Repository
Dataset Description
This dataset contains 5,184 surgical frames with classification annotations for surgical… See the full description on the dataset page: https://huggingface.co/datasets/mmrech/pitvqa-unified-vlm.siglip-doc-understanding-classifier
SigLIP Doc Understanding — Unanswerable Question Detection Dataset
A mixed answerable / unanswerable benchmark dataset built from DocVQA and MP-DocVQA, used to
train and evaluate the siglip-doc-understanding-classifier
unanswerable-question detector.
Each row pairs a document image with a question. Half of the questions are the original,
answerable DocVQA/MP-DocVQA questions; the other half are corrupted versions of those same
questions — modified so the document image no longer… See the full description on the dataset page: https://huggingface.co/datasets/giacolees/siglip-doc-understanding-classifier.medical-vlm-unlearning-incremental-subset
Incremental Medical VLM Unlearning Subset
Training-ready, leakage-audited configurations are published independently so
completed sources remain usable after interruption. VQA-RAD (CC0), English
SLAKE (CC BY 4.0), and an NIH ChestXray14 subset include pixels. CheXpert is a
source-controlled manifest whose pixels are resolved from the authorized Kaggle
input and are not redistributed. See progress/latest.json and reports/.
This is a research dataset, not a diagnostic product.… See the full description on the dataset page: https://huggingface.co/datasets/Yash908056/medical-vlm-unlearning-incremental-subset.Burberry.Product.prices.United.States
Burberry web scraped data
About the website
Burberry operates within the luxury fashion industry in the United States, which is a segment of the wider retail industry. The American market is highly competitive and renowned for its considerable consumer spending. E-commerce has become a vital platform for luxury brands like Burberry to extend their customer reach, especially amid changing shopping habits among consumers. The online luxury fashion market in the United… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Burberry.Product.prices.United.States.paleo-hebrew-seals-unambiguous
PaleoHebrew-Seals Real Benchmark (Unambiguous Subset)
This repository hosts the real benchmark part of PaleoHebrew-Seals, a dataset suite for multimodal recognition of Paleo-Hebrew seal inscriptions from photographs.
Why this dataset is needed
Paleo-Hebrew seal inscriptions are difficult for standard OCR systems: the signs are sparse, shallow, frequently worn, and embedded in irregular seal impressions captured under uncontrolled lighting and viewpoint changes.… See the full description on the dataset page: https://huggingface.co/datasets/mr3vial/paleo-hebrew-seals-unambiguous.Loro.Piana.Product.prices.United.States
Loro Piana web scraped data
About the website
Loro Piana operates in the luxury fashion industry in the United States, focusing particularly on high-end, quality fabrics and materials. The brand is popular among wealthy Americans who crave its stylish, high-quality goods ranging from clothing to accessories. In this digital era, Loro Piana has also been proactively involved in the Ecommerce space, maximizing their online presence across several platforms. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Loro.Piana.Product.prices.United.States.Burberry.Product.prices.United.Arab.Emirates
Burberry web scraped data
About the website
The luxury fashion industry in the EMEA region, particularly in the United Arab Emirates, is a robust and rapidly growing market. This growth is primarily fuelled by the affluent consumer base and tourists, who have a keen interest in high-end and prestigious fashion labels. Middle Eastern consumers often look at luxury goods as a status symbol, thus driving up demand for premium brands like Burberry. In recent years, Ecommerce… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Burberry.Product.prices.United.Arab.Emirates.figmirror-unified
Unified FigMirror Dataset
Canonical release with 550 samples.
dataset_augmentation: 500
paper_derivative: 50
paper_derivative verified_pass: 50
Data unit:
One row in data/train.jsonl is one task / one data point.
Asset files under assets/ are supporting files, not separate data points.
Semantic task families:
chart_style_augmentation: input is a reference/source chart; output is an augmented chart.
paper_figure_reproduction: input is a paper figure reference; output is a… See the full description on the dataset page: https://huggingface.co/datasets/zcahjl3/figmirror-unified.Chanel.Product.prices.United.States
Chanel web scraped data
About the website
The global luxury goods industry, specifically the high-end fashion sector, is a competitive marketplace where brands like Chanel thrive. The American market, especially the United States, plays a critical role in this industry, as it is one of the worlds biggest consumers of luxury products. With its affluent consumers propensity for luxury and up-scale products, the US market is a major driver of growth in this sector. The… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Chanel.Product.prices.United.States.Farfetch.Product.prices.United.Kingdom
Farfetch web scraped data
About the website
Farfetch operates in the dynamic and rapidly evolving E-commerce industry in the EMEA, particularly in the United Kingdom. This sector is marked by intense digital transformation with a growing shift towards online shopping. Notably, the fashion and lifestyle segment of e-commerce is witnessing massive growth. The UK E-commerce sector is marked by high internet penetration rates, favourable consumer attitudes, and advances in… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Farfetch.Product.prices.United.Kingdom.Louis.Vuitton.Product.prices.United.Kingdom
Louis Vuitton web scraped data
About the website
Louis Vuitton operates within the luxury fashion industry in the EMEA region, particularly in the United Kingdom. This industry is characterised by high-end products ranging from clothing, accessories to leather goods. It is mainly driven by factors such as brand identity, quality of products, and latest fashion trends. With the rise of digitalisation, an significant portion of sales in this industry shifted to E-commerce… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Louis.Vuitton.Product.prices.United.Kingdom.Celine.Product.prices.United.Kingdom
Celine web scraped data
About the website
The Ecommerce industry in the EMEA region, particularly in the United Kingdom, has shown a significant growth rate in recent years, becoming a pivotal element in the regions economic landscape. The rise of digital technologies and change in consumer behavior has further accelerated this upward trend. In this digital marketplace, Celine, a well-known high-end fashion brand, has maintained its mark. The dataset at-hand encompasses… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Celine.Product.prices.United.Kingdom.Net.a.Porter.Product.prices.United.Kingdom
Net-a-Porter web scraped data
About the website
Net-a-Porter operates in the highly competitive e-commerce industry in EMEA, particularly in the United Kingdom. This vibrant industry is characterized by a dynamic digital environment encompassing various segments, notably fashion, luxury goods, and retail. These segments continue to witness significant growth, driven by advancements in technology and changing consumer behaviors. Net-a-Porter, being a premium online luxury… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Net.a.Porter.Product.prices.United.Kingdom.radiology-ready-v2
Radiology Ready v2
Training-ready dataset for fine-tuning the Radiology agent of an AI Medical Department,
built on MedGemma 1.5-4B-IT.
Each row is a (prompt, response) pair ready for SFT. CT scan images are not stored here —
they live in UngLong/openm3chest-npy-v2
and are fetched at training time via {pids}/{keys}.npy.
Dataset Summary
Rows
500
Screening rows
302 (Pool A + Pool B)
Detail rows
198 (Pool A — nodule present)
Unique CT volumes
477… See the full description on the dataset page: https://huggingface.co/datasets/UngLong/radiology-ready-v2.Rue.La.La.Product.prices.United.States
Rue La La web scraped data
About the website
Rue La La operates in the thriving Ecommerce industry in the United States. This online-driven marketplace sector is significantly fuelled by the continuous innovations in technology, the growing adoption of mobile devices, and the changing consumption patterns of consumers. In particular, Rue La La is recognized for its flash sales model, selling designer apparels, accessories, footwear, and home decor among other things. As… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Rue.La.La.Product.prices.United.States.University_CommonLostItemsunhinged-cast-20
Unhinged Cast — 20 Original Cinema Character Reference Sheets
20 fully-original unhinged movie characters, each as a complete concept-art reference sheet
(front view, side profile, back view, action pose, 4 expression headshots) on a clean neutral
gray studio background — the standard training format for character LoRAs / consistent-character
image+video pipelines.
All sheets generated with Grok Imagine Image 2.0 (grok-imagine-image-2.0, 1:1, 1K) from
detailed fixed prompts.… See the full description on the dataset page: https://huggingface.co/datasets/TheMindExpansionNetwork/unhinged-cast-20.
