datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anything-v3.0-glazed
Dataset Card for Anything v3.0 Glazed Samples
Dataset Description
Dataset Summary
This dataset contains image samples originally generated by Linaqruf/anything-v3.0
and subsequently processed by Glaze tool.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/anything-v3.0-glazed.stable-diffusion-v1-5-glazed
Dataset Card for Stable Diffusion v1.5 Glazed Samples
Dataset Description
Dataset Summary
This dataset contains image samples originally generated by runwayml/stable-diffusion-v1-5
and subsequently processed by Glaze tool.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/stable-diffusion-v1-5-glazed.Glaucoma_Dataset
Glaucoma Dataset
Dataset Summary
The Glaucoma Dataset is a comprehensive collection of retinal fundus images designed for the automated detection and classification of glaucoma. Containing between 10,000 and 100,000 high-quality images, this dataset aims to support the development and evaluation of machine learning and deep learning models in the field of ophthalmic medical imaging.
The dataset is organized using the standard imagefolder format, making it highly… See the full description on the dataset page: https://huggingface.co/datasets/Nj-1111/Glaucoma_Dataset.glaucoma-expert-cot-final
Glaucoma Expert Chain-of-Thought
Ophthalmologist six-step reasoning reports for fundus photographs, each paired with
a binary glaucoma label. 1,074 cases from LAG and Papila.
Files
file
rows
glaucoma / not
train.jsonl
823
304 / 519
val.jsonl
92
46 / 46
test.jsonl
159
79 / 80
images/
1,074
<source>_<id>.jpg
Record schema
{
"id": "1689",
"source": "LAG",
"image": "LAG_1689.jpg",
"split": "train",
"final_diagnosis_GT":… See the full description on the dataset page: https://huggingface.co/datasets/yuzhench/glaucoma-expert-cot-final.glaucoma-expert-cot-raw-1077
Glaucoma Expert Chain-of-Thought
Ophthalmologist six-step reasoning reports for fundus photographs, each paired with
a binary glaucoma label. 1,074 cases from LAG and Papila.
Files
file
rows
split
expert_cot_trainval.jsonl
915
train (823) + val (92)
expert_cot_test.jsonl
159
test
images/
1,074
<source>_<id>.jpg
Record schema
{
"id": "1689",
"source": "LAG",
"image": "LAG_1689.jpg",
"split": "train"… See the full description on the dataset page: https://huggingface.co/datasets/yuzhench/glaucoma-expert-cot-raw-1077.corruption-glass_blur
Corruption Dataset: Glass_Blur
Dataset Description
This dataset contains corrupted versions of ImageNet-1K images using glass_blur corruption. It is part of the ImageNet-C benchmark for evaluating model robustness to common image corruptions.
Dataset Structure
Train: 1,281,167 corrupted images
Validation: 50,000 corrupted images
Classes: 1000 ImageNet-1K classes
Format: Arrow (Hugging Face Datasets)
Corruption Type: Glass_Blur
Applies glass blur… See the full description on the dataset page: https://huggingface.co/datasets/MarMaster/corruption-glass_blur.GLAMI-Entity-Matching-Dataset
GLAMI Duplication Detection
Product-duplicate detection over GLAMI e-commerce listings: ~1.3M product
images plus multilingual titles, descriptions and attributes, with labelled
groups of items that do or do not refer to the same physical product.
Released under the Apache License 2.0 — see LICENSE.
TODO: describe how the labels were produced.
Structure
Config
Files
Contents
images
images/shard-*.parquet
itemId → image bytes, one row per product image… See the full description on the dataset page: https://huggingface.co/datasets/zidcenek/GLAMI-Entity-Matching-Dataset.social-commerce-screenshot-claim-glance
Social Commerce Screenshot Claim Glance
This is a 100-image public preview subset for research on catalog-grounded visual retrieval and exact-claim authorization in social-commerce screenshots.
The dataset name for Hugging Face should be:
social-commerce-screenshot-claim-glance
What It Contains
The preview contains synthetic/redacted app-style screenshots that resemble noisy buyer or social-commerce inputs. The examples include platform-like UI framing, cropped… See the full description on the dataset page: https://huggingface.co/datasets/Sonjoy/social-commerce-screenshot-claim-glance.SynGallery-abl4-tex-light-glass-frame
SynGallery-abl4-tex-light-glass-frame: + frame variety
Rung 4 of the SynGallery instance-level artwork-recognition ablation ladder. 4,898 MET paintings × 5 camera viewpoints = 24,490 synthetic RGB images at 512×512, paired with their source photos and museum metadata.
In this rung, the scene varies textures, lighting, glass and frame molding variant + color/roughness/metallic, while freezing camera pose (the only frozen factor). Same schema, source images and index↔painting… See the full description on the dataset page: https://huggingface.co/datasets/patryk-bartkowiak/SynGallery-abl4-tex-light-glass-frame.Glacier-Dataset
Polar Glacier Bitemporal Remote Sensing Dataset
Dataset Overview
This dataset contains bitemporal remote sensing images from two representative polar regions:
Southeastern Coast of Greenland
(Latitude 64°–66°N, Longitude 51°–56°W):Dominated by glaciers and icefields, this area features exposed bedrock mountains and narrow coastal vegetation zones. It is a key region for studying glacier dynamics, with typical crevasse systems on the glacier surface and… See the full description on the dataset page: https://huggingface.co/datasets/cuibinge/Glacier-Dataset.SynGallery-abl3-tex-light-glass
SynGallery-abl3-tex-light-glass: + glass
Rung 3 of the SynGallery instance-level artwork-recognition ablation ladder. 4,898 MET paintings × 5 camera viewpoints = 24,490 synthetic RGB images at 512×512, paired with their source photos and museum metadata.
In this rung, the scene varies textures, lighting and a glass sheet present with probability 0.25, while freezing frame variant/color, camera pose. Same schema, source images and index↔painting mapping as every other rung — they… See the full description on the dataset page: https://huggingface.co/datasets/patryk-bartkowiak/SynGallery-abl3-tex-light-glass.glami-1m-mteb
GLAMI-1M MTEB multimodal classification
This is an MTEB-ready derivative of the official
glami/glami-1m
release for multilingual image+text fashion classification. The source is
pinned at revision befda45d8d4e8b8082bb8a1912d1f9eb9483991c and remains
licensed under Apache-2.0.
Each example contains the official product image, name and description
joined as text, and the official category ID as label. The complete
116,004-row human-labeled test split is unchanged.
To keep… See the full description on the dataset page: https://huggingface.co/datasets/artist/glami-1m-mteb.glaucoma-detection
Glaucoma Detection
Retinal fundus images for glaucoma stage classification.
Splits
Split
Samples
train
2847
validation
1259
test
1272
Columns
image: retinal image
class: glaucoma stage label
Classes
Class
Description
normal
No glaucoma visible in the image.
early
Early-stage glaucoma findings.
advanced
Advanced glaucoma findings.
rock-glacier-datasetTODO: Add a description...OCT_And_Fundus_Glaucoma_Dataset
Data on OCT and Fundus Images
Dataset Description
This dataset is a copy of Data on OCT and Fundus Images which is shared with the license CC BY 4.0.
This dataset contains 50 samples under its train split. The dataset includes image data.
Splits
train: 50 samples
Data Fields
The dataset includes the following columns:
Image_Name: String data
Opt_1, Opt_2, Opt_3, and Opt_4: Struct/Dict data with two fields:
CDR: Float64 data
Glaucoma: String data… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/OCT_And_Fundus_Glaucoma_Dataset.index-card-blank-content
Index-card blank / content / divider classifier — dataset
Cropped single archival index cards labelled blank, content, or divider, for
training a tiny CPU pre-filter that skips blank/divider cards before expensive VLM metadata
extraction in card-catalogue digitisation pipelines.
Two collections: Boston Public Library (BPL) FRC shelf-list cards and National Library
of Scotland (NLS) Advocates Library cards. Styles differ, so evaluate per collection.
How it was made… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/index-card-blank-content.water_glassbottle_aesthetics_rated
Dataset Card for Dataset Name
This dataset holds 121 images of glass bottles for drinking water. The aesthetics were rated by five participants from Germany across different demographics.
