datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anything-v3.0-glazed
Dataset Card for Anything v3.0 Glazed Samples
Dataset Description
Dataset Summary
This dataset contains image samples originally generated by Linaqruf/anything-v3.0
and subsequently processed by Glaze tool.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/anything-v3.0-glazed.stable-diffusion-v1-5-glazed
Dataset Card for Stable Diffusion v1.5 Glazed Samples
Dataset Description
Dataset Summary
This dataset contains image samples originally generated by runwayml/stable-diffusion-v1-5
and subsequently processed by Glaze tool.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/stable-diffusion-v1-5-glazed.Glaucoma_Dataset
Glaucoma Dataset
Dataset Summary
The Glaucoma Dataset is a comprehensive collection of retinal fundus images designed for the automated detection and classification of glaucoma. Containing between 10,000 and 100,000 high-quality images, this dataset aims to support the development and evaluation of machine learning and deep learning models in the field of ophthalmic medical imaging.
The dataset is organized using the standard imagefolder format, making it highly… See the full description on the dataset page: https://huggingface.co/datasets/Nj-1111/Glaucoma_Dataset.glaucoma-expert-cot-final
Glaucoma Expert Chain-of-Thought
Ophthalmologist six-step reasoning reports for fundus photographs, each paired with
a binary glaucoma label. 1,074 cases from LAG and Papila.
Files
file
rows
glaucoma / not
train.jsonl
823
304 / 519
val.jsonl
92
46 / 46
test.jsonl
159
79 / 80
images/
1,074
<source>_<id>.jpg
Record schema
{
"id": "1689",
"source": "LAG",
"image": "LAG_1689.jpg",
"split": "train",
"final_diagnosis_GT":… See the full description on the dataset page: https://huggingface.co/datasets/yuzhench/glaucoma-expert-cot-final.VTON-Synthetic-Pairs
Dataset Card for Glance VTON Synthetic Dataset
Dataset Summary
The Glance VTON Synthetic Dataset is a large-scale, fully synthetic paired dataset designed for training and evaluating Virtual Try-On (VTON) models. It was created to overcome the scarcity, high acquisition cost, restrictive commercial licensing, and demographic/stylistic biases of existing real-world paired VTON datasets.
The dataset features carefully curated pairs consisting of an isolated garment image… See the full description on the dataset page: https://huggingface.co/datasets/glanceai/VTON-Synthetic-Pairs.glami-1m-t2i-mteb
GLAMI-1M text-to-image retrieval
This MTEB-formatted derivative uses the complete 116,004-row official GLAMI-1M test split. Product names and descriptions are text queries and product images are the corpus. Repeated image IDs and exact repeated texts are deduplicated within each language, and qrels retain every observed text-image association.
The unchanged source archives are already hosted by the original authors in glami/glami-1m. GLAMI-1M-dataset--test-only.zip is pinned at… See the full description on the dataset page: https://huggingface.co/datasets/artist/glami-1m-t2i-mteb.noteflow-research-pilots
Keep the failed attempts. Check the artifact.
Versioned public development evidence from Robot Reel × Skills Anywhere × EvalArc, recorded 14 September 2026 on an NVIDIA L40S, with separate scripted Harbor controls on CPU and separate GPU context-control and agent-requested MCP handoff cohorts recorded 19 September 2026. This is an inspectable engineering casebook, not a held-out benchmark or training corpus with established efficacy.
Configuration
Actual experiment
What… See the full description on the dataset page: https://huggingface.co/datasets/glayguo/noteflow-research-pilots.scopes_test
Dataset Card for "scopes_test"
More Information needed
glaucoma-expert-cot-raw-1077
Glaucoma Expert Chain-of-Thought
Ophthalmologist six-step reasoning reports for fundus photographs, each paired with
a binary glaucoma label. 1,074 cases from LAG and Papila.
Files
file
rows
split
expert_cot_trainval.jsonl
915
train (823) + val (92)
expert_cot_test.jsonl
159
test
images/
1,074
<source>_<id>.jpg
Record schema
{
"id": "1689",
"source": "LAG",
"image": "LAG_1689.jpg",
"split": "train"… See the full description on the dataset page: https://huggingface.co/datasets/yuzhench/glaucoma-expert-cot-raw-1077.corruption-glass_blur
Corruption Dataset: Glass_Blur
Dataset Description
This dataset contains corrupted versions of ImageNet-1K images using glass_blur corruption. It is part of the ImageNet-C benchmark for evaluating model robustness to common image corruptions.
Dataset Structure
Train: 1,281,167 corrupted images
Validation: 50,000 corrupted images
Classes: 1000 ImageNet-1K classes
Format: Arrow (Hugging Face Datasets)
Corruption Type: Glass_Blur
Applies glass blur… See the full description on the dataset page: https://huggingface.co/datasets/MarMaster/corruption-glass_blur.rayban-meta-glasses
Ray-Ban Meta Glasses
Brand: Ray-Ban Meta
Item: link
ImageNet-C-glass_blur-severity_5glami-1m
GLAMI-1M contains 1.1 million fashion items, 968 thousand unique images and 1 million unique texts. It contains 13 languages, mostly European. And 191 fine-grained categories, for example we have 15 shoe types. It contains high quality annotations from professional curators and it also presents a difficult production industry problem.
Each sample contains an image, country code, name in corresponding language, description, target category and source of the label which can be of multiple types… See the full description on the dataset page: https://huggingface.co/datasets/glami/glami-1m.metacam-glasses
metacam-datasets
Currently holding glasses images
social-commerce-screenshot-claim-glance
Social Commerce Screenshot Claim Glance
This is a 100-image public preview subset for research on catalog-grounded visual retrieval and exact-claim authorization in social-commerce screenshots.
The dataset name for Hugging Face should be:
social-commerce-screenshot-claim-glance
What It Contains
The preview contains synthetic/redacted app-style screenshots that resemble noisy buyer or social-commerce inputs. The examples include platform-like UI framing, cropped… See the full description on the dataset page: https://huggingface.co/datasets/Sonjoy/social-commerce-screenshot-claim-glance.SynGallery-abl4-tex-light-glass-frame
SynGallery-abl4-tex-light-glass-frame: + frame variety
Rung 4 of the SynGallery instance-level artwork-recognition ablation ladder. 4,898 MET paintings × 5 camera viewpoints = 24,490 synthetic RGB images at 512×512, paired with their source photos and museum metadata.
In this rung, the scene varies textures, lighting, glass and frame molding variant + color/roughness/metallic, while freezing camera pose (the only frozen factor). Same schema, source images and index↔painting… See the full description on the dataset page: https://huggingface.co/datasets/patryk-bartkowiak/SynGallery-abl4-tex-light-glass-frame.glasses_celebGlacier-Dataset
Polar Glacier Bitemporal Remote Sensing Dataset
Dataset Overview
This dataset contains bitemporal remote sensing images from two representative polar regions:
Southeastern Coast of Greenland
(Latitude 64°–66°N, Longitude 51°–56°W):Dominated by glaciers and icefields, this area features exposed bedrock mountains and narrow coastal vegetation zones. It is a key region for studying glacier dynamics, with typical crevasse systems on the glacier surface and… See the full description on the dataset page: https://huggingface.co/datasets/cuibinge/Glacier-Dataset.TWIN
TWIN
This repository contains the TWIN dataset introduced in the paper Same or Not? Enhancing Visual Perception in Vision-Language Models. TWIN contains 561K challenging (image, question, answer) tuples emphasizing fine-grained image understanding.
For evaluating on the dataset with LMMS-eval, please refer to this repo.
Citation
If you use the TWIN dataset in your research, please use the following BibTeX entry.
@misc{marsili2025notenhancingvisualperception… See the full description on the dataset page: https://huggingface.co/datasets/glab-caltech/TWIN.MICCAI_FLARE_glaucomarobocasa_20260430T030150Z_full_run_setup_wine_glasses ---
pretty_name: RoboCasa Trajectories Single
configs:
- config_name: default
data_files:
- split: train
path: data/train-*
---
# RoboCasa Trajectories Single
This dataset contains one row per RoboCasa trajectory / episode.
## Structure
Each row is one trajectory / episode.
Episode-level JSON is stored inline:
adapted_trajectory
original_trajectory
execution_metadata
Step-level data is stored in aligned sequence columns:… See the full description on the dataset page: https://huggingface.co/datasets/DorianAtSchool/robocasa_20260430T030150Z_full_run_setup_wine_glasses.terrain-referenced-3d-glacier-mapping-product
Terrain-referenced Glacier Mapping Product
This repository provides the terrain-referenced glacier-area mapping product generated for the manuscript. The product is openly available through Hugging Face with DOI: 10.57967/hf/9900.
The archive contains regional mapping outputs, oblique terrain-visualization products, metadata files, and tabular glacier-area summaries. Glacier masks generated by Prithvi-SDT are linked with Copernicus DEM terrain information and RGI 7.0 glacier… See the full description on the dataset page: https://huggingface.co/datasets/yyhw/terrain-referenced-3d-glacier-mapping-product.SynGallery-abl3-tex-light-glass
SynGallery-abl3-tex-light-glass: + glass
Rung 3 of the SynGallery instance-level artwork-recognition ablation ladder. 4,898 MET paintings × 5 camera viewpoints = 24,490 synthetic RGB images at 512×512, paired with their source photos and museum metadata.
In this rung, the scene varies textures, lighting and a glass sheet present with probability 0.25, while freezing frame variant/color, camera pose. Same schema, source images and index↔painting mapping as every other rung — they… See the full description on the dataset page: https://huggingface.co/datasets/patryk-bartkowiak/SynGallery-abl3-tex-light-glass.GLAMI-1M
This is fork of original dataset converted to dataset format.
GLAMI-1M contains 1.1 million fashion items, 968 thousand unique images and 1 million unique texts. It contains 13 languages, mostly European. And 191 fine-grained categories, for example we have 15 shoe types. It contains high quality annotations from professional curators and it also presents a difficult production industry problem.
Each sample contains an image, country code, name in corresponding language… See the full description on the dataset page: https://huggingface.co/datasets/pySilver/GLAMI-1M.scopes
Dataset Card for "scopes"
More Information needed
synthetic-glass-with-liquid-filled
🥃 Glass Half Full — Synthetic Glass with Liquid Filled
8,000 synthetic images of drinking glasses with varying liquid fill levels,
rendered with Blender Cycles (physically-based path tracer) at 256×256
resolution. Every image ships with perfect YOLO-format bounding-box labels
for two classes — glass and liquid — computed directly from 3D geometry
(no human annotation).
Built for the Existential Glass Analyzer,
a browser-based model that answers the timeless question: is your… See the full description on the dataset page: https://huggingface.co/datasets/Aspirin4/synthetic-glass-with-liquid-filled.Synthetic-Glass-Transparent-Packaging-Dataset-Sample
Transparent Packaging & Glass Benchmark
Watch our benchmark breakdown: Why Object Detection Fails on Glass.
Synthetic Transparent Glass & Packaging Dataset
A photorealistic synthetic computer vision dataset for transparent glass and
packaging object detection and instance segmentation. The dataset is
designed for models dealing with challenging transparent and reflective
materials, including transparency, reflections, refractions, specular
highlights, and harsh… See the full description on the dataset page: https://huggingface.co/datasets/Ji0134ch/Synthetic-Glass-Transparent-Packaging-Dataset-Sample.glami-1m-mteb
GLAMI-1M MTEB multimodal classification
This is an MTEB-ready derivative of the official
glami/glami-1m
release for multilingual image+text fashion classification. The source is
pinned at revision befda45d8d4e8b8082bb8a1912d1f9eb9483991c and remains
licensed under Apache-2.0.
Each example contains the official product image, name and description
joined as text, and the official category ID as label. The complete
116,004-row human-labeled test split is unchanged.
To keep… See the full description on the dataset page: https://huggingface.co/datasets/artist/glami-1m-mteb.imnet1k_sunglasses_dark_glasses_shadesidktest
