datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SNAP
SNAP Benchmark
Code and annotations: [https://github.com/ykotseruba/SNAP]
SNAP (stands for Shutter speed, ISO seNsitivity, and APerture) is a new benchmark consisting of images of objects taken under controlled lighting conditions and with densely sampled camera settings.
This benchmark allows testing the effects of capture bias, which includes camera settings and illumination, on performance of vision algorithms.
SNAP contains 37,558 images of 100 scenes (10 scenes per 10 object… See the full description on the dataset page: https://huggingface.co/datasets/ykotseruba/SNAP.gelsight-mini-pretrain
GelSight Mini Pretrain
~853K GelSight Mini tactile RGB frames, 12 public sources, one parquet schema. Built for self-supervised representation learning (VAE / MAE / SimCLR / DINO) — every frame contact-filtered, channel-normalized, and re-encoded as JPEG q92.
Frames
Sources
Real
536K
FoTA (labeled+unlabeled), 3DCal, FEATS, GelSLAM, TactileTracking, RTM, FeelAnyForce, UniT, TacQuad
Sim
317K
sim_tactile_mnist, sim_starstruck (Taxim-rendered, Mini-calibrated)
NC… See the full description on the dataset page: https://huggingface.co/datasets/yxma/gelsight-mini-pretrain.MethaneUnion
MethaneUnion
This dataset product is organized as:
datasets/
temporal_split/
original_scale/{train,test}.csv
120m_GSD/{train,test}.csv
360m_GSD/{train,test}.csv
480m_GSD/{train,test}.csv
960m_GSD/{train,test}.csv
geo_split/
original_scale/{train,test}.csv
120m_GSD/{train,test}.csv
360m_GSD/{train,test}.csv
480m_GSD/{train,test}.csv
960m_GSD/{train,test}.csv
Each CSV contains exactly these columns: id, label, latitude, longitude… See the full description on the dataset page: https://huggingface.co/datasets/yuyao42/MethaneUnion.AU-OPG
AU-OPG: Panoramic Dental Radiographs With Oriented Tooth-Level Annotations
AU-OPG (Ajman University Orthopantomography) is a dataset of 901 panoramic dental radiographs annotated for:
oriented tooth detection;
radiographic diagnosis; and
radiographic-evidence-based treatment planning.
The release contains 7,006 tooth-level annotations. Each annotated tooth has a tooth-aligned oriented bounding box, one diagnostic condition, and one corresponding treatment label. The predefined… See the full description on the dataset page: https://huggingface.co/datasets/YSFF/AU-OPG.NIH-CXR14-BiomedCLIP-Features
NIH-CXR14-BiomedCLIP-Features Dataset
This dataset is derived from the NIH Chest X-ray Dataset (NIH-CXR14) and processed using the BiomedCLIP-PubMedBERT_256-vit_base_patch16_224 model from Microsoft. It contains image and text features extracted from chest X-ray images and their corresponding textual findings.
Dataset Description
The original NIH-CXR14 dataset comprises 112,120 chest X-ray images with disease labels from 30,805 unique patients. This processed dataset… See the full description on the dataset page: https://huggingface.co/datasets/Yasintuncer/NIH-CXR14-BiomedCLIP-Features.gelsight-mini-pretrain-nc
GelSight Mini Pretrain · Non-Commercial Extension
⚠️ Non-commercial use only. This repository is licensed
CC-BY-NC-4.0 because it includes upstream sources whose licenses
restrict commercial use. For commercial-friendly Mini tactile data,
see the main yxma/gelsight-mini-pretrain repo
(CC-BY-4.0).
This dataset is the CC-BY-NC extension of yxma/gelsight-mini-pretrain.
It contains only the GelSight Mini sources whose upstream licenses are
not compatible with CC-BY-4.0 aggregation.… See the full description on the dataset page: https://huggingface.co/datasets/yxma/gelsight-mini-pretrain-nc.MOUNT-Cattle
Updates/News 📣
🎉 News (Feb. 2026): The dataset paper FSMC-Pose has been accepted for CVPR 2026 Findings!
🔗 News: Please find the open-source dataset on Hugging Face: MOUNT-Cattle.
🔥 Downloads reached 2.4k within 7 days of release.
📌 Overview
Mounting posture is an important visual indicator of estrus in dairy cattle. MOUNT-Cattle is a mounting dataset, covering 1,176 mounting… See the full description on the dataset page: https://huggingface.co/datasets/y1665065879/MOUNT-Cattle.gelsight-mini-pretrain-video
GelSight Mini Pretrain · Video / Sequence Subset
🎬 Companion to yxma/gelsight-mini-pretrain.
Where the main repo treats every kept frame as an independent image, this repo
preserves temporal sequences — one row per frame, ordered, with explicit
sequence-id + position metadata, for video tactile pretraining.
Why this repo
The main repo's pipeline applies perceptual-hash dedupe within each capture
to drop near-identical adjacent frames. That's great for image-level… See the full description on the dataset page: https://huggingface.co/datasets/yxma/gelsight-mini-pretrain-video.Agrotools
ThinkGeoBench Question-4
This repository contains an upload-ready Hugging Face release folder for the question-4 split of ThinkGeoBench.
Included files
metadata.jsonl: normalized table for Hugging Face Dataset Viewer
images/: image assets referenced by question-4
assets/AppleSizeEstimate/: depth .npy assets referenced by Apple size estimation samples
question-4.original.json: original source question file
question_taxonomy_summary.md: taxonomy and template notes… See the full description on the dataset page: https://huggingface.co/datasets/yz686868/Agrotools.
