datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EMID-Emotion-Matching
EMID-Emotion-Matching
orrzohar/EMID-Emotion-Matching is a derived dataset built on top of
the Emotionally paired Music and Image Dataset (EMID) from ECNU (ecnu-aigc/EMID).
It is designed for music ↔ image emotion matching with Qwen-Omni–style models.
Each example contains:
audio: mono waveform stored as datasets.Audio (HF Hub preview can play it)
sampling_rate: sampling rate used when decoding (typically 16 kHz)
image: a single image (datasets.Image)
same: bool, whether the audio… See the full description on the dataset page: https://huggingface.co/datasets/orrzohar/EMID-Emotion-Matching.tibetan-page-orientation-classifier-dataset
Tibetan Page Orientation Dataset
Covers 7 Tibetan script families: Danyig, Druma, Gyuyig, Multi-Scripts, Pedri, Tsugdri, Uchen.
Dataset composition
Each manuscript page appears twice: once as the original scan (non_flipped) and once rotated 180° (flipped). The model's task is to distinguish these two orientations.
Scripts are balanced — each of the 7 script families contributes the same number of pages (downsampled to the smallest family).
Script (script)… See the full description on the dataset page: https://huggingface.co/datasets/BDRC/tibetan-page-orientation-classifier-dataset.Watermark-or-Not-20K
Watermark-or-Not-20K Dataset
Overview
The Watermark-or-Not-20K dataset consists of 20,000 images annotated with binary labels indicating the presence or absence of a watermark. It is designed to support training and evaluation of models focused on watermark detection, which is useful for content filtering, copyright protection, and image moderation tasks.
Dataset Structure
Split: train
Number of samples: 20,000
Label Type: Categorical (2 classes)
Image… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Watermark-or-Not-20K.dental-implant-surgery-sample
Dental Implant Surgery — Multimodal Annotated Video (Sample Case)
A public sample from one complete All-on-4 full-arch mandibular dental implant
surgery: two synchronised camera angles, the operating surgeon narrating while
he works, and six layers of structured clinical annotation (L0–L5) tied frame by
frame to what he said.
This is a showcase slice, not the whole case. What is here is enough to judge
the structure, the annotation quality and the honesty of the documentation.… See the full description on the dataset page: https://huggingface.co/datasets/OralSurgery/dental-implant-surgery-sample.3d-printed-or-not
3d-printed-or-not: An Image Dataset of 3D-printed Prototypes
This dataset is a collection of images that are particularly relevant to engineering and design, consisting of two categories: 3D-printed prototypes, and non-3D-printed prototypes This data was collected through a hybrid approach that entailed both web scraping and direct collection from engineering labs and workspaces at Penn State University. The initial data was then augmented using several data augmentation techniques… See the full description on the dataset page: https://huggingface.co/datasets/cmudrc/3d-printed-or-not.2026-24679-HW1-Multimodal-Original
Straight-member torque: image and structured statics data
eandujar/2026-24679-HW1-Multimodal-Original
100 synthetic planar-statics cases containing a rendered diagram and
structured numerical/categorical features describing the member geometry,
supports, and applied loads.
The prediction task has two outputs:
Torque direction — clockwise or counterclockwise.
Torque magnitude — absolute moment about the pin in N m.
This therefore supports both classification and regression… See the full description on the dataset page: https://huggingface.co/datasets/eandujar/2026-24679-HW1-Multimodal-Original.fashion-lookmatch-dataset
👔 Fashion LookMatch Synthetic Dataset
Overview
This dataset contains 1,000 synthetic fashion images generated using Stable Diffusion XL (SDXL).
It was created as part of the "Build Your Own AI Application" final project.
The dataset is designed for building a Visual Retrieval & Outfit Recommendation System.
📊 Exploratory Data Analysis (EDA)
To ensure the quality and balance of the dataset, we performed a rigorous EDA.
1. Category Balance
We ensured… See the full description on the dataset page: https://huggingface.co/datasets/orianrivlin/fashion-lookmatch-dataset.face-or-not
Face or Not
theoriclabs/face-or-not is a balanced binary image-classification dataset of
4,000 128x128 RGB crops. The label answers one narrow question: does this crop
contain an Open Images Human face annotation?
This is classification, not face localization, identification, recognition,
or biometric matching. It has no names or identity labels.
Splits
split
no_face
face
total
train
1,600
1,600
3,200
validation
200
200
400
test
200
200
400… See the full description on the dataset page: https://huggingface.co/datasets/theoriclabs/face-or-not.Orchid2024_raw
Dataset Card for Orchid2024
The Orchid2024 dataset is a fine-grained classification dataset specifically designed for cultivars of Chinese Cymbidium orchids (Chinese orchids). The dataset's samples come from 20 cities across 12 provincial-level administrative regions in China, covering 1,269 cultivars from 8 Cymbidium species, and 6 additional categories, totaling 156,630 images. The dataset nearly includes all common Chinese orchid cultivars currently found in China. Its… See the full description on the dataset page: https://huggingface.co/datasets/HelloPlant/Orchid2024_raw.dustbin-or-not-dataset
Dustbin vs. Not Dustbin Image Dataset
Dataset Summary
A small binary image classification dataset for predicting whether an image contains a dustbin/trash can.
has_dustbin = 1: dustbin visible
has_dustbin = 0: no dustbin visible
30+ original photographs
Images resized to 224 × 224 RGB
Data Collection
Original photographs were captured by the author and organized into dustbin/ and not_dustbin/ folders. No people, faces, or sensitive personal… See the full description on the dataset page: https://huggingface.co/datasets/srivathsanb14/dustbin-or-not-dataset.Orchid2024
Dataset Card for Orchid2024
The Orchid2024 dataset is a fine-grained classification dataset specifically designed for cultivars of Chinese Cymbidium orchids (Chinese orchids). The dataset's samples come from 20 cities across 12 provincial-level administrative regions in China, covering 1,269 cultivars from 8 Cymbidium species, and 6 additional categories, totaling 156,630 images. The dataset nearly includes all common Chinese orchid cultivars currently found in China. Its… See the full description on the dataset page: https://huggingface.co/datasets/HelloPlant/Orchid2024.Belarusian_ornament
