datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AU-OPG
AU-OPG: Panoramic Dental Radiographs With Oriented Tooth-Level Annotations
AU-OPG (Ajman University Orthopantomography) is a dataset of 901 panoramic dental radiographs annotated for:
oriented tooth detection;
radiographic diagnosis; and
radiographic-evidence-based treatment planning.
The release contains 7,006 tooth-level annotations. Each annotated tooth has a tooth-aligned oriented bounding box, one diagnostic condition, and one corresponding treatment label. The predefined… See the full description on the dataset page: https://huggingface.co/datasets/YSFF/AU-OPG.osm-europe-1k
OSM-Europe-1k
A 1,000-image street-level geolocation benchmark for Europe, sampled from
the OpenStreetView-5M (OSV-5M)
test split. Intended as a contamination-free, openly-licensed reference set
for evaluating image→GPS models. Companion benchmark to a master's thesis at
FH JOANNEUM (Florian Leber, 2026).
What's in this repo:
The 1,000 OSM/Mapillary image bytes are bundled directly under
images/ (66 MB) — re-distribution is allowed by the upstream
CC-BY-SA 4.0 license, with… See the full description on the dataset page: https://huggingface.co/datasets/lebfla11/osm-europe-1k.OsteosarcomaTumorAssessment
Osteosarcoma data from UT Southwestern/UT Dallas for Viable and Necrotic Tumor Assessment (Osteosarcoma-Tumor-Assessment)
Unofficial fork.
Folder Structure
ML_Features_1144.csv # Contains 1144 rows for all the image tiles and 69 columns for filename, classification, and 65 machine learning features.
OsteosarcomaTumorAssessment.tar.zst
|-- Osteosarcoma-UT.sums
|-- Training-Set-1 # 11 folders with 547 images. Each folder contains 48~50 image tiles and 1 csv for… See the full description on the dataset page: https://huggingface.co/datasets/CAIR-M3LLM/OsteosarcomaTumorAssessment.dataset-openmoji
Dataset OpenMoji
Creator: https://www.kaggle.com/krayc81This is base on https://openmoji.org/License https://creativecommons.org/licenses/by-sa/4.0
Files:
README.md this :)
data.csv containing all data see bellow description
openmoji folder containing the image files
The data.csv contains:
idx the character as int
character text representation
bytes representation
hex representation (replace Ox with U+ for unicode)
description of the emoji
path_black path to the bw image… See the full description on the dataset page: https://huggingface.co/datasets/Kray-C/dataset-openmoji.flickr8k-sau-pace-annotated
Annotation
Annotated this dataset by clasifying the images into slow, medium or fast depending on the suitable paced background music.
Disaster-Type_Classification_Dataset_for_Automated_Fact-Checking
DTCD-AFC: Disaster-Type Classification Dataset for Automated Fact-Checking
Overview
The DTCD-AFC is a dataset designed for disaster-type classification evaluation for automated fact-checking.
It consists of multimodal social media posts collected based on past natural disasters, each labeled with the disaster type to which its content relates.
The social media posts are sourced from CrisisMMD.
Files
disaster_type_classification_dataset_for_afc.csv: The CSV… See the full description on the dataset page: https://huggingface.co/datasets/o-yas/Disaster-Type_Classification_Dataset_for_Automated_Fact-Checking.
