datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FedJam
FedJam Dataset
The FedJam dataset is a multimodal dataset for jamming detection and classification in wireless networks, combining time–frequency spectrogram images with
cross-layer network KPI time series. Each sample includes aligned vision and time-series modalities, allowing joint analysis of physical-layer signal behavior
and network-layer performance. The data are collected from a real over-the-air experimental testbed, under a variety of operating conditions, including… See the full description on the dataset page: https://huggingface.co/datasets/panitsasi/FedJam.random_streetview_images_pano_v0.0.2
Dataset Card for panoramic street view images (v.0.0.2)
Dataset Summary
The random streetview images dataset are labeled, panoramic images scraped from randomstreetview.com. Each image shows a location
accessible by Google Streetview that has been roughly combined to provide ~360 degree view of a single location. The dataset was designed with the intent to geolocate an image purely based on its visual content.
Supported Tasks and Leaderboards
None as of now!… See the full description on the dataset page: https://huggingface.co/datasets/stochastic/random_streetview_images_pano_v0.0.2.hateful-memes-data
Hateful Memes (CS5242 submission mirror)
Mirror of the Facebook Hateful Memes Challenge dataset (Kiela et al., 2020)
used for reproducibility of our CS5242 (NUS) submission.
Contents
img/ — 10,000 PNG images of memes
train.jsonl (8,500), dev_seen.jsonl (500), dev_unseen.jsonl (540),
test_seen.jsonl (1,000), test_unseen.jsonl (2,000)
Provenance
This mirror merges two existing mirrors of the original Meta release:
Label files and most images from… See the full description on the dataset page: https://huggingface.co/datasets/panjiyarsunil/hateful-memes-data.PANDA_MIL
PANDA - Multiple Instance Learning (MIL)
Important. This dataset is part of the torchmil library.
This repository provides an adapted version of the Prostate cANcer graDe Assessment (PANDA) dataset tailored for Multiple Instance Learning (MIL). It is designed for use with the PANDAMILDataset class from the torchmil library. PANDA is a widely used benchmark in MIL research, making this adaptation particularly valuable for developing and evaluating MIL models.
About the… See the full description on the dataset page: https://huggingface.co/datasets/torchmil/PANDA_MIL.PANDA-PLUS-Bench
PANDA-PLUS-Bench
A benchmark dataset for evaluating WSI-specific feature collapse in pathology foundation models.
Dataset Description
PANDA-PLUS-Bench contains expert-annotated prostate biopsy patches from 9 whole slide images (9 unique patients) with pixel-level Gleason pattern annotations.
Dataset Summary
Patches: ~2,770 per augmentation condition
Resolution: 224×224 pixels at 20× magnification
Classes: Benign (0), GP3 (1), GP4 (2), GP5 (3)
Slides: 9 (one… See the full description on the dataset page: https://huggingface.co/datasets/dellacorte/PANDA-PLUS-Bench.pangaea2-vhr
Mirror Notice
This repository contains unofficial mirrors of the following datasets, provided solely for hash-based versioning and reproducibility of research results. This is NOT the official source.
PureForest: https://huggingface.co/datasets/IGNF/PureForest
mpv4ger: https://huggingface.co/datasets/recursix/geo-bench-1.0
xView2: https://xview2.org/dataset
SpaceNet 3 Roads: https://spacenet.ai/spacenet-roads-dataset/ (S3: s3://spacenet-dataset/spacenet/SN3_roads/)
Legal Notice:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijrajsharma/pangaea2-vhr.classified_fr_road_signs
France road signs classification dataset
This dataset contains a total of 66000+ detected road signs from the Panoramax street level pictures.
The detection model used is available at https://huggingface.co/Panoramax/detect_face_plate_sign
250+ classes of road signs have been created, each matching a sign official type:
Axx signs = danger or warning
Bxx signs = restrictions / forbiden
Cxx signs = information
CExx signs = touristic information
etc.
Check… See the full description on the dataset page: https://huggingface.co/datasets/Panoramax/classified_fr_road_signs.Induction-Cooker-Ceramic-Panel-Crack-Identification-Dataset
Induction Cooker Ceramic Panel Crack Identification Dataset
In the current industrial field, the crack problem of induction cooker ceramic panels poses a threat to product safety, leading to potential explosion risks. Existing detection methods mostly rely on manual inspection, which is inefficient and prone to errors. This dataset aims to provide high-quality crack image data to train machine learning models, automating the detection process and improving detection efficiency and… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Induction-Cooker-Ceramic-Panel-Crack-Identification-Dataset.classified_nl_road_signsThis dataset has been created using Panoramax pictures from the Netherlands on which the https://huggingface.co/Panoramax/detect_face_plate_sign model has been used to detect road road signs and crop them.
It contain 24000+ photos of NL road signs in 150+ classes.
Additional "bad" or "other" classes contain non road signs (detection false positives or not yet classed signs).
For some classes, additionnal road signs have been added coming from european countries using similar signs.
The file… See the full description on the dataset page: https://huggingface.co/datasets/Panoramax/classified_nl_road_signs.classified_de_road_signsThis dataset has been created using Panoramax pictures from Germany on which the https://huggingface.co/Panoramax/detect_face_plate_sign model has been used to detect road road signs and crop them.
It contain 20000+ photos of DE road signs in 130+ classes.
Additional "bad" or "other" classes contain non road signs (detection false positives or not yet classed signs).
For some classes, additionnal road signs have been added coming from european countries using similar signs.
The file names are… See the full description on the dataset page: https://huggingface.co/datasets/Panoramax/classified_de_road_signs.classified_be_road_signsThis dataset has been created using Panoramax pictures from Belgium on which the https://huggingface.co/Panoramax/detect_face_plate_sign model has been used to detect road road signs and crop them.
It contain 24000+ photos of BE road signs in 140+ classes.
Additional "bad" or "other" classes contain non road signs (detection false positives or not yet classed signs).
For some classes, additionnal road signs have been added coming from european countries using similar signs.
The file names are… See the full description on the dataset page: https://huggingface.co/datasets/Panoramax/classified_be_road_signs.Pansy-Recognition-Image-Dataset
Pansy Recognition Image Dataset
Currently, garden management faces the challenge of efficiently and accurately identifying flower varieties. Traditional manual identification relies on experience and is inefficient. Existing image recognition technologies still need improvement in the accuracy of specific flower types, especially in complex backgrounds. This dataset aims to address common accuracy deficiencies in pansy recognition by providing a large number of high-quality images… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Pansy-Recognition-Image-Dataset.Geotagged_France_Panoramas
Dataset
DisasterM3
DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response
Junjue Wang*,
Weihao Xuan*,
Heli Qi, Zhihao Liu, Kunyi Liu, Yuhan Wu, Hongruixuan Chen,
Jian Song
Junshi Xia, Zhuo Zheng, Naoto Yokoya†
* Equal Contributions
† Corresponding Author
Paper: https://arxiv.org/abs/2505.21089
Code: https://github.com/Junjue-Wang/DisasterM3
Highlights
DisasterM3 includes 26,988 bi-temporal satellite images and 123k instruction pairs across… See the full description on the dataset page: https://huggingface.co/datasets/Pandamx/DisasterM3.Pansy-Recognition-Image-Dataset
Pansy Recognition Image Dataset
Currently, garden management faces the challenge of efficiently and accurately identifying flower varieties. Traditional manual identification relies on experience and is inefficient. Existing image recognition technologies still need improvement in the accuracy of specific flower types, especially in complex backgrounds. This dataset aims to address common accuracy deficiencies in pansy recognition by providing a large number of high-quality images… See the full description on the dataset page: https://huggingface.co/datasets/shangzx/Pansy-Recognition-Image-Dataset.opencs2_panorama
OpenCS2 — Panorama Dataset
360° panoramas captured in-engine from Counter-Strike 2 at positions
sampled from real player movement in HLTV demos. For each position the
dataset ships six cube-map faces (90° HFOV, 1024×1024) and the
stitched 4096×2048 equirectangular preview, plus the camera pose in
Source 2 hammer-unit coordinates (the same convention as
blanchon/opencs2_dataset).
Stat
Value
Maps
de_ancient, de_anubis, de_dust2, de_inferno, de_mirage, de_nuke… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/opencs2_panorama.stickers-binary-v2-cleaned
Stickers Binary v2 — Cleaned
Binary SFW/NSFW sticker classification dataset. This version has been cleaned
of likely label errors using cross-validated out-of-fold model predictions
combined with cleanlab's
find_label_issues.
Structure
This dataset has exactly two columns:
Column
Type
Description
image
image
The sticker image, 256x256, letterboxed (see below).
label
int64
0 = SFW, 1 = NSFW.
Class distribution
Split
Count… See the full description on the dataset page: https://huggingface.co/datasets/Pankaj8922/stickers-binary-v2-cleaned.biome-themed-pantanal-biopark-fish-tanks
Biome-Themed Pantanal Biopark Fish Tanks
This dataset contains 3654 images taken in the Pantanal Biopark.
A paper with further details and baseline results is to be published soon in PLOS ONE.
heb_synth_pangoline
Dataset Card for Hebrew Synthetic Pangoline Dataset
INFO: I'm not giving access to users with 0 models/0 datasets/0 activity - sharing is both ways
Dataset Summary
The Hebrew Synthetic Pangoline Dataset is a comprehensive collection of synthetic Hebrew document images generated using a custom implementation of Pangoline, a text-to-image synthesis tool. The dataset contains high-quality synthetic Hebrew text rendered as images, along with corresponding ground truth… See the full description on the dataset page: https://huggingface.co/datasets/johnlockejrr/heb_synth_pangoline.test
Dataset Card for Imagenette
Dataset Summary
A smaller subset of 10 easily classified classes from Imagenet, and a little more French.
This dataset was created by Jeremy Howard, and this repository is only there to share his work on this platform. The repository owner takes no credit of any kind in the creation, curation or packaging of the dataset.
Supported Tasks and Leaderboards
image-classification: The dataset can be used to train a model for Image… See the full description on the dataset page: https://huggingface.co/datasets/Pankajric22/test.yid_synth_pangolineINFO: I'm not giving access to users with 0 models/0 datasets/0 activity - sharing is both ways
Dataset Summary
The Yiddish Synthetic Pangoline Dataset is a comprehensive collection of synthetic Yiddish document images generated using a custom implementation of Pangoline, a text-to-image synthesis tool. The dataset contains high-quality synthetic Yiddish text rendered as images, along with corresponding ground truth text and ALTO-XML layout annotations. This dataset is designed… See the full description on the dataset page: https://huggingface.co/datasets/johnlockejrr/yid_synth_pangoline.
