datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
weapon-detection-runs-backup-2026-07-24weapon-detection-workerssd-backup-2026-07-24CIDER
CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment
Paper | Code
Dataset for the COLM 2026 paper CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment
CIDER is a dataset of privacy disclosure decisions collected from real users. It consists of 14,850 annotations from 169 users, forming 1,650 contextual disclosure boundary sets across 60 interpersonal communication scenarios.
What can you do with… See the full description on the dataset page: https://huggingface.co/datasets/peach-lab/CIDER.Arabic-VLM-Full-Pearl
💎 The Arabic VLM Dataset (Full Pearl Edition)
This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper.
Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.PEACE
PEACE: Empowering Geologic Map Holistic Understanding with MLLMs
[Code] [Paper] [Data]
Introduction
We construct a geologic map benchmark, GeoMap-Bench, to evaluate the performance of MLLMs on geologic map understanding across different abilities, the overview of it is as shown in below Table.
Property
Description
Source
USGS(English)
CGS(Chinese)
Content
Image-question pair… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/PEACE.peanuts-opt-6.7b
Peanut Comic Strip Dataset (Snoopy & Co.)
This is a dataset Peanuts comic strips from 1950/10/02 to 2000/02/13.
There are 77,457 panels extracted from 17,816 comic strips.
The dataset size is approximately 4.4G.
Each row in the dataset contains the following fields:
image: PIL.Image containing the extracted panel.
panel_name: unique identifier for the row.
characters: tuple[str, ...] of characters included in the comic strip the panel is part of.
themes: tuple[str, ...] of theme… See the full description on the dataset page: https://huggingface.co/datasets/afmck/peanuts-opt-6.7b.StreamGaze_v2
StreamGaze Dataset
StreamGaze is a comprehensive streaming video benchmark for evaluating MLLMs on gaze-based QA tasks across past, present, and future contexts.
Companion dataset: The EgoGazeVQA dataset is hosted separately at Peanuttoad/gaze_dataset.
📁 Dataset Structure
streamgaze/
├── metadata/
│ ├── egtea.csv # EGTEA fixation metadata
│ ├── egoexolearn.csv # EgoExoLearn fixation metadata
│ └── holoassist.csv # HoloAssist… See the full description on the dataset page: https://huggingface.co/datasets/Peanuttoad/StreamGaze_v2.PEARLSignature-Verification-Dataset
Multilingual Signature Verification Dataset
Dataset Summary
The Multilingual Signature Verification Dataset is a curated collection of handwritten signatures designed for offline signature verification and related computer vision tasks.
The dataset contains more than 7,000 signature images spanning three major writing systems:
Hindi
Bengali
English
The English portion includes samples from the well-known CEDAR Signature Dataset, while additional Hindi and… See the full description on the dataset page: https://huggingface.co/datasets/Peachyy2208/Signature-Verification-Dataset.tcg-frame-removal-dataset
TCG Frame Removal Dataset
547 paired examples for training instruction-editing models that strip the frame, text,
and UI elements from trading-card images and extend the artwork to a seamless full-bleed
illustration. This is the training set for the
TCG Frame Removal LoRA (FLUX.2-Klein 4B) model
(weights).
Game
Pairs
Magic: The Gathering
304
Digimon
154
Pokémon
63
Yu-Gi-Oh!
26
Fields
id (string): unique card slug, prefixed by game (mtg-… See the full description on the dataset page: https://huggingface.co/datasets/pearsonkyle/tcg-frame-removal-dataset.PEaCEpeabody-testPEaCE-512pxpeanuts-flan-t5-xl
Peanut Comic Strip Dataset (Snoopy & Co.)
This is a dataset Peanuts comic strips from 1950/10/02 to 2000/02/13.
There are 77,456 panels extracted from 17,816 comic strips.
The dataset size is approximately 4.4G.
Each row in the dataset contains the following fields:
image: PIL.Image containing the extracted panel.
panel_name: unique identifier for the row.
characters: tuple[str, ...] of characters included in the comic strip the panel is part of.
themes: tuple[str, ...] of theme… See the full description on the dataset page: https://huggingface.co/datasets/afmck/peanuts-flan-t5-xl.pear-pests-dataset
🍐 Pear Leaf Pest & Disease Detection Dataset
This dataset supports object detection for pear leaf health monitoring: locating and classifying an insect pest and fungal/bacterial disease lesions directly on pear leaf images. It targets the timely detection and localization of foliar pear pests and diseases central to precision agriculture, where manual agronomist inspection is labor-intensive, subjective, and hard to scale.
The dataset consists of 2,210 annotated images (1,542… See the full description on the dataset page: https://huggingface.co/datasets/salahkhenfer/pear-pests-dataset.terrariumfloorplan-room-segmentation
Floorplans Dataset
This dataset is derived from the Floorplans Diff dataset and has been curated by removing all unannotated images to ensure clean and consistent training data.
It is designed for semantic image segmentation, specifically focusing on identifying and segmenting rooms within floorplan images.
Each sample consists of an image paired with a corresponding segmentation mask, enabling models to learn pixel-level classification for the room class.
Origin
This… See the full description on the dataset page: https://huggingface.co/datasets/peaceAsh/floorplan-room-segmentation.10132025_pick_redspatula_peachtableclothThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 50,
"total_frames": 9956,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/10132025_pick_redspatula_peachtablecloth.scans
scans
Facsimiles of primary sources of the French Revolution. Scans, and nothing
else — 39,668 files, 155 GB: 39,536 PDFs, 94 JP2 map sheets, 30 loose page
images and 8 DjVu volumes.
The records that describe these files, the transcriptions and OCR of what is
printed on them, and the tool that fetched and bound them all live in the corpus
that cites this archive: https://github.com/peasanttide/history. This
repository is the bytes.
That split is the point. A facsimile is stable… See the full description on the dataset page: https://huggingface.co/datasets/peasanttide/scans.harmful_memePEARL
💎 PEARL Benchmark
Dataset ID: UBC-NLP/PEARL
Description:
The PEARL Benchmark is a meticulously curated subset of 6,867 high-quality Question/Answer pairs derived from the larger PEARL dataset. It is specifically designed for evaluating Vision-Language Models' (VLMs) understanding of Arab cultural content. The benchmark covers ten significant cultural domains (e.g., architecture, clothing, cuisine) and thirteen diverse question types, tailored for robust model assessment of… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/PEARL.pear
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/etahamad/pear.image-n_peaks-sft-swift09182025_realsense_peachflowertable_bluedinosaurmugThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 50,
"total_frames": 36167,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/09182025_realsense_peachflowertable_bluedinosaurmug.color-fundus-eyeINFINITY project on cfp images gathered from public
This dataset contains retinal fundus images.Files are organized by class in separate folders.
harmful_memesbanners-directv_peach
Dataset Card for "banners-directv_peach"
More Information needed
09182025_realsense_peachflowertable_bluedinosaurmug_bak2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 51,
"total_frames": 36348,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/09182025_realsense_peachflowertable_bluedinosaurmug_bak2.index-cards-peabody-newspaper
Peabody Newspaper Index Cards (Peabody Institute Library, MA)
3,694 typewritten index cards from the Peabody Institute Library — Sutton
Room Local History Resource Center (Peabody, Massachusetts), indexing people,
events, and news in South Danvers / Peabody as recorded in local newspapers.
The information was typed onto cards over decades by library staff as the local
newspaper-of-record archive's principal finding aid.
Plus a companion "Poor Family" genealogy index from the same… See the full description on the dataset page: https://huggingface.co/datasets/biglam/index-cards-peabody-newspaper.09182025_realsense_peachflowertable_bluedinosaurmug_bakThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 50,
"total_frames": 36167,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/rli14/09182025_realsense_peachflowertable_bluedinosaurmug_bak.
