datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Defactify_Image_Dataset
Defactify_Image_Dataset
This dataset is associated with the paper A Comprehensive Dataset for Human vs. AI Generated Image Detection.
📝 Dataset Description
Dataset Summary
The Defactify_Image_Dataset (A Comprehensive Dataset for Human vs. AI Generated Image Detection) is a high-quality collection of 96,000 images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. Built using the MS… See the full description on the dataset page: https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Image_Dataset.MetaPKLot-Dataset
MetaPKLot
A Large-Scale Benchmark for Vision-Based Parking Lot Management
2,265,974 labeled samples · 1,366,185 new annotations · 3 research challenges · COCO-style annotations
MetaPKLot is a large-scale, harmonized dataset designed for research on vision-based parking lot management.
It extends and standardizes three existing parking datasets:
PKLot
CNRPark-EXT
PLds
MetaPKLot introduces new annotations, revises existing parking-space annotations, standardizes… See the full description on the dataset page: https://huggingface.co/datasets/DSBD-Research/MetaPKLot-Dataset.sawhill-dataset
Sawhill Numismatic Collection Dataset
Dataset Description
This dataset contains video recordings and extracted images of coins from the MacKenzie Art Gallery's Sawhill Numismatic Collection. The dataset is designed for research in automated coin identification, cultural heritage digitization, and computer vision applications in numismatics.
Dataset Summary
Source: MacKenzie Art Gallery Sawhill Numismatic Collection
Content: Handheld video recordings of coins… See the full description on the dataset page: https://huggingface.co/datasets/COIN-Research-Group/sawhill-dataset.auto-research-bench-data
auto-research-bench-data
同行评审语料 + 论文 PDF,覆盖五个机器学习会议 2022–2026 年。
用于研究「AI 生成的论文与人类论文有何差异」,特别是图表质量与实验管理两个维度。
数据来自 OpenReview,用官方 API 抓取。
内容
1. 评审元数据(18 个 jsonl,3.11 GB)
每行一篇投稿,字段如下:
字段
说明
forum
OpenReview 论文 ID,与 PDF 文件名一致,是关联两部分数据的 key
title / abstract / keywords
论文元信息
venue / venueid
录用层级。注意层级只在 venue 里(如 ICLR 2024 oral),venueid 对所有录用论文都是 .../Conference
reviews
全部 Official_Review,含评分、置信度、正文各字段
comments
作者 rebuttal 与其他… See the full description on the dataset page: https://huggingface.co/datasets/chengwanru/auto-research-bench-data.lanternfly_research_dataset
Lantern Fly Research Dataset
This dataset contains human-verified spotted lanternfly sightings collected through the Lantern Fly Tracker app. Each entry includes high-quality photos, precise geolocation data, and comprehensive metadata for ecological research.
🎯 Purpose
This dataset supports:
Ecological research on spotted lanternfly distribution and spread patterns
Machine learning model training with verified, high-quality data
Temporal and spatial analysis of… See the full description on the dataset page: https://huggingface.co/datasets/rlogh/lanternfly_research_dataset.Agricultural-Climate-Adaptation-Research-Dataset
Agricultural Climate Adaptation Research Dataset
Agriculture is currently facing challenges posed by climate change, particularly the increasing impact of drought on crop yields. Existing research data often lacks detailed analysis under specific climate conditions, leading to ineffective agricultural management measures. This dataset aims to fill this gap by including images of farmland drought and vegetation recovery, assisting AI models in researching agriculture's ability to… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Agricultural-Climate-Adaptation-Research-Dataset.
