datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cocollama3-jailbreaksgemma2-jailbreaksindoor-safety-hazard-detection-and-work-zone-monitoring
Indoor Safety Hazard Detection & Work-Zone Monitoring
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/indoor-safety-hazard-detection-and-work-zone-monitoring.lie-detection-rollouts
Lie Detection Rollouts
Assistant completions across many open-weight models on the lie-detection
evaluation suite used by the
deception research pipeline. One subset per model,
one split per task.
Columns
messages — list of OpenAI-style messages. Each message has:
role: system | user | assistant
content: final message text
reasoning_content: chain-of-thought for reasoning models, None otherwise
is_lie — ground-truth label from the is_deceptive scorer:
lie |… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/lie-detection-rollouts.Safety-Toxicity-Detection
SEA Toxicity Detection
SEA Toxicity Detection evaluates a model's ability to identify toxic content such as hate speech and abusive language in text. It is sampled from MLHSD for Indonesian, TTD for Thai, and ViHSD for Vietnamese.
Supported Tasks and Leaderboards
SEA Toxicity Detection is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Indonesian (id)
Thai… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Safety-Toxicity-Detection.bedroom-change-detection-state-monitoring
Bedroom Change Detection & State Monitoring
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by the sync… See the full description on the dataset page: https://huggingface.co/datasets/physicl/bedroom-change-detection-state-monitoring.CS2CD.Counter-Strike_2_Cheat_Detection
Counter Strike 2 Cheat Detection Dataset
Overview
The CS2CD (Counter-Strike 2 Cheat Detection) dataset is an anonymised dataset comprised of Counter-Strike 2(CS2) gameplay at a variety of skill-levels with cheater annotations. This dataset contains 478 CS2 matches with no cheater present, and 317 matches CS2 matches with at least one cheater present.
Dataset structure
The dataset is partitioned into data with at least one cheater present, and data with no… See the full description on the dataset page: https://huggingface.co/datasets/CS2CD/CS2CD.Counter-Strike_2_Cheat_Detection.code_x_glue_cc_defect_detection
Dataset Card for "code_x_glue_cc_defect_detection"
Dataset Summary
CodeXGLUE Defect-detection dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/Defect-detection
Given a source code, the task is to identify whether it is an insecure code that may attack software systems, such as resource leaks, use-after-free vulnerabilities and DoS attack. We treat the task as binary classification (0/1), where 1 stands for insecure code and 0 for secure… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_defect_detection.indoor-anomaly-detection-path-obstruction-monitoring
Indoor Anomaly Detection & Path Obstruction Monitoring
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled… See the full description on the dataset page: https://huggingface.co/datasets/physicl/indoor-anomaly-detection-path-obstruction-monitoring.deepfake-audio-detection
Deepfake Audio Detection Dataset (v4)
Dataset Description
This dataset contains 1,866 audio samples (933 real, 933 synthetic) for training deepfake audio detection models. It is specifically designed for binary classification tasks to distinguish between authentic human speech and AI-generated synthetic audio.
What's New in v4
52% larger: Increased from 1,224 to 1,866 samples (642 new samples)
Expanded TTS coverage: Added Hume AI as 6th synthetic voice… See the full description on the dataset page: https://huggingface.co/datasets/garystafford/deepfake-audio-detection.drone-audio-detection-samples
Dataset Description
Drone Audio Detection Samples (DADS) is currently the largest publicly available drone audio database, specifically designed for developing drone detection systems using deep learning techniques. All audio files are standardized to a sample rate of 16,000 Hz, 16-bit depth, mono-channel, and vary in length from 500 milliseconds to several minutes.
Most drone audio files were manually trimmed to ensure that a drone was always present in the recording. However, some… See the full description on the dataset page: https://huggingface.co/datasets/geronimobasso/drone-audio-detection-samples.AIGC-Detection-Benchmark
AIGC Detection Benchmark Dataset
📝 Dataset Description
Dataset Summary
The AIGC Detection Benchmark Dataset is a high-quality collection of images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. The dataset contains a mix of real-world images and images generated by a wide array of prominent AI models, including diffusion models (like Stable Diffusion, DALL-E 2, Midjourney, ADM) and GANs… See the full description on the dataset page: https://huggingface.co/datasets/TheKernel01/AIGC-Detection-Benchmark.llm-bias-detection
LLM Bias Detection Evaluation Traces
Evaluation data accompanying Navigating the digital spectrum: Assessing
political bias, stability, and downstream fairness in Large Language Models
(arXiv:2609.08637).
Licence scope: CC BY 4.0 covers the authors' original documentation,
templates, selection/arrangement and author-generated tables. It does not
relicense source text or annotations. IBM retains CC BY-SA 3.0; hate-corpus
components retain CC BY 4.0, CC0 or MIT as documented in… See the full description on the dataset page: https://huggingface.co/datasets/nishan-chatterjee/llm-bias-detection.fire-smoke-detection-corpus-v1
FireViewer Fire/Smoke Detection Corpus v1
Status
Active strict-clean detection corpus. Current catalogue state: 102,257 rows, split 60,981 train / 19,209 validation / 22,067 test.
The corpus stores source-specific provenance, hashes, grouping/de-duplication information, validation status and annotation metadata. It is the current training reference for the strict FireViewer detector releases.
Rights
There is no single licence covering all source… See the full description on the dataset page: https://huggingface.co/datasets/fireviewer/fire-smoke-detection-corpus-v1.upper-level-trough-ridge-detection-data
Detection of Upper-Level Troughs and Ridges Using Deep Learning — Dataset
Application in the Mediterranean
Expert annotations, model-ready ERA5 fields, and the derived climatology accompanying Detection of Upper-Level Troughs and Ridges Using Deep Learning – Application in the Mediterranean by Ofir Ariel, Omer Sela, Hadas Saaroni, and Baruch Ziv.
Project resources
Resource
Link
Paper
EGUsphere preprint
Interactive project page… See the full description on the dataset page: https://huggingface.co/datasets/Omer-Sela/upper-level-trough-ridge-detection-data.fashionpedia
Dataset Card for Fashionpedia
Dataset Summary
Fashionpedia is a dataset mapping out the visual aspects of the fashion world.
From the paper:
Fashionpedia is a new dataset which consists of two parts: (1) an ontology built by fashion experts containing 27 main apparel categories, 19 apparel parts, 294 fine-grained attributes and their relationships; (2) a dataset with everyday and celebrity event fashion images annotated with segmentation masks and their associated… See the full description on the dataset page: https://huggingface.co/datasets/detection-datasets/fashionpedia.code_x_glue_cc_clone_detection_big_clone_bench
Dataset Card for "code_x_glue_cc_clone_detection_big_clone_bench"
Dataset Summary
CodeXGLUE Clone-detection-BigCloneBench dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/Clone-detection-BigCloneBench
Given two codes as the input, the task is to do binary classification (0/1), where 1 stands for semantic equivalence and 0 for others. Models are evaluated by F1 score.
The dataset we use is BigCloneBench and filtered following the paper… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_clone_detection_big_clone_bench.language-detection
Dataset Card for "language-detection"
More Information needed
detect-waste
Dataset Card for detect-waste
Dataset Summary
AI4Good project for detecting waste in environment. www.detectwaste.ml.
Our latest results were published in Waste Management journal in article titled Deep learning-based waste detection in natural and urban environments.
You can find more technical details in our technical report Waste detection in Pomerania: non-profit project for detecting waste in environment.
Did you know that we produce 300 million tons of plastic every… See the full description on the dataset page: https://huggingface.co/datasets/Yorai/detect-waste.date_fruit_maturity_detection
Date Fruit Maturity Detection
A dataset for detection of Date Fruit Maturity. The dataset contains 9,010 images with 15,238 bounding box annotations across 4 categories.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{zarouit2024date,
title={Date fruit detection dataset for automatic harvesting},
author={Zarouit, Yousra and Zekkouri, Hassan and Ouhda, Mohamed and Aksasse, Brahim},
journal={Data… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/date_fruit_maturity_detection.tweets_hate_speech_detection
Dataset Card for Tweets Hate Speech Detection
Dataset Summary
The objective of this task is to detect hate speech in tweets. For the sake of simplicity, we say a tweet contains hate speech if it has a racist or sexist sentiment associated with it. So, the task is to classify racist or sexist tweets from other tweets.
Formally, given a training sample of tweets and labels, where label ‘1’ denotes the tweet is racist/sexist and label ‘0’ denotes the tweet is not… See the full description on the dataset page: https://huggingface.co/datasets/tweets-hate-speech-detection/tweets_hate_speech_detection.anomaly_detection_cmsl1t
Trigger Anomaly Detection for New Physics at the Large Hadron Collider
This dataset is a mirror of the Zenodo record: https://doi.org/10.5281/zenodo.21787779
This dataset contains Level-1 Trigger objects from the CMS experiment at the CERN Large
Hadron Collider, assembled for research on unsupervised anomaly detection in the
trigger. The goal of unsupervised anomaly detection in this context is the discovery of
new physics. This data set does not contain new physics. It is meant… See the full description on the dataset page: https://huggingface.co/datasets/CERN/anomaly_detection_cmsl1t.willow-surface-code-detection-events
Willow Surface-Code Detection Events (ingested)
Detection events and logical-observable flips derived from Google's Willow
below-threshold surface-code dataset (Zenodo
10.5281/zenodo.13273331), rotated
surface code at distances 3, 5, 7 in X and Z memory.
Each row is one experimental shot. Detection events are derived from the raw
device measurement records with Stim's measurement-to-detector converter, using
the per-shot sweep bits, and validated to reproduce the dataset's… See the full description on the dataset page: https://huggingface.co/datasets/ShayManor/willow-surface-code-detection-events.government-ai-detection
Government AI Text Detection — Results
Completed AI-detection output from the government-ai
pipeline, which measures the prevalence of AI-generated/edited text across
four kinds of US government media, 2000–2026:
source
what
bills
Congressional bill text (as-introduced versions), from govinfo
speeches
Floor speeches + Extensions of Remarks from the Congressional Record
comments
Public comments on regulations.gov (via the Mirrulations mirror)
documents
The… See the full description on the dataset page: https://huggingface.co/datasets/ksasse/government-ai-detection.NJU-HARD-Detection
NJU-HARD-Detection
Object Detection in 122 MP Wide-Area UAV Imagery
🤗 Hugging Face · 🟣 ModelScope · 📊 Statistics: HF / MS
English | 中文: Hugging Face · ModelScope
🌍 Overview
NJU-HARD-Detection provides the object-detection release of HARD, with full-resolution frames and per-frame pedestrian and vehicle boxes. Detection training and evaluation use the categories and boxes; the retained source identity fields are… See the full description on the dataset page: https://huggingface.co/datasets/RL-MIND/NJU-HARD-Detection.language_detection_trainmaize_weed_detection
Maize Weed Detection
A dataset for detection of maize and weeds. The dataset contains 500 images with 5,764 bounding box annotations across 2 categories.
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{olaniyi2023development,
title={Development of maize plant dataset for intelligent recognition and weed control},
author={Olaniyi, Olayemi Mikail and Salaudeen, Muhammadu Tajudeen and Daniya… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/maize_weed_detection.hyperpartisan_news_detection_bypublisher_promptsourceai-text-detection-pile
Dataset Card for AI Text Dectection Pile
Dataset Summary
This is a large scale dataset intended for AI Text Detection tasks, geared toward long-form text and essays. It contains samples of both human text and AI-generated text from GPT2, GPT3, ChatGPT, GPTJ.
Here is the (tentative) breakdown:
Human Text
Dataset
Num Samples
Link
Reddit WritingPromps
570k
Link
OpenAI Webtext
260k
Link
HC3 (Human Responses)
58k
Link
ivypanda-essays
TODO
TODO… See the full description on the dataset page: https://huggingface.co/datasets/artem9k/ai-text-detection-pile.
