datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hatespeech_detection
HaSpeeDe2
The HaSpeeDe2 dataset collects 8,012 tweets and 500 news headlines annotated for the presence of hate speech, stereotypes and nominal utterance.
The dataset has been used in the context of the HaSpeeDe task (http://www.di.unito.it/~tutreeb/haspeede-evalita20/index.html), organized as part of the EVALITA 2020 evaluation campaign (http://www.evalita.it/2020).
In order to meet the GDPR requirements, texts have been pseudonymized replacing all original IDs in both datasets… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/hatespeech_detection.gesture_detectionscientific-exaggeration-detection
Dataset Card for Scientific Exaggeration Detection
Dataset Summary
Public trust in science depends on honest and factual communication of scientific papers. However, recent studies have demonstrated a tendency of news media to misrepresent scientific papers by exaggerating their findings. Given this, we present a formalization of and study into the problem of exaggeration detection in science communication. While there are an abundance of scientific papers and popular… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/scientific-exaggeration-detection.reasoning-drift-onset-detection-v0.1
Important Evaluation Limitation
Version 0.1 uses a highly regular trajectory structure in which the first drift step is frequently located at Step 4 and visible failure commonly appears at Step 5.
This creates a positional shortcut: a model may achieve inflated onset-detection performance by learning the dataset construction pattern rather than analysing the reasoning trajectory.
Version 0.1 should therefore be treated as a task-definition and scorer-validation release, not as a… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/reasoning-drift-onset-detection-v0.1.birds-object-detection-1600x896
Birds Object Detection 1600x896
A single-class bird object-detection dataset prepared for Ultralytics at a
1600 x 896 camera resolution. The dataset uses a deterministic 90% training
and 10% validation split.
Contents
data/yolo_birds_1600x896.tar.gz: complete Ultralytics dataset archive.
manifest.json: archive size, SHA-256 checksum, and dataset metadata.
After extraction, the archive contains data.yaml and matching image and YOLO
label trees:
data.yaml… See the full description on the dataset page: https://huggingface.co/datasets/PBatch23888/birds-object-detection-1600x896.satellite-water-detectionucf-anomaly-detection-mapped
UCF-Crime: Precomputed I3D Features with Temporal Annotations
This dataset provides pre-extracted 1024-dimensional I3D RGB features along with frame-level temporal anomaly labels for videos from the UCF-Crime dataset.
Dataset Characteristics
Features
1024-dimensional I3D RGB feature vectors
Extracted from 64 uniformly sampled frames per video
Feature tensor shape: [64, 1024]
Temporal Annotations
Mapped from original anomaly intervals… See the full description on the dataset page: https://huggingface.co/datasets/Rahima411/ucf-anomaly-detection-mapped.han-humanoid-vision-object-detection-metrics-v1
Humanoid Vision Detection Metrics Dataset
Overview
Dataset performa sistem visi humanoid saat mendeteksi objek.
Features
image_brightness_level
object_distance_cm
detection_confidence_score
camera_noise_index
frame_processing_time_ms
Target
detection_accuracy_percent
Football-detectionomission-detection-realdata-pilot
Omission Detection — Real-Data Pilot
What is Omission Detection?
Large language models (LLMs) in agentic pipelines often omit information
present in their context window — they fail to surface a relevant fact even
when it is theoretically visible. This dataset captures 372 controlled
trials from the real-data pilot, extending the synthetic sweep to
real-world documents and agent frameworks.
Each trial fetches a document from a real source (PubMed, HAPI FHIR, SEC… See the full description on the dataset page: https://huggingface.co/datasets/Santhiyarajan/omission-detection-realdata-pilot.deckergui-ui-detections
UI element detection dataset for YOLO-based visual perception in DeckerGUI. Contains annotated screenshots with bounding boxes for buttons, inputs, menus, sidebars, and headers.
Dataset Details
Repository: ctaxnagomi/deckergui-ui-detections
License: MIT
DeckerGUI Version: v2.0.0
Created: 2026-08-17
Dataset Schema
See metadata.json for the full schema definition.
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ctaxnagomi/deckergui-ui-detections.omission-detection-synthetic
Omission Detection — Synthetic Sweep
What is Omission Detection?
Large language models (LLMs) in agentic pipelines often omit information
present in their context window — they fail to surface a relevant fact even
when it is theoretically visible. This dataset captures 75,876 controlled
trials designed to measure and attribute these omissions across 9 taxonomic
layers (L0–L8).
Each trial generates a synthetic clinical document, embeds a "needle" fact at a… See the full description on the dataset page: https://huggingface.co/datasets/Santhiyarajan/omission-detection-synthetic.CouplingSmells-Detection-Data261_DetectionAI_Gemini2.5_datamultiturn-injection-detection
Multi-Turn Distributed Prompt Injection Detection Dataset
Dataset Description
27,180 synthetic multi-turn conversations (18,754 train / 3,296 val / 5,130 test) designed for training and evaluating temporal prompt injection detectors. Each conversation consists of 6-9 user turns with assistant responses.
Shared-Prefix Design
Every attack conversation is paired with a benign conversation that shares identical opening turns. A conversational prefix of 3-5 user… See the full description on the dataset page: https://huggingface.co/datasets/rockCO78/multiturn-injection-detection.aml_fraud_detection_and_reasoning_50k_syntetichallucination_detection_transformers
ToolACE Hallucination Dataset
This dataset was generated for the assignment Hallucination Detection in Tool Calling.
It is based on ToolACE tool-calling dialogues and uses a RAGTruth-style schema:
query: user question
context: tool output / grounding evidence
output: assistant final answer
hallucination_labels: character-level hallucination spans
Files:
File
Rows
Description
toolace_clean_ragtruth.jsonl
1347
clean ToolACE tool-use answers… See the full description on the dataset page: https://huggingface.co/datasets/HASSANI8046/hallucination_detection_transformers.bug_detection_dataset261_DetectionAI_GPT_data
Dataset README
Overview
This dataset contains training, validation and test data withh paired text samples for human vs. AI-generated text classification. The prompt we use to generate AI-generated samples is "Please rewrite the following text and only return me the rewritten text: text". Each pair shares the same id:
label = 0 → human-written text
label = 1 → AI-generated rewrite (e.g., from GPT-4 models)
Data Format
Each entry has the following… See the full description on the dataset page: https://huggingface.co/datasets/Clement1290/261_DetectionAI_GPT_data.uav_drone_detections_final
