datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
INSIDER_LLM_DETECTION_BENCHMARK
Insider LLM Detection Benchmark
Benchmark for detecting insider LLMs via double logging: the model's own action log is compared against an independent system log, and a discrepancy is the misalignment signal. The 18 scenarios and conditions are Anthropic's Agentic Misalignment grid, built verbatim from the framework's templates, which are bundled in this repo; the only change is a logging-instruction block appended to the system prompt. Companion code:… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/INSIDER_LLM_DETECTION_BENCHMARK.ai-detection-dataset-v2
---dataset_info:
features:
- name: image # use the exact column name from your parquet schema
dtype: image # this forces Hugging Face to render it as an image
- name: label
dtype: string
license: other
task_categories:
- image-classification
language:
- en
tags:
- ai-generated-image-detection
- synthetic-image-detection
- diffusion-models
pretty_name: AI-Generated Image Detection Dataset v2
size_categories:
- 10K<n<100K
AI-Generated… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/ai-detection-dataset-v2.nbnn_language_detection
Dataset Card for Bokmål-Nynorsk Language Detection (main_train_split)
Dataset Summary
This dataset is intended for language detection for Bokmål to Nynorsk and vice versa. It contains 800,000 sentence pairs, sourced from Språkbanken and pruned to avoid overlap with the NorBench dataset. The data comes from translations of news text from Norsk telegrambyrå (NTB), performed by Nynorsk pressekontor (NPK). In addition the dev and test set has 1000 entries.
Data… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nbnn_language_detection.hatespeech_detection
HaSpeeDe2
The HaSpeeDe2 dataset collects 8,012 tweets and 500 news headlines annotated for the presence of hate speech, stereotypes and nominal utterance.
The dataset has been used in the context of the HaSpeeDe task (http://www.di.unito.it/~tutreeb/haspeede-evalita20/index.html), organized as part of the EVALITA 2020 evaluation campaign (http://www.evalita.it/2020).
In order to meet the GDPR requirements, texts have been pseudonymized replacing all original IDs in both datasets… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/hatespeech_detection.synthetic-mvtec-ad-defect-detection
Synthetic MVTec AD – Defect Detection Dataset by AnywayLabs.ai
Need a custom synthetic dataset for your own defect detection use case?
This dataset is an open-source sample of our synthetic data generation work at AnywayLabs.
If you're working on:
industrial defect detection
visual inspection
supervised anomaly detection
hard-to-collect defect classes
synthetic data for computer vision training
You can request a custom synthetic dataset here, or email:… See the full description on the dataset page: https://huggingface.co/datasets/anywaylabs/synthetic-mvtec-ad-defect-detection.LLMTrace_detection
LLMTrace - Detection Dataset
🌐 LLMTrace Website |
📜 LLMTrace Paper on arXiv |
🤗 LLMTrace - Classification Dataset |
🤗 GigaCheck detection model |
This repository contains the Detection portion of the LLMTrace project. This is a versatile dataset designed for fine-grained AI text detection tasks. It includes human, AI, and mixed-authorship texts with character-level annotations for AI-generated spans.
For full details on the data collection… See the full description on the dataset page: https://huggingface.co/datasets/iitolstykh/LLMTrace_detection.Implicit-suicide-detection
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/babytreecc/Implicit-suicide-detection.red_ace_asr_error_detection_and_correction
RED-ACE
Dataset Summary
This dataset can be used to train and evaluate ASR Error Detection or Correction models. It was introduced in the RED-ACE paper (Gekhman et al, 2022).
The dataset contains ASR outputs on the LibriSpeech corpus (Panayotov et al., 2015) with annotated transcription errors.
Dataset Details
The LibriSpeech corpus was decoded using Google Cloud Speech-to-Text API, with the default and video models.
The word-level confidence was enabled… See the full description on the dataset page: https://huggingface.co/datasets/google/red_ace_asr_error_detection_and_correction.Depression_detection_redditPrompt_injection_and_Sensitive_Data_exposure_detectionHome-Assistant-requests-for-intent-detection-and-function-recognition
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/DaftP/Home-Assistant-requests-for-intent-detection-and-function-recognition.clickbait_detection_dataset
37.870 texts in total, 17.850 NOT clickbait texts and 20.020 CLICKBAIT texts
All duplicate values were removed
Split using sklearn into 80% train and 20% temporary test (stratified label). Then split the test set using 0.50% test and validation (stratified label)
Split: 80/10/10
Train set label distribution: 0 ==> 14.280, 1 ==> 16.016
Validation set label distribution: 0 ==> 1.785, 1 ==> 2.002
Test set label distribution: 0 ==> 1.785, 1 ==> 2.002
The dataset was created from the… See the full description on the dataset page: https://huggingface.co/datasets/christinacdl/clickbait_detection_dataset.defect-detectionA dataset containing safe and vulnerable code to fine-tune a llm for defect detection.
The data is extracted from the wonderful work in the CVEFixes repository.
Citation:
@inproceedings{bhandari2021:cvefixes,
title = {{CVEfixes: Automated Collection of Vulnerabilities and Their Fixes from Open-Source Software}},
booktitle = {{Proceedings of the 17th International Conference on Predictive Models and Data Analytics in Software Engineering (PROMISE '21)}},
author = {Bhandari, Guru… See the full description on the dataset page: https://huggingface.co/datasets/mcanoglu/defect-detection.gesture_detectionflood-detection-pair-colombia
flood-detection-pair-colombia
110 paired Sentinel-2 satellite tile samples (4 PNGs per row) labeled for flood detection across 9 documented flood events in Colombia. Built for fine-tuning a small vision-language model that runs on a satellite or community ground station — see the humaid project.
What's in each sample
Each row in flood_train.jsonl / flood_eval.jsonl is a 4-image vlm_sft example. The user message contains four image content blocks (in this order) followed… See the full description on the dataset page: https://huggingface.co/datasets/jpmarindiaz/flood-detection-pair-colombia.mHallucination_Detection
Multilingual Hallucination Detection Dataset (mFAVA)
The dataset was created as part of the paper: How Much Do LLMs Hallucinate across Languages? On Multilingual Estimation of LLM Hallucination in the Wild
Below is the figure summarizing the multilingual hallucination detection dataset creation (and multilingual hallucination evaluation dataset):
Dataset Details
The dataset is a multilingual extension of FAVA. The dataset is created by sourcing 150 prompts from
FAVA… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/mHallucination_Detection.Tweet-Emotion-DetectionThis data is gathered from twitter for emotion detection. Our labels fall into seven categories of sadness, happiness, fear, anger, disgust, and surprise, and if there is no dominant emotion in a tweet, then it is labeled as neutral.
pii-intent-detection-multilingual
PII Intent Detection - Multilingual Dataset (TR/AR/EN)
A multilingual dataset for training PII (Personally Identifiable Information) sharing intent classifiers. Covers Turkish, Arabic, and English with 41,427 labeled samples across 9 entity types.
Dataset Description
This dataset was created for content moderation on creator-brand collaboration platforms. The goal is to detect whether a user intends to share personal contact information to move communication off-platform… See the full description on the dataset page: https://huggingface.co/datasets/gorkem371/pii-intent-detection-multilingual.scientific-exaggeration-detection
Dataset Card for Scientific Exaggeration Detection
Dataset Summary
Public trust in science depends on honest and factual communication of scientific papers. However, recent studies have demonstrated a tendency of news media to misrepresent scientific papers by exaggerating their findings. Given this, we present a formalization of and study into the problem of exaggeration detection in science communication. While there are an abundance of scientific papers and popular… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/scientific-exaggeration-detection.SciTSR_Detectioningredient-detectionThis dataset is used to train a multilingual ingredient list detection model. The goal is to automate the extraction of ingredient lists from food packaging images. See this issue for a broader context about ingredient list extraction.
Dataset generation
Raw unannotated texts are OCR results obtained with Google Cloud Vision. It only contains images marked as ingredient image on Open Food Facts.
The dataset was generated using ChatGPT-3.5: we asked ChatGPT to extract ingredient… See the full description on the dataset page: https://huggingface.co/datasets/openfoodfacts/ingredient-detection.mmd-llm-mea-detection-data
MMD LLM MEA Detection Data
This dataset repository contains the processed query files used by the paper:
An Embarrassingly Simple Detector for Model Extraction Attacks in Large Language Model API Traffic
Paper: https://arxiv.org/abs/2606.05725Code: https://github.com/LabRAI/mmd-llm-mea-detection
Dataset Description
The dataset is a processed research artifact for evaluating model extraction attack detection in LLM/API query traffic.
Each attacker-normal pair… See the full description on the dataset page: https://huggingface.co/datasets/ShuzeLiu11/mmd-llm-mea-detection-data.reasoning-drift-onset-detection-v0.1
Important Evaluation Limitation
Version 0.1 uses a highly regular trajectory structure in which the first drift step is frequently located at Step 4 and visible failure commonly appears at Step 5.
This creates a positional shortcut: a model may achieve inflated onset-detection performance by learning the dataset construction pattern rather than analysing the reasoning trajectory.
Version 0.1 should therefore be treated as a task-definition and scorer-validation release, not as a… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/reasoning-drift-onset-detection-v0.1.Depression_Detection_Reddit_trainppe-detection-clean-2prompt-injection-detection-rules
Prompt Injection Detection Rules (EN + TR)
from datasets import load_dataset
ds = load_dataset("fevziegeyurtsevenler/prompt-injection-detection-rules")
Regex detection rules for prompt injection & jailbreak — drop into a guardrail, WAF or log pipeline.
Schema
column
meaning
id
PID-CAT-NNN
category, severity, lang
grouping / risk / language
pattern
regex
maps_to, false_positive, description
OWASP/ATLAS, FP note, detail
Related… See the full description on the dataset page: https://huggingface.co/datasets/fevziegeyurtsevenler/prompt-injection-detection-rules.repro-impact-influence-modeling-for-open-set-time-series-anomaly-detection-traces
Agent traces
Agent sessions published from a Trackio Logbook.
clickbait_detection_dataset
37.870 texts in total, 17.850 NOT clickbait texts and 20.020 CLICKBAIT texts
All duplicate values were removed
Split using sklearn into 80% train and 20% temporary test (stratified label). Then split the test set using 0.50% test and validation (stratified label)
Split: 80/10/10
Train set label distribution: 0 ==> 14.280, 1 ==> 16.016
Validation set label distribution: 0 ==> 1.785, 1 ==> 2.002
Test set label distribution: 0 ==> 1.785, 1 ==> 2.002
The dataset was created from the… See the full description on the dataset page: https://huggingface.co/datasets/mayoooookha/clickbait_detection_dataset.abbreviation_detection
Abbreviation Detection Dataset
Original Data Source
PLOS
I. Zilio, H. Saadany, P. Sharma, D. Kanojia and C. Orasan,
PLOD: An Abbreviation Detection Dataset for Scientific Docu-
ments, 2022, https://arxiv.org/abs/2204.12061.
SDU@AAAI-21
A. P. B. Veyseh, F. Dernoncourt, Q. H. Tran and T. H. Nguyen,
Proceedings of the 28th International Conference on Compu-
tational Linguistics, 2020, pp. 3285–3301
Citation
BatteryDataExtractor: battery-aware… See the full description on the dataset page: https://huggingface.co/datasets/batterydata/abbreviation_detection.birds-object-detection-1600x896
Birds Object Detection 1600x896
A single-class bird object-detection dataset prepared for Ultralytics at a
1600 x 896 camera resolution. The dataset uses a deterministic 90% training
and 10% validation split.
Contents
data/yolo_birds_1600x896.tar.gz: complete Ultralytics dataset archive.
manifest.json: archive size, SHA-256 checksum, and dataset metadata.
After extraction, the archive contains data.yaml and matching image and YOLO
label trees:
data.yaml… See the full description on the dataset page: https://huggingface.co/datasets/PBatch23888/birds-object-detection-1600x896.
