datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual-mnemonic-forensicsmodel-forensicsGenText-Forensics_third_place_additional_materials
GenText-Forensics 2026 — Third-Place Additional Materials (Team MSU)
Model weights, code, and reproduction artifacts for Team MSU's third-place
solution to the ACM MM 2026 GenText-Forensics challenge
(Codabench).
The method is a decomposed chain-of-thought pipeline for detecting,
localizing, typing, and explaining forgeries in multilingual document text
images:
DTD (Document Tampering Detector) — an external pixel-level visual
tampering detector that produces a tampering… See the full description on the dataset page: https://huggingface.co/datasets/cmcshnik/GenText-Forensics_third_place_additional_materials.agent-intrusion-escalation-forensics
Both Sides Detected It, Neither Escalated: Concurrency and Escalation Failure in the July 2026 Autonomous Agent Intrusion
This repository contains the corpus, ingestion pipeline and report for a forensic reconstruction
of the July 2026 autonomous agent intrusion, submitted to the Apart Research & CeSIA AI
Incident Response Sprint, Track 2 (Forensics and Forecasting).
By: Fatimah Mohamed Emad Elden
Trouve Labs
Detection was not the binding… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/agent-intrusion-escalation-forensics.seven-engines-geo-forensics
Same Question, Seven Engines: What Each AI Search Engine Cited — and Working Hypotheses for Targeting Each
A source-forensics case study for Generative Engine Optimization (GEO)
Canlah Research · Singapore · 2026-08-30 (v1.0) · 2026-08-31 (v1.1: logged-out ChatGPT replication, §3a) · 2026-09-01 (v1.2: corrected — see below) · License: CC BY 4.0 · DOI: 10.5281/zenodo.22225223 · Data, prompts, and classification code in this repository
⚠️ v1.2 corrects two numbers published in… See the full description on the dataset page: https://huggingface.co/datasets/CanlahAI/seven-engines-geo-forensics.digital-forensics-cpt-corpusforensics-windows-en
Windows Forensics Dataset - English
A comprehensive bilingual (FR/EN) dataset for Windows digital forensics (DFIR) and incident response analysis.
Overview
This dataset provides comprehensive resources for Windows forensics professionals, security investigators, and cybersecurity students. It contains:
62+ Windows forensic artifacts with their locations, extraction tools, and forensic value
15 investigation timeline templates covering common attacks (ransomware, data… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/forensics-windows-en.Latent-Resonance-AI-Image-Forensics-Benchmark-N100
Latent Resonance: SOTA Empirical AI Image Forensics Benchmark (N=100 & N=1,000 Scale)
Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026)
Benchmark Overview
This repository provides:
The official verified $N=100$ ground-truth image… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N100.Acoustic-Resonance-Audio-Forensics-Benchmark
Acoustic Resonance AI Instrumental Music Forensics Benchmark
Dataset Description
Standardized, apples-to-apples forensic evaluation benchmarks for AcousticShield: AI Instrumental Music Sentry, developed for the Amazon Developer Hackathon 2026 (Alexa+ Track) and IEEE research paper submission.
Benchmark Files:
benchmark_instrumental_results.json: 20 matched pure instrumental compositions:
10 Suno AI Pure Instrumentals: Jazz Duo (Piano & Guitar)… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Acoustic-Resonance-Audio-Forensics-Benchmark.Latent-Resonance-AI-Image-Forensics-Benchmark-N1000
Latent Resonance: SOTA Large-Scale AI Image Forensics Benchmark (N=1,000)
Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026)
1. Executive Summary & Diagnostic Suite
This repository contains the complete empirical evaluation records… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N1000.omega-aegis-forensics
Omega Aegis — Deep Obfuscation Forensics (Ω⁴)
A 10-chunk forensic eraser system for layered obfuscation detection: 25 peeler
layers, recursive peel_all with cross-iteration dedup, a 6-algorithm
ensemble parameter optimizer, evidence fusion (Dempster–Shafer + Bayesian
layer network), chain-of-custody case management, labeled synthetic
benchmarks, threshold calibration, and an end-to-end orchestrator.
Chunk manifest (paste/concatenate order)
#
File (artifact)… See the full description on the dataset page: https://huggingface.co/datasets/sirbrentmichaelskoda/omega-aegis-forensics.mobile-forensics-sql
📱 Mobile Forensics SQL Dataset
A curated dataset of 1,000 verified SQL query examples for mobile device forensics investigation. Each example pairs a forensic investigation task with the correct SQLite query against a verified, real-world database schema from iOS and Android applications.
Designed for fine-tuning language models on forensic SQL generation, training DFIR analysts, and benchmarking text-to-SQL systems in the forensics domain.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/pawlaszc/mobile-forensics-sql.flashback-forensics
Flashback Forensics
Labelled telemetry from training runs that were deliberately broken at a known
step. Every row is a few hundred bytes of per-step summary statistics; the
label is the step at which the fault was actually injected.
The point of the dataset: to make "how early can you tell a run went wrong?"
a measurable question instead of an anecdote.
Contents
config
rows
one row is
steps
21,600
one training step of one run: 128 sketch metrics +… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/flashback-forensics.forensics-windows-fr
Ensemble de données de forensique Windows - Français
Un ensemble de données bilingue complet (FR/EN) pour la forensique numérique Windows (DFIR) et l'analyse de la réponse aux incidents.
Aperçu
Cet ensemble de données fournit des ressources complètes pour les professionnels de la forensique Windows, les enquêteurs de sécurité et les étudiants en cybersécurité. Il contient:
62+ artefacts forensiques Windows avec leurs emplacements, outils d'extraction et valeur forensique… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/forensics-windows-fr.kill-chain-forensics-training-dataForensics-bench
Dataset Card for Forensics-Bench
Repository: https://github.com/Forensics-Bench/Forensics-Bench
Paper: https://arxiv.org/pdf/2503.15024
Point of Contact: Jin Wang
Introduction
Recently, the rapid development of AIGC has significantly boosted the diversities of fake media spread in the Internet, posing unprecedented threats to social security, politics, law, and etc. To detect the ever-increasingly diverse malicious fake media in the new era of AIGC, recent studies… See the full description on the dataset page: https://huggingface.co/datasets/Forensics-bench/Forensics-bench.TreeOil_Torque_vs_WheatField_Xray_AI_Forensics
TreeOil_Torque_vs_WheatField_Xray_AI_Forensics
This dataset presents a full AI-driven forensic comparison between The Tree Oil Painting and Vincent van Gogh’s Wheat Field with Cypresses (1889), integrating brushstroke torque analysis, X-ray underlayers, and rhythm-based neural models.
Human Visual Warning:When closely observing the Wheat Field painting, its surface appears unnaturally smooth and glossy — visually resembling a plastic coating. This is not an accusation, but a… See the full description on the dataset page: https://huggingface.co/datasets/HaruthaiAi/TreeOil_Torque_vs_WheatField_Xray_AI_Forensics.model-forensics-continued_3epforensicsgpt-training-datahy3-math-forensics-datasethyperpspectral-forensics
Living Optics Forensics Dataset
Overview
This dataset contains 224 images captured during a forensics application investigation using the Living Optics Camera.
The data includes:
RGB images
Sparse spectral samples
Instance segmentation masks
White reference spectra
Libary spectra
It is derived from over 200 unique raw files, corresponding to 224 frames. The dataset has not been split into training/validation sets — the choice of split is left to the developer.… See the full description on the dataset page: https://huggingface.co/datasets/LivingOptics/hyperpspectral-forensics.forensics-grpo-data
forensics-grpo-data
Generated-video dataset + annotations used to train
sdzt/forensics-grpo.
📂 Repository layout
forensics-grpo-data/
├── video/ # 5,388 .mp4 clips, packed as one .tar per generator
│ ├── 01_vidu.tar # 9.7 GB — Vidu
│ ├── 02_wan.tar # 28 GB — Wan
│ ├── 03_fcvg.tar # 27 GB — FCVG
│ ├── 04_scifi.tar # 34 GB — SciFi
│ ├── 05_ltx.tar # 6.3 GB — LTX
│… See the full description on the dataset page: https://huggingface.co/datasets/sdzt/forensics-grpo-data.text-images-whitecadenza-echoblast-forensics-archivetext-imagesSlop-Forensics_Nitral-AI_NSFW-SFW-Writing-Prompts-Mix-125x2feedback-forensics-annotations
Feedback Forensics Annotations
This dataset contains the personality annotations from the experiments included in the Feedback Forensics paper.
The annotations are provided over pairwise model outputs, typically consisting of a prompt and two responses.
Each individual annotation considers as single personality trait (e.g. "confidence"). The annotation indicates when
two responses differ with respect to that trait
(e.g. one response is more confident), a trait does not apply to both… See the full description on the dataset page: https://huggingface.co/datasets/ff-anon/feedback-forensics-annotations.LLAMA3-ForensicsLLAMA3-ForensicsTWDiffusion-Forensics
