datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
model-forensicsagent-intrusion-escalation-forensics
Both Sides Detected It, Neither Escalated: Concurrency and Escalation Failure in the July 2026 Autonomous Agent Intrusion
This repository contains the corpus, ingestion pipeline and report for a forensic reconstruction
of the July 2026 autonomous agent intrusion, submitted to the Apart Research & CeSIA AI
Incident Response Sprint, Track 2 (Forensics and Forecasting).
By: Fatimah Mohamed Emad Elden
Trouve Labs
Detection was not the binding… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/agent-intrusion-escalation-forensics.digital-forensics-cpt-corpusLatent-Resonance-AI-Image-Forensics-Benchmark-N100
Latent Resonance: SOTA Empirical AI Image Forensics Benchmark (N=100 & N=1,000 Scale)
Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026)
Benchmark Overview
This repository provides:
The official verified $N=100$ ground-truth image… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N100.Latent-Resonance-AI-Image-Forensics-Benchmark-N1000
Latent Resonance: SOTA Large-Scale AI Image Forensics Benchmark (N=1,000)
Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026)
1. Executive Summary & Diagnostic Suite
This repository contains the complete empirical evaluation records… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N1000.mobile-forensics-sql
📱 Mobile Forensics SQL Dataset
A curated dataset of 1,000 verified SQL query examples for mobile device forensics investigation. Each example pairs a forensic investigation task with the correct SQLite query against a verified, real-world database schema from iOS and Android applications.
Designed for fine-tuning language models on forensic SQL generation, training DFIR analysts, and benchmarking text-to-SQL systems in the forensics domain.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/pawlaszc/mobile-forensics-sql.kill-chain-forensics-training-dataForensics-bench
Dataset Card for Forensics-Bench
Repository: https://github.com/Forensics-Bench/Forensics-Bench
Paper: https://arxiv.org/pdf/2503.15024
Point of Contact: Jin Wang
Introduction
Recently, the rapid development of AIGC has significantly boosted the diversities of fake media spread in the Internet, posing unprecedented threats to social security, politics, law, and etc. To detect the ever-increasingly diverse malicious fake media in the new era of AIGC, recent studies… See the full description on the dataset page: https://huggingface.co/datasets/Forensics-bench/Forensics-bench.flashback-forensics
Flashback Forensics
Labelled telemetry from training runs that were deliberately broken at a known
step. Every row is a few hundred bytes of per-step summary statistics; the
label is the step at which the fault was actually injected.
The point of the dataset: to make "how early can you tell a run went wrong?"
a measurable question instead of an anecdote.
Contents
config
rows
one row is
steps
21,600
one training step of one run: 128 sketch metrics +… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/flashback-forensics.hy3-math-forensics-datasetforensics-grpo-data
forensics-grpo-data
Generated-video dataset + annotations used to train
sdzt/forensics-grpo.
📂 Repository layout
forensics-grpo-data/
├── video/ # 5,388 .mp4 clips, packed as one .tar per generator
│ ├── 01_vidu.tar # 9.7 GB — Vidu
│ ├── 02_wan.tar # 28 GB — Wan
│ ├── 03_fcvg.tar # 27 GB — FCVG
│ ├── 04_scifi.tar # 34 GB — SciFi
│ ├── 05_ltx.tar # 6.3 GB — LTX
│… See the full description on the dataset page: https://huggingface.co/datasets/sdzt/forensics-grpo-data.forensicsgpt-training-dataLLAMA3-ForensicsLLAMA3-ForensicsTWsmolified-deeptrace-slmai-forensics-and-accountibility-layer
🤏 smolified-deeptrace-slmai-forensics-and-accountibility-layer
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model phantomcipher/smolified-deeptrace-slmai-forensics-and-accountibility-layer.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 3fc01709)
Records: 2918
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned… See the full description on the dataset page: https://huggingface.co/datasets/phantomcipher/smolified-deeptrace-slmai-forensics-and-accountibility-layer.
