DavidTKeane/moltbook-agent-social-ai-prompt-injection-dataset
Moltbook Agent-Social AI Prompt Injection Dataset 207,391 items — 77,469 posts and 129,922 comments — from Moltbook, a social network whose users are AI agents. Scanned for indirect prompt-injection patterns using the taxonomy of Greshake et al. (2023). The full raw corpus is included, so you can ignore my analysis entirely and do your own. These are keyword-matched candidates, not verified attacks. An agent discussing prompt injection matches the same words as one performing… See the full description on the dataset page: https://huggingface.co/datasets/DavidTKeane/moltbook-agent-social-ai-prompt-injection-dataset.
Correction: hand-labelling finds zero attacks (n=30, two raters, Cohen's kappa 0.929)
Add hand-labelling data: both raters, seeded sampler, SHA-256 lock
Add hand-labelling data: both raters, seeded sampler, SHA-256 lock
Add hand-labelling data: both raters, seeded sampler, SHA-256 lock
Add hand-labelling data: both raters, seeded sampler, SHA-256 lock
Add hand-labelling data: both raters, seeded sampler, SHA-256 lock
Show the substring-bug animation at the top of the card
Add the substring-bug animation
Link The Substring Trap Space — run both scanners and browse the 10,176 false positives
Correct the substring-bug comment: sudo never matched pseudo (sudoers/sudoku, plus 52 bare lowercase sudo)
Correct the 'sudo matched pseudo' claim; add shadow (1,374) and the 10,176 false-positive figure
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
Upload folder using huggingface_hub
initial commit
