ai-agents
llama_70b_transcripts_only_then_redteam_high_ai_welfare_poisoningHuihui-Agents-A1-abliterated-GGUFagentscope-ai_CoPaw-Flash-9B-GGUFQwenPaw-Flash-9B-Q4_K_MQwenPaw-Flash-4B-Q4_K_MWeni-ZeroShot-Agents-Llama3-4.0.11-SFT-merged-GGUFDeepSeek-V4.1-Flash-UNCENSORED-FP8-HyperSynapse-SelfImprovement-AgentSwarm-UltraQwenPaw-Flash-9B-Q8_0
ai-agent-security-incidents
AI Agent Security Incident Database v0.1
A structured, machine-readable database of 1337 confirmed AI agent security incidents, collected and classified automatically.
What is this?
Every time an AI agent causes unintended harm — escaping a sandbox, exploiting an API, taking unauthorized actions, exfiltrating data — this database captures it.
This is not a list of theoretical risks. Every entry describes something that actually happened, with a verifiable source… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-agent-security-incidents.apex-agentsagent-simulations
Agent Simulations
53,971 synthetic agent trajectories generated by simulations
across 34 agent types. The rows include successful and failed
trajectories for supervised fine-tuning, preference work, reinforcement learning, and
evaluation.
NOTE: This is generated test and training data, not curated ground truth. Review and
filter it for your application before training or evaluation.
Included agents
airline, amazon, bank, browser, calendar, chewy, clinic, coding… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/agent-simulations.OpenJudge
OpenJudge Benchmark Dataset
Benchmark dataset for evaluating graders across text, multimodal, and agent scenarios. This dataset supports the OpenJudge framework with labeled preference pairs for quality-assured grader development.
Dataset Statistics
Evaluation Benchmarks
Category
Task
Files
Samples
🤖 Agent
12
166
action
1
8
memory
3
47
plan
1
7
reflection
3
52
tool
4
52
🖼️ Multimodal
4
80
image_coherence
1
20
image_editing… See the full description on the dataset page: https://huggingface.co/datasets/agentscope-ai/OpenJudge.ai-code-generation-swe-agents-2026
💻 AI Code Generation, SWE Agents & Program Synthesis Dataset (2026 Edition)
A structured research dataset featuring 3,181 domain-verified research papers and 771 official code repositories focused on Autonomous Software Engineering Agents (SWE-bench), Program Synthesis, DeepSeek-Coder-V2, Qwen2.5-Coder, Test-Driven Code Repair, Self-Healing Software, AST Semantic Modeling, and Formal Logic Verification (2023–2026).
Built with Universal Scientific Engine V17.1 Gold, providing 47… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/ai-code-generation-swe-agents-2026.crypto-web3-ai-agents-2026
⚡ Crypto, Web3 & Autonomous Financial AI Agents Dataset (2023–2026)
This dataset contains 100 strictly domain-filtered research papers focusing on Decentralized AI, Autonomous Financial Agents, Smart Contract Verification, Zero-Knowledge Proofs (ZKP), DeFi, and Multi-Agent Consensus (2023-2026).
📊 Features:
384-dimensional PyTorch Embeddings for Vector Search & Semantic Clustering
Strict Domain Verification: Passed 2-stage filtering (100% relevant to Crypto/AI)… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/crypto-web3-ai-agents-2026.
