CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01McGill-NLP /agent-reward-bench AgentRewardBench 💾Code 📄Paper 🌐Website 🤗Dataset 💻Demo 🏆Leaderboard AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor Loading dataset You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/agent-reward-bench.imagerobotics1K<n<10K4 likes21k downloads1y agoHugging Face02Lakera /b3-agent-security-benchmark-weak[paper] [blogpost] [game] b3 AI Security Benchmark: Breaking Agent Backbones Highly contextalized prompt injections crowd-sourced during the Gandalf Agent Breaker Challenge. This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents. The high quality dataset was used to evaluate the security of more than 30 LLMs. Dataset Summary Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.tabulartext-classificationn<1K6 likes920 downloads11mo agoHugging Face03vals-ai /finance_agent_benchmark Finance Agent Benchmark Dataset We present the Finance Agent Benchmark, featuring challenging and diverse real-world finance research problems which require LLMs to perform complex analysis with the use of of recent SEC filings. We construct the benchmark using a taxonomy of nine financial task categories, developed in consultation with experts from banks, hedge funds, and private equity firms. The dataset includes 537 expert-authored questions, covering tasks from information… See the full description on the dataset page: https://huggingface.co/datasets/vals-ai/finance_agent_benchmark.textn<1K9 likes676 downloads1y agoHugging Face04jamesdborin /Nemotron-SFT-Agentic-v2-prompt-only Nemotron-SFT-Agentic-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Agentic-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Agentic-v2-prompt-only.tabular100K<n<1M0 likes580 downloads3mo agoHugging Face05yuanyyaa /agent-reward-bench AgentRewardBench 💾Code 📄Paper 🌐Website 🤗Dataset 💻Demo 🏆Leaderboard AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor Loading dataset You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/yuanyyaa/agent-reward-bench.imagerobotics1K<n<10K0 likes568 downloads6mo agoHugging Face06gemmozero /ai-agent-security-incidents AI Agent Security Incident Database v0.1 A structured, machine-readable database of 1365 confirmed AI agent security incidents, collected and classified automatically. What is this? Every time an AI agent causes unintended harm — escaping a sandbox, exploiting an API, taking unauthorized actions, exfiltrating data — this database captures it. This is not a list of theoretical risks. Every entry describes something that actually happened, with a verifiable source… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-agent-security-incidents.tabulartext-classification1K<n<10K1 likes562 downloads2h agoHugging Face07Yunhao-Feng /AgentHazard AgentHazard A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents 🌐 Website | 📊 Dataset | 📄 Paper | 📖 Appendix 🎯 Overview AgentHazard is a comprehensive benchmark for evaluating harmful behavior in computer-use agents. Unlike traditional prompt-level safety benchmarks, AgentHazard focuses on execution-level failures that emerge through the composition of locally plausible steps across multi-turn, tool-mediated trajectories. Key Features… See the full description on the dataset page: https://huggingface.co/datasets/Yunhao-Feng/AgentHazard.tabular1K<n<10K0 likes362 downloads6mo agoHugging Face08BothBosu /multi-agent-scam-conversation Synthetic Multi-Turn Scam and Non-Scam Phone Conversation Dataset with Agentic Personalities Dataset Description The Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset with Agentic Personalities is an enhanced collection of simulated phone conversations between two AI agents, one acting as a scammer or non-scammer and the other as an innocent receiver. Each dialogue is labeled as either a scam or non-scam interaction. This dataset is designed to help develop… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/multi-agent-scam-conversation.texttext-classification1K<n<10K10 likes275 downloads2y agoHugging Face09cx-cmu /deepresearchgym-agentic-search-logs DeepResearchGym Agentic Search Logs This repository hosts the dataset accompanying the paper “Agentic Search in the Wild” (arXiv: https://arxiv.org/abs/2601.17617). The dataset contains 14M+ search queries collected via DeepResearchGym (DRGym), an open-source search API designed for DeepResearch-style agentic search. For more background on DRGym, see: https://arxiv.org/abs/2505.19253. All records have been anonymized and shuffled to prevent re-identification, and we additionally… See the full description on the dataset page: https://huggingface.co/datasets/cx-cmu/deepresearchgym-agentic-search-logs.tabulartext-retrieval10M<n<100M16 likes223 downloads8mo agoHugging Face10replit /agent-challenge Replit Agent Challenge For comprehensive details about the challenge, visit our GitHub repository. Dataset Overview This dataset comprises a collection of instructions and file states specifically curated for the agent challenge. It is derived from a subset of SWE-Bench-Lite. Schema Structure The dataset follows this schema: - File_before: [Initial state of the file] - Instructions: [Steps to transform the file to its final state] - File_after: [Resulting state… See the full description on the dataset page: https://huggingface.co/datasets/replit/agent-challenge.textn<1K3 likes205 downloads2y agoHugging Face11agenticx /DrugbankVocabularytext10K<n<100K0 likes175 downloads1y agoHugging Face12AdityaaXD /Multi-Agent_Reinforcement_Learning_Trading_System_Data 📊 Multi-Agent RL Trading System - Dataset This dataset contains historical OHLCV (Open, High, Low, Close, Volume) data for AAPL, MSFT, and GOOGL, pre-processed for Reinforcement Learning based trading systems. 📁 Dataset Content The dataset consists of CSV files downloaded via yfinance: AAPL.csv: Apple Inc. daily data (Jan 2018 - Dec 2024). MSFT.csv: Microsoft Corp. daily data (Jan 2018 - Dec 2024). GOOGL.csv: Alphabet Inc. daily data (Jan 2018 - Dec 2024). 📝… See the full description on the dataset page: https://huggingface.co/datasets/AdityaaXD/Multi-Agent_Reinforcement_Learning_Trading_System_Data.tabulartime-series-forecasting1K<n<10K4 likes157 downloads8mo agoHugging Face13sanjaydoss /Multi-Agent_Reinforcement_Learning_Trading_System_Data 📊 Multi-Agent RL Trading System - Dataset This dataset contains historical OHLCV (Open, High, Low, Close, Volume) data for AAPL, MSFT, and GOOGL, pre-processed for Reinforcement Learning based trading systems. 📁 Dataset Content The dataset consists of CSV files downloaded via yfinance: AAPL.csv: Apple Inc. daily data (Jan 2018 - Dec 2024). MSFT.csv: Microsoft Corp. daily data (Jan 2018 - Dec 2024). GOOGL.csv: Alphabet Inc. daily data (Jan 2018 - Dec 2024).… See the full description on the dataset page: https://huggingface.co/datasets/sanjaydoss/Multi-Agent_Reinforcement_Learning_Trading_System_Data.tabulartime-series-forecasting1K<n<10K8 likes154 downloads19d agoHugging Face14agentic-learning-ai-lab /daily-oracle Daily Oracle 📰 Project Website📝 Paper - Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle Daily Oracle is a continuous evaluation benchmark using automatically generated QA pairs from daily news to assess how the future prediction capabilities of LLMs evolve over time. Dataset Details Question Type: True/False (TF) & Multiple Choice (MC) Current Version* Time Span: 2020.01.01 - 2026.07.18 Size: 20,376 TF questions and 18,557 MC… See the full description on the dataset page: https://huggingface.co/datasets/agentic-learning-ai-lab/daily-oracle.textquestion-answering10K<n<100K4 likes151 downloads2mo agoHugging Face15witcheer /agentic-score-leaderboard 🛠️ Agentic Score Leaderboard — one RTX 5090 How well do local models actually drive a tool-using agent loop? Not single-call function-calling benchmarks — a real loop: native OpenAI tool-calling through llama-server, multi-step deterministic tasks, programmatic verification. Everything runs on a single RTX 5090 32GB. Updated 2026-06-17 · llama.cpp b9562 · --jinja native tool-calling · temp 0. Leaderboard # model params Agentic Score success tool-eff… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/agentic-score-leaderboard.tabularn<1K3 likes149 downloads3mo agoHugging Face16Koplos /finance_agent_benchmark Finance Agent Benchmark Dataset We present the Finance Agent Benchmark, featuring challenging and diverse real-world finance research problems which require LLMs to perform complex analysis with the use of of recent SEC filings. We construct the benchmark using a taxonomy of nine financial task categories, developed in consultation with experts from banks, hedge funds, and private equity firms. The dataset includes 537 expert-authored questions, covering tasks from information… See the full description on the dataset page: https://huggingface.co/datasets/Koplos/finance_agent_benchmark.textn<1K1 likes140 downloads2mo agoHugging Face17FatimahEmadEldin /agent-intrusion-escalation-forensics Both Sides Detected It, Neither Escalated: Concurrency and Escalation Failure in the July 2026 Autonomous Agent Intrusion This repository contains the corpus, ingestion pipeline and report for a forensic reconstruction of the July 2026 autonomous agent intrusion, submitted to the Apart Research & CeSIA AI Incident Response Sprint, Track 2 (Forensics and Forecasting). By: Fatimah Mohamed Emad Elden Trouve Labs Detection was not the binding… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/agent-intrusion-escalation-forensics.documentn<1K0 likes133 downloads8d agoHugging Face18invariantlabs /agent-ctf24-publictext10K<n<100K1 likes123 downloads2y agoHugging Face19Xieji-Li /Derm1M-AgentAuggated Derm1M-AgentAug Knowledge-enriched captions for 413,369 dermatological images, generated by MAGEN (Multi-Agent data GENeration) and used to pretrain O-MAKE. MAGEN rewrites part of the corpus through a foundation-model-assisted captioning agent with a diagnostic tool, verifying each result by retrieval; captions it did not improve on keep the original Derm1M text, and the agent_generated column records which is which. Every caption is additionally decomposed into distinct… See the full description on the dataset page: https://huggingface.co/datasets/Xieji-Li/Derm1M-AgentAug.textzero-shot-image-classification100K<n<1M4 likes109 downloads1mo agoHugging Face20BothBosu /single-agent-scam-conversations Synthetic Multi-Turn Scam and Non-Scam Phone Conversation Dataset Dataset Description The dataset is designed to help develop and evaluate models for detecting and classifying various types of phone-based scams. Dataset Structure The dataset consists of three columns: dialogue: The transcribed conversation between the caller and receiver. type: The specific type of scam or non-scam interaction. labels: A binary label indicating whether the conversation is a… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/single-agent-scam-conversations.texttext-classification1K<n<10K2 likes100 downloads2y agoHugging Face21AgenticFinLab /PortBench-Market PortBench Market Base Dataset Dataset Description A ten-year (Jan 2015–Dec 2025) daily financial dataset covering 183 instruments across six heterogeneous asset classes, designed for multi-asset portfolio management research and LLM evaluation. Asset Coverage Asset Class Instruments Data Fields Sources Equities 126 OHLCV + return Yahoo Finance (ETFs: broad market, sector, factor, international) Bonds 16 Close + return (ETFs); yield… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PortBench-Market.tabulartime-series-forecasting1K<n<10K3 likes94 downloads4mo agoHugging Face22jamesdborin /Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1-prompt-only Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1-prompt-only.tabular1K<n<10K0 likes87 downloads3mo agoHugging Face23nwhite-systems /responsible-agent-workflow-evaluation Responsible Agent Workflow Evaluation Version 1.0.0 contains 130 wholly synthetic scenarios for evaluating whether an AI agent respects safety, permission and accountability boundaries in operational settings. Thirteen categories contain ten scenarios each. Every record includes an intentionally unsafe request, contextual facts, expected safe behaviour, explicitly prohibited behaviour, severity, evaluation criteria and reviewer guidance. This is a red-team and… See the full description on the dataset page: https://huggingface.co/datasets/nwhite-systems/responsible-agent-workflow-evaluation.textn<1K1 likes87 downloads2mo agoHugging Face24agentlans /big-five-personality-traits Big Five Personality Traits Dataset This dataset contains AI-generated descriptions of personality traits based on the Big Five (OCEAN) model. For each trait and intensity level (1–5), five descriptions were produced by ten different chatbots: Grok, Gemini, Claude, KimiK2 (via HuggingChat), Deepseek, MetaAI, Perplexity, LeChat, ChatGPT, and Copilot. Overview The dataset can support tasks such as persona creation, comparative language analysis, and research on how AI… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/big-five-personality-traits.texttext-classification1K<n<10K0 likes85 downloads10mo agoHugging Face25Dharun72 /llm-agentic-precomputed-v3textn<1K0 likes84 downloads3mo agoHugging Face26agenticx /DILIranktext1K<n<10K0 likes83 downloads1y agoHugging Face27AgenticFinLab /PyFi-600K Dataset Card for PyFi-600K This dataset card aims to be a introduction for PyFi-600K, A financial VLM dataset containing 600K question-answer pairs generated via Adversarial agents. AgenticFinLab/PyFi-600K/ ├── README.md # Dataset documentation and description ├── images.zip # Compressed image files ├── PyFi-600K-dataset.csv # Q&A pairs in CSV format ├── PyFi-600K-dataset.json # Q&A pairs in JSON format ├── PyFi-600K-chain-dataset.json # Chain of Thought Q&A pairs dataset └──… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PyFi-600K.imagequestion-answering100K<n<1M1 likes81 downloads9mo agoHugging Face28jamesdborin /Nemotron-RL-Agentic-SWE-Pivot-v1-prompt-only Nemotron-RL-Agentic-SWE-Pivot-v1-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Agentic-SWE-Pivot-v1. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-SWE-Pivot-v1-prompt-only.tabular10K<n<100K0 likes79 downloads3mo agoHugging Face29Sepideh2027 /AgentYear: 2025License: MITAuthor: Sepideh Moafi PathogenAgentAI Instruction Dataset Dataset Description A ClinVar-derived dataset developed as part of the PathogenAgentAI research software project. The dataset is released in two parallel formats: Tabular version (train.csv, valid.csv, test.csv) — structured genomic-variant data for classical ML and analysis. BioGPT instruction version (biogpt_train.csv, biogpt_valid.csv, biogpt_test.csv) — instruction-style data… See the full description on the dataset page: https://huggingface.co/datasets/Sepideh2027/Agent.texttext-generation1M<n<10M0 likes78 downloads3d agoHugging Face30CanlahAI /agent-readiness-2026 Agent-Readiness of 50 Cross-Border DTC Brands (2026) Open dataset · CC BY 4.0 · published by Canlah AI (CANLAH AI PTE. LTD., Singapore) Canonical citation — cite the DOI: https://doi.org/10.5281/zenodo.22103177 ⚠️ v1.0.2 (2026-09-01) withdraws two claims from earlier versions — that three manifests had gone dark, and a ~12% churn rate derived from them. Both were an artifact of this repo's verification script probing the wrong path. The headline finding (25/25 platform-issued)… See the full description on the dataset page: https://huggingface.co/datasets/CanlahAI/agent-readiness-2026.tabularn<1K0 likes76 downloads21d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.