CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Exgentic /agent-llm-traces-v2 Exgentic Agent LLM Traces v2 — Agent Chat Only OpenTelemetry-shaped execution traces for 10,057 agent runs across 6 benchmarks (AppWorld, SWE-bench, BrowseCompPlus, τ²-bench Airline/Retail/Telecom), filtered to the agent under test's chat-only LLM calls. This is the dataset for replay testing, behavioral analysis, or any task where you care about what the benchmarked model actually did — not the eval scaffolding around it. This v2 release expands upon Exgentic/agent-llm-traces… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces-v2.tabulartext-generation10K<n<100K0 likes2.4k downloads3mo agoHugging Face02yatin-superintelligence /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M29 likes1.2k downloads7mo agoHugging Face03yatin-superintelligence /Audio-Video-Engineering-Agentic-Tasks-1M Audio/Video Engineering Agentic Tasks (1M) Abstract A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.tabulartext-generation1M<n<10M14 likes1.1k downloads7mo agoHugging Face04r0b0tlab /deepseek-v4-pro-0813-agentic DeepSeek-V4-Pro 0813 Agentic (DS4) A standalone, verifiable-first agentic training corpus: 19,072 training traces plus 2,135 held-out evaluation rows (validation 1,070 / test 1,065), generated by DeepSeek-V4-Pro 0813 (deepseek-v4-pro-0813, official API, thinking mode) across 13 verifiable task families, each row admitted only after passing a deterministic programmatic verifier. The corpus is designed to be directly usable for SFT, GRPO/RLVR, and NeMo Gym / NeMo RL (verified… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/deepseek-v4-pro-0813-agentic.tabulartext-generation10K<n<100K23 likes1.1k downloads1mo agoHugging Face05ru-dataset /agent-think-tool_use Agent Think Tool Use Датасет многошаговых агентных сессий для дообучения моделей работе с кодом, инструментами и инженерными задачами. Записи содержат пользовательские требования, комментарии агента во время работы, decision summaries, вызовы инструментов, результаты запусков, обработку ошибок и финальную проверку. Каждый shard представляет отдельную связанную сессию, а не отдельный вопрос и ответ. Данные охватывают исследование задачи, работу с документацией, проектирование… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/agent-think-tool_use.tabulartext-generationn<1K2 likes818 downloads5d agoHugging Face06Multi-Agent-LLMs /DEBATE DEBATE: Diverse Multi-Agent Debates This dataset is presented in the paper "MALLM: Multi-Agent Large Language Models Framework". Citation comming soon. tabulartext-generation10K<n<100K2 likes710 downloads1y agoHugging Face07yzhuang /Agentic-Long-Context-Understanding-QA 📖 Agentic Long Context Understanding 📖 Self-Taught Agentic Long Context Understanding (Arxiv). AgenticLU refines complex, long-context queries through self-clarifications and contextual grounding, enabling robust long-document understanding in a single pass. Installation Requirements This codebase is largely based on OpenRLHF and Helmet, kudos to them. The requirements are the same pip install openrlhf pip install -r ./HELMET/requirements.txt… See the full description on the dataset page: https://huggingface.co/datasets/yzhuang/Agentic-Long-Context-Understanding-QA.tabularquestion-answering100K<n<1M18 likes575 downloads1y agoHugging Face08rAVEUK /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M3 likes511 downloads6mo agoHugging Face09AgentCrush /agents-index AgentCrush Agent Index Evidence-ranked index of the AI agent economy. Updated daily from agentcrush.xyz. Overview 1,443 agents indexed across categories: developer tools, tokenized agents, service agents, model families 207 evidence-ranked with verified multi-signal scores Updated: 2026-09-24 Configs Config Description Rows agents All indexed agents with metadata ~1,443 evidence_ranked Evidence-ranked tier only ~207 snapshots_latest… See the full description on the dataset page: https://huggingface.co/datasets/AgentCrush/agents-index.tabulartext-classification1K<n<10K0 likes497 downloads8h agoHugging Face10ServiceNow-AI /AgentJudgeBench AgentJudgeBench: Evaluating LLM Judge Reliability on Agentic Tool-Calling A benchmark for systematically evaluating how reliably LLM judges assess agentic tool-calling workflows across structured, dependency-driven tasks. Why this benchmark? AgentJudgeBench measures how reliably LLM judges assess agentic tool-calling outputs. It provides 3,808 benchmark records spanning six DAG topologies and three difficulty… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/AgentJudgeBench.tabularquestion-answering100K<n<1M0 likes493 downloads23d agoHugging Face11FineEnvs /data-agent 📈 Data Agent Data-analysis tasks as a plain, load-and-go dataset — no runtime, no framework required. Each row is one self-contained task: a real tabular dataset, a question about it, and a deterministically-checkable gold answer. Load it, prompt any model however you like, and grade the result with the bundled grader. Where it comes from Built from the jupyter-agent dataset — real data-science notebooks over Kaggle datasets. Every question–answer pair was… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/data-agent.tabularquestion-answering1K<n<10K0 likes419 downloads24d agoHugging Face12LulaCola /AgentProcessBench AgentProcessBench AgentProcessBench is a benchmark for process-level evaluation of tool-using agents. Each example is a full agent trajectory with multi-turn messages, tool definitions, tool-use traces, reference outputs, and step-wise process labels. The benchmark contains 1,000 trajectories in total, with 250 examples from each subset: bfcl gaia_dev hotpotqa tau2 arxiv.org/abs/2603.14465 tabularquestion-answering1K<n<10K18 likes349 downloads6mo agoHugging Face13kryp1234 /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M1 likes297 downloads6mo agoHugging Face14AgentVQA /AgentVQA AgentVQA: A Multi-Domain Visual Question Answering Dataset AgentVQA is a comprehensive dataset for training and evaluating visual agents across multiple domains, including GUI interaction, spatial reasoning, game understanding, robot manipulation, and video perception. The dataset contains multiple-choice questions based on screenshots, images, and videos. Dataset Structure The dataset is organized into five main domains: 1. Web Agents Domain (GUI Interaction)… See the full description on the dataset page: https://huggingface.co/datasets/AgentVQA/AgentVQA.imagevisual-question-answering10K<n<100K1 likes235 downloads9mo agoHugging Face15anonymous-nips2026 /Agent-ValueBench Agent-ValueBench Agent-ValueBench constitutes the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions). This Hugging Face release contains both structured JSONL tables for dataset viewing and Croissant metadata generation, and the original raw benchmark artifacts. Repository Structure… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nips2026/Agent-ValueBench.tabularquestion-answering1K<n<10K0 likes228 downloads5mo agoHugging Face16Value4AI /Agent-ValueBench Agent-ValueBench Paper | Project Page | GitHub Agent-ValueBench is the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions). Repository Structure README.md data/ cases.jsonl rubrics.jsonl environments.jsonl raw/ case/ rubric/ environment/ Data Files… See the full description on the dataset page: https://huggingface.co/datasets/Value4AI/Agent-ValueBench.tabularother1K<n<10K3 likes164 downloads4mo agoHugging Face17AmanPriyanshu /tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified ToolACE - Tool-Use Agent Data Cleaned & Rectified 👥 Follow the Author Aman Priyanshu Overview This dataset is a cleaned and restructured version of the Team-ACE/ToolACE dataset. ToolACE is a high-quality conversational tool-use dataset containing 11,300+ examples of natural language interactions requiring function calling across diverse domains. This version converts the original OpenAI function-call format into a standardized multi-turn tool-use… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified.tabulartext-generation10K<n<100K0 likes132 downloads7mo agoHugging Face18AgenticFinLab /PortBench-QA PortBench QA Dataset Dataset Description 6,269 structured question-answer pairs probing correlation-based financial reasoning for multi-asset portfolio management, generated from the PortBench Market Base Dataset. Task Templates Template Task Complexity Pairs T1 Return prediction — direction for next N days 1 (single asset) 1,000 T2 Risk assessment — VaR at given confidence level 1 1,000 T3 Position sizing — given max drawdown… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PortBench-QA.tabularquestion-answering1K<n<10K3 likes128 downloads4mo agoHugging Face19SentientAGI /crypto-agent-safe-function-calling CrAI-SafeFuncCall Dataset 📄 Paper: Real AI Agents with Fake Memories: Fatal Context Manipulation Attacks on Web3 Agents 🤗 Dataset: CrAI-SafeFuncCall 📊 Benchmark: CrAI-Bench Overview The CrAI-SafeFuncCall dataset is designed to enhance the security of AI agents when performing function calls in the high-stakes domain of cryptocurrency and financial applications. It focuses on the critical challenge of detecting and mitigating memory injection attacks. Derived from the… See the full description on the dataset page: https://huggingface.co/datasets/SentientAGI/crypto-agent-safe-function-calling.tabularquestion-answering1K<n<10K3 likes127 downloads1y agoHugging Face20OxRML /AgentSLR AgentSLR: Priority Pathogens Dataset Paper Codebase Project Website This dataset accompanies the paper Evaluating AI-based Scientific Knowledge Synthesis with Epidemiological Systematic Reviews. It provides the data component of the AgentSLR evaluation harness. This covers article metadata, human abstract and full text screening labels, and structured human data extractions for epidemiological parameters… See the full description on the dataset page: https://huggingface.co/datasets/OxRML/AgentSLR.tabularquestion-answering10K<n<100K3 likes126 downloads5mo agoHugging Face210xKitkat /AgentForge-1152 AgentForge-1152: Tool-Verified General + Cyber Reasoning 1,152 evidence-grounded agent trajectories. 144 disjoint task families. A clean 50/50 split between general reasoning and defensive cybersecurity. AgentForge-1152 is a model-agnostic English corpus for supervised fine-tuning and evaluating assistants that use tools, recover from failed checks, and ground final answers in captured evidence. It uses portable messages and JSON function definitions rather than a model-specific… See the full description on the dataset page: https://huggingface.co/datasets/0xKitkat/AgentForge-1152.tabulartext-generation1K<n<10K0 likes121 downloads14d agoHugging Face22PurpleFlea /agent-financial-interactions Purple Flea Agent Financial Interactions A dataset of API interactions for autonomous AI agents using financial infrastructure (Purple Flea). Useful for training agents that handle crypto, trading, gambling, domain registration, escrow, and free onboarding via faucet. Research This dataset is referenced in: "Purple Flea: A Multi-Agent Financial Infrastructure Protocol for Autonomous AI Systems" https://doi.org/10.5281/zenodo.18808440 The paper covers the economic model… See the full description on the dataset page: https://huggingface.co/datasets/PurpleFlea/agent-financial-interactions.tabulartext-generationn<1K1 likes118 downloads7mo agoHugging Face23ArkhAngelLifeJiggy /deepseek-v4-pro-0813-agentic DeepSeek-V4-Pro 0813 Agentic (DS4) A standalone, verifiable-first agentic training corpus: 19,072 training traces plus 2,135 held-out evaluation rows (validation 1,070 / test 1,065), generated by DeepSeek-V4-Pro 0813 (deepseek-v4-pro-0813, official API, thinking mode) across 13 verifiable task families, each row admitted only after passing a deterministic programmatic verifier. The corpus is designed to be directly usable for SFT, GRPO/RLVR, and NeMo Gym / NeMo RL (verified… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/deepseek-v4-pro-0813-agentic.tabulartext-generation10K<n<100K0 likes95 downloads27d agoHugging Face24finaspirant /wearable-agent-trajectory-annotations Wearable Agent Trajectory Annotation Dataset Dataset Summary 50 wearable agent trajectories annotated by 5 LLM-simulated annotator personas using the agenteval-schema-v1 JSON schema, across two calibration phases (500 annotation records total). Designed to benchmark annotation-quality pipelines for agentic AI systems. Each trajectory captures a wearable AI agent responding to a real-time sensor event (health alert, privacy-sensitive context, location trigger… See the full description on the dataset page: https://huggingface.co/datasets/finaspirant/wearable-agent-trajectory-annotations.tabulartext-classificationn<1K1 likes94 downloads17d agoHugging Face25AgenticFinLab /PyFi-600K Dataset Card for PyFi-600K This dataset card aims to be a introduction for PyFi-600K, A financial VLM dataset containing 600K question-answer pairs generated via Adversarial agents. AgenticFinLab/PyFi-600K/ ├── README.md # Dataset documentation and description ├── images.zip # Compressed image files ├── PyFi-600K-dataset.csv # Q&A pairs in CSV format ├── PyFi-600K-dataset.json # Q&A pairs in JSON format ├── PyFi-600K-chain-dataset.json # Chain of Thought Q&A pairs dataset └──… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PyFi-600K.imagequestion-answering100K<n<1M1 likes90 downloads9mo agoHugging Face26jang1563 /sci-agent-verification-cascade Scientific Agent Verification Cascade Public evaluation fixtures and verified aggregate results for testing whether scientific claims keep their source, meaning, uncertainty, and verification requirements as they move between AI agents. This dataset accompanies the Scientific Agent Verification Cascade codebase. Version 0.2.0 contains synthetic evaluation data and aggregate-only results. It contains no raw hosted-model response, private holdout identifier, source-record… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/sci-agent-verification-cascade.tabulartext-generationn<1K0 likes85 downloads14d agoHugging Face27AgentsSci /EMNLP_Cost-Aware-Protocol-Routing Cost-Aware Protocol Routing: Matched Protocol Outcomes The short version. We ran the same 6,803 reasoning problems through four different LLM collaboration setups — from a single direct answer up to a four-agent deliberation — and recorded, for every problem, which ones got it right. Then we asked whether a model can look at a problem beforehand and predict which setup is worth paying for. It can predict whether it will fail. It cannot predict which collaboration protocol will… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/EMNLP_Cost-Aware-Protocol-Routing.tabulartabular-classification10K<n<100K0 likes84 downloads6h agoHugging Face28tmskss /eu-tenders-with-questions-for-agentic-checklist-filling eu-tenders-with-questions-for-agentic-checklist-filling Dataset Description This dataset contains questions and answers for evaluating Retrieval-Augmented Generation (RAG) systems in the context of generative agentic checklist-filling. The dataset is designed to benchmark various RAG architectures (Hybrid RAG, Graph RAG, Multi-Hop/Agentic RAG) on document analysis tasks. Dataset Summary Total Questions: 97 Document Families: 7 Languages: EN Domain: Procurement… See the full description on the dataset page: https://huggingface.co/datasets/tmskss/eu-tenders-with-questions-for-agentic-checklist-filling.documentquestion-answeringn<1K1 likes81 downloads5mo agoHugging Face29phynics /agentic-publication-protocol-dataset APP compare-app benchmark Paired reader conversations and blinded evaluations comparing an Agentic Publication Protocol (APP) paper agent against a general repository-aware agent, on 11 quantum-physics papers. For each paper, a neutral reader asks the same scripted questions to both agents; the two transcripts are anonymized and scored by a blinded evaluator on accuracy, informativeness, grounding, and honesty (1-10). Evaluator: Codex CLI, gpt-5.5, reasoning effort xhigh… See the full description on the dataset page: https://huggingface.co/datasets/phynics/agentic-publication-protocol-dataset.tabularquestion-answeringn<1K1 likes77 downloads3mo agoHugging Face30AdithyaSK /data_agent data_agent Plain, Harbor-free version of the data-analysis agent tasks — usable directly via load_dataset. Splits: train 5000, test 250, eval 144. Deterministic grading, no LLM judge. Columns task_id, source_row_id — ids question — the question to answer answer — gold answer; reward_mode (numeric/exact_short/exact_bool/list/list_csv/flexible), atol/rtol — how to grade difficulty_level (1-5), difficulty_tier (easy/medium/hard) kaggle_dataset — source Kaggle… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/data_agent.tabularquestion-answering1K<n<10K0 likes77 downloads24d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.