CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Agents-X /TIR-Bench TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning Introduction: TIR-Bench is a comprehensive benchmark designed to evaluate the "thinking-with-images" capabilities of Multimodal Large Language Models (MLLMs), addressing a gap left by existing benchmarks like Visual Search which only test basic operations. As models like OpenAI o3 begin to intelligently create and operate tools to transform images for problem-solving, TIR-Bench provides 13… See the full description on the dataset page: https://huggingface.co/datasets/Agents-X/TIR-Bench.imagequestion-answering1K<n<10K3 likes1.7k downloads9mo agoHugging Face02AgentCrush /agents-index AgentCrush Agent Index Evidence-ranked index of the AI agent economy. Updated daily from agentcrush.xyz. Overview 1,445 agents indexed across categories: developer tools, tokenized agents, service agents, model families 207 evidence-ranked with verified multi-signal scores Updated: 2026-09-25 Configs Config Description Rows agents All indexed agents with metadata ~1,445 evidence_ranked Evidence-ranked tier only ~207 snapshots_latest… See the full description on the dataset page: https://huggingface.co/datasets/AgentCrush/agents-index.tabulartext-classification1K<n<10K0 likes491 downloads10h agoHugging Face03agentscope-ai /OpenJudge OpenJudge Benchmark Dataset Benchmark dataset for evaluating graders across text, multimodal, and agent scenarios. This dataset supports the OpenJudge framework with labeled preference pairs for quality-assured grader development. Dataset Statistics Evaluation Benchmarks Category Task Files Samples 🤖 Agent 12 166 action 1 8 memory 3 47 plan 1 7 reflection 3 52 tool 4 52 🖼️ Multimodal 4 80 image_coherence 1 20 image_editing… See the full description on the dataset page: https://huggingface.co/datasets/agentscope-ai/OpenJudge.texttext-generation1K<n<10K2 likes307 downloads7mo agoHugging Face04llm-agents /CriticBench Dataset Card for Dataset Name CriticBench is a comprehensive benchmark designed to assess LLMs' abilities to generate, critique/discriminate and correct reasoning across a variety of tasks. CriticBench encompasses five reasoning domains: mathematical, commonsense, symbolic, coding, and algorithmic. It compiles 15 datasets and incorporates responses from three LLM families. Dataset Details Dataset Description Curated by: THU Funded by [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/llm-agents/CriticBench.textquestion-answering1K<n<10K16 likes213 downloads3y agoHugging Face05OxRML /AgentSLR AgentSLR: Priority Pathogens Dataset Paper Codebase Project Website This dataset accompanies the paper Evaluating AI-based Scientific Knowledge Synthesis with Epidemiological Systematic Reviews. It provides the data component of the AgentSLR evaluation harness. This covers article metadata, human abstract and full text screening labels, and structured human data extractions for epidemiological parameters… See the full description on the dataset page: https://huggingface.co/datasets/OxRML/AgentSLR.tabularquestion-answering10K<n<100K3 likes120 downloads5mo agoHugging Face06AgentsSci /EMNLP_Cost-Aware-Protocol-Routing Cost-Aware Protocol Routing: Matched Protocol Outcomes The short version. We ran the same 6,803 reasoning problems through four different LLM collaboration setups — from a single direct answer up to a four-agent deliberation — and recorded, for every problem, which ones got it right. Then we asked whether a model can look at a problem beforehand and predict which setup is worth paying for. It can predict whether it will fail. It cannot predict which collaboration protocol will… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/EMNLP_Cost-Aware-Protocol-Routing.tabulartabular-classification10K<n<100K0 likes91 downloads1d agoHugging Face07AgentsSci /AAAI_Wrong-but-Useful Wrong but Useful — Trajectory Value Dataset Companion dataset to Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages (arXiv:2608.14375 · PDF) · Code · Project page READ BEFORE USE — licensing and do-not-train terms This dataset redistributes third-party benchmark content (question text and gold answers) alongside model-generated text and the trajectory-value measurements that are this paper's contribution. The benchmark content… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/AAAI_Wrong-but-Useful.tabularquestion-answering1M<n<10M0 likes81 downloads21h agoHugging Face08emgena /omnimcp_agentselfheal_mcp_pro_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_agentselfheal_mcp_pro_teaser.texttext-generationn<1K0 likes79 downloads8d agoHugging Face09alirezaaminzadeh /agentshield-bench AgentShield-Bench Structured benchmark dataset for evaluating security of AI agents in tool-calling and MCP environments. Overview AgentShield-Bench provides 720 scenarios spanning 11 adversarial attack categories plus benign control cases. Each scenario includes trusted instructions, untrusted content, available tools, canary secrets, expected safe behaviors, and attack success conditions. Schema Field Type Description scenario_id string… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/agentshield-bench.texttext-classification1K<n<10K0 likes73 downloads2mo agoHugging Face10Narmeen07 /agent-session-handoff Agent Session Handoff Synthetic operations-agent transcripts for studying knowledge retention across a model swap: does a compressed KV-cache memory keep more of a session than a text summary when the next model takes over? Each episode is built from a sampled fact record (service owners, ports, branches, config values, ticket states, decisions) rendered by an LLM writer into a realistic user / assistant / tool-output session. The writer never sees the questions. A swap point… See the full description on the dataset page: https://huggingface.co/datasets/Narmeen07/agent-session-handoff.tabularquestion-answering1K<n<10K0 likes69 downloads8d agoHugging Face11AYI-NEDJIMI /ai-agents-fr Agents IA - Dataset Francais Dataset bilingue complet sur les Agents IA, les Frameworks Multi-Agents et le Model Context Protocol (MCP). Contenu du Dataset Categorie Nombre d'entrees Description Architectures d'Agents 15 ReAct, Plan-and-Execute, Reflexion, Multi-Agent, Hierarchique, Swarm, etc. Frameworks 12 CrewAI, AutoGen, LangGraph, Semantic Kernel, Haystack, Dify, etc. MCP & Tool Use 15 Architecture MCP, transports, function calling, securite… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-agents-fr.textquestion-answeringn<1K0 likes63 downloads7mo agoHugging Face12AYI-NEDJIMI /ai-agents-en AI Agents - English Dataset Comprehensive bilingual dataset on AI Agents, Multi-Agent Frameworks, and the Model Context Protocol (MCP). Dataset Contents Category Entry Count Description Agent Architectures 15 ReAct, Plan-and-Execute, Reflexion, Multi-Agent, Hierarchical, Swarm, etc. Frameworks 12 CrewAI, AutoGen, LangGraph, Semantic Kernel, Haystack, Dify, etc. MCP & Tool Use 15 MCP architecture, transports, function calling, security, ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-agents-en.textquestion-answeringn<1K0 likes63 downloads7mo agoHugging Face13searchsim /agentsim-atc AgentSim Agent-Trace Corpus (ATC) Grounded reasoning traces of retrieval-augmented question-answering agents, generated by the AgentSim platform. 103,567 reasoning steps spanning three established IR benchmarks (Quasar-T 38,915 + CausalQA 36,192 + MSMARCO 28,460), with 20,548 supervised query-document-answer triples extracted for fine-tuning, and 199,968 unique retrieved documents. Every reasoning step traces back to specific documents in the source corpus, enabling step-level… See the full description on the dataset page: https://huggingface.co/datasets/searchsim/agentsim-atc.tabularquestion-answering100K<n<1M1 likes59 downloads5mo agoHugging Face14values-md /when-agents-act Dataset Card for "When Agents Act" Dataset Summary This dataset contains 702 ethical decision judgements from 9 frontier LLMs (Claude Opus 4.5, GPT-5, GPT-5 Nano, Claude Sonnet 4.5, Claude Haiku 4.5, Gemini 3 Pro, Gemini 2.5 Flash, Grok-4, Grok-4 Fast) across 10 rigorously curated AI-relevant ethical dilemmas. Models were tested in both theory mode (hypothetical reasoning) and action mode (tool-enabled agents believing actions would execute). Key Finding: Models reverse… See the full description on the dataset page: https://huggingface.co/datasets/values-md/when-agents-act.tabulartext-classificationn<1K1 likes54 downloads10mo agoHugging Face15AgentsSci /NeurIPS26_Precise-but-Uncoupled Precise but Uncoupled Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Accepted to NeurIPS 2026 — Main Track Protocol traces, process metrics and derived tables for the paper. Reviewer detection quality and successful critique uptake are empirically separable: a multi-agent protocol can identify errors accurately and still fail to change the answer it carries forward. Resource Link 📄 Paper arXiv:2607.15388 · PDF 🌐… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/NeurIPS26_Precise-but-Uncoupled.tabularquestion-answering100K<n<1M0 likes43 downloads19h agoHugging Face16AgentsSci /IEEE2026_BigData_MAS-4-Science-Matching SciAgentTrace An execution-layer trace resource for scientific-agent workload characterization. A protocol fixes who reasons, what each role can see, when feedback returns, and when a workflow stops. Those choices determine the sequence of model requests that produces an answer, so protocol design is also workload design. Two workflows that consume similar token totals can issue very different request sequences. SciAgentTrace records that difference. The matched core runs the… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/IEEE2026_BigData_MAS-4-Science-Matching.tabulartext-generation1M<n<10M0 likes40 downloads1d agoHugging Face17searchsim /agentsim-atc-multihop AgentSim Agent-Trace Corpus — Multi-hop A multi-hop sibling of the AgentSim Agent-Trace Corpus (agentsim-atc) with an evolved schema designed for student model distillation. 1 490 accepted SFT trajectories plus 2 980 step-level DPO preference pairs, generated over 5 multi-hop QA datasets through a 7-action agentic executor with an Always-Search Policy filter. This corpus accompanies a follow-up technical report to "AgentSim: A Platform for Verifiable Agent-Trace Simulation"… See the full description on the dataset page: https://huggingface.co/datasets/searchsim/agentsim-atc-multihop.textquestion-answering1K<n<10K0 likes20 downloads5mo agoHugging Face18AgentsSci /scientific-agent-protocol-traces SciAgentTrace Matched cross-domain dataset of scientific-agent protocols. The central comparison contains the same 6,653 problems under two actor models and four protocols: 53,224 trajectories in 40 complete model--benchmark--protocol groups. The broader table-first package contains 68,892 trajectories. Begin with trajectories, outcomes, or matched_outcomes, then follow stable identifiers to messages and compressed raw traces. Repository:… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/scientific-agent-protocol-traces.tabulartext-generation1M<n<10M0 likes16 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.