CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Agents-X /TIR-Bench TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning Introduction: TIR-Bench is a comprehensive benchmark designed to evaluate the "thinking-with-images" capabilities of Multimodal Large Language Models (MLLMs), addressing a gap left by existing benchmarks like Visual Search which only test basic operations. As models like OpenAI o3 begin to intelligently create and operate tools to transform images for problem-solving, TIR-Bench provides 13… See the full description on the dataset page: https://huggingface.co/datasets/Agents-X/TIR-Bench.imagequestion-answering1K<n<10K3 likes1.7k downloads9mo agoHugging Face02AgentCrush /agents-index AgentCrush Agent Index Evidence-ranked index of the AI agent economy. Updated daily from agentcrush.xyz. Overview 1,443 agents indexed across categories: developer tools, tokenized agents, service agents, model families 207 evidence-ranked with verified multi-signal scores Updated: 2026-09-24 Configs Config Description Rows agents All indexed agents with metadata ~1,443 evidence_ranked Evidence-ranked tier only ~207 snapshots_latest… See the full description on the dataset page: https://huggingface.co/datasets/AgentCrush/agents-index.tabulartext-classification1K<n<10K0 likes497 downloads6h agoHugging Face03agentscope-ai /OpenJudge OpenJudge Benchmark Dataset Benchmark dataset for evaluating graders across text, multimodal, and agent scenarios. This dataset supports the OpenJudge framework with labeled preference pairs for quality-assured grader development. Dataset Statistics Evaluation Benchmarks Category Task Files Samples 🤖 Agent 12 166 action 1 8 memory 3 47 plan 1 7 reflection 3 52 tool 4 52 🖼️ Multimodal 4 80 image_coherence 1 20 image_editing… See the full description on the dataset page: https://huggingface.co/datasets/agentscope-ai/OpenJudge.texttext-generation1K<n<10K2 likes300 downloads7mo agoHugging Face04llm-agents /CriticBench Dataset Card for Dataset Name CriticBench is a comprehensive benchmark designed to assess LLMs' abilities to generate, critique/discriminate and correct reasoning across a variety of tasks. CriticBench encompasses five reasoning domains: mathematical, commonsense, symbolic, coding, and algorithmic. It compiles 15 datasets and incorporates responses from three LLM families. Dataset Details Dataset Description Curated by: THU Funded by [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/llm-agents/CriticBench.textquestion-answering1K<n<10K16 likes212 downloads3y agoHugging Face05nebulonsai /Multilingual-Representative-Agents-Dataset We have developed chat data containing over 2,300 files for support representatives responding to customers in different fields, with a total of 10 languages, 230 files for each language. Introducing v1. Each file contains over 10 conversations. These data have been created in strict compliance with privacy principles, using random human names and random company names. Supported languages: en:hi:zh:ar:ru:uk:es:tr:fr:az The format is Customer:Agent, and as specified in the metadata… See the full description on the dataset page: https://huggingface.co/datasets/nebulonsai/Multilingual-Representative-Agents-Dataset.question-answering1K<n<10K3 likes127 downloads9mo agoHugging Face06OxRML /AgentSLR AgentSLR: Priority Pathogens Dataset Paper Codebase Project Website This dataset accompanies the paper Evaluating AI-based Scientific Knowledge Synthesis with Epidemiological Systematic Reviews. It provides the data component of the AgentSLR evaluation harness. This covers article metadata, human abstract and full text screening labels, and structured human data extractions for epidemiological parameters… See the full description on the dataset page: https://huggingface.co/datasets/OxRML/AgentSLR.tabularquestion-answering10K<n<100K3 likes123 downloads5mo agoHugging Face07alirezaaminzadeh /agentshield-bench AgentShield-Bench Structured benchmark dataset for evaluating security of AI agents in tool-calling and MCP environments. Overview AgentShield-Bench provides 720 scenarios spanning 11 adversarial attack categories plus benign control cases. Each scenario includes trusted instructions, untrusted content, available tools, canary secrets, expected safe behaviors, and attack success conditions. Schema Field Type Description scenario_id string… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/agentshield-bench.texttext-classification1K<n<10K0 likes93 downloads2mo agoHugging Face08emgena /omnimcp_agentselfheal_mcp_pro_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_agentselfheal_mcp_pro_teaser.texttext-generationn<1K0 likes79 downloads7d agoHugging Face09Narmeen07 /agent-session-handoff Agent Session Handoff Synthetic operations-agent transcripts for studying knowledge retention across a model swap: does a compressed KV-cache memory keep more of a session than a text summary when the next model takes over? Each episode is built from a sampled fact record (service owners, ports, branches, config values, ticket states, decisions) rendered by an LLM writer into a realistic user / assistant / tool-output session. The writer never sees the questions. A swap point… See the full description on the dataset page: https://huggingface.co/datasets/Narmeen07/agent-session-handoff.tabularquestion-answering1K<n<10K0 likes68 downloads7d agoHugging Face10AYI-NEDJIMI /ai-agents-fr Agents IA - Dataset Francais Dataset bilingue complet sur les Agents IA, les Frameworks Multi-Agents et le Model Context Protocol (MCP). Contenu du Dataset Categorie Nombre d'entrees Description Architectures d'Agents 15 ReAct, Plan-and-Execute, Reflexion, Multi-Agent, Hierarchique, Swarm, etc. Frameworks 12 CrewAI, AutoGen, LangGraph, Semantic Kernel, Haystack, Dify, etc. MCP & Tool Use 15 Architecture MCP, transports, function calling, securite… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-agents-fr.textquestion-answeringn<1K0 likes65 downloads7mo agoHugging Face11searchsim /agentsim-atc AgentSim Agent-Trace Corpus (ATC) Grounded reasoning traces of retrieval-augmented question-answering agents, generated by the AgentSim platform. 103,567 reasoning steps spanning three established IR benchmarks (Quasar-T 38,915 + CausalQA 36,192 + MSMARCO 28,460), with 20,548 supervised query-document-answer triples extracted for fine-tuning, and 199,968 unique retrieved documents. Every reasoning step traces back to specific documents in the source corpus, enabling step-level… See the full description on the dataset page: https://huggingface.co/datasets/searchsim/agentsim-atc.tabularquestion-answering100K<n<1M1 likes64 downloads5mo agoHugging Face12AYI-NEDJIMI /ai-agents-en AI Agents - English Dataset Comprehensive bilingual dataset on AI Agents, Multi-Agent Frameworks, and the Model Context Protocol (MCP). Dataset Contents Category Entry Count Description Agent Architectures 15 ReAct, Plan-and-Execute, Reflexion, Multi-Agent, Hierarchical, Swarm, etc. Frameworks 12 CrewAI, AutoGen, LangGraph, Semantic Kernel, Haystack, Dify, etc. MCP & Tool Use 15 MCP architecture, transports, function calling, security, ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-agents-en.textquestion-answeringn<1K0 likes57 downloads7mo agoHugging Face13values-md /when-agents-act Dataset Card for "When Agents Act" Dataset Summary This dataset contains 702 ethical decision judgements from 9 frontier LLMs (Claude Opus 4.5, GPT-5, GPT-5 Nano, Claude Sonnet 4.5, Claude Haiku 4.5, Gemini 3 Pro, Gemini 2.5 Flash, Grok-4, Grok-4 Fast) across 10 rigorously curated AI-relevant ethical dilemmas. Models were tested in both theory mode (hypothetical reasoning) and action mode (tool-enabled agents believing actions would execute). Key Finding: Models reverse… See the full description on the dataset page: https://huggingface.co/datasets/values-md/when-agents-act.tabulartext-classificationn<1K1 likes52 downloads10mo agoHugging Face14uninhibited-scholar /agent-safety-bench-zh agent-safety-bench-zh · 中文 Agent 工具调用风险 / 提示注入 评测基准 评测一个 LLM / 安全护栏能否正确识别 agent 拟执行动作中的危险操作与被注入的恶意指令,并给出 allow / block 决策与风险分级。纯防御 / 安全教育 / 评测研究。 当前 v0.2,105 条(test split),三类:benign(应 allow) / prompt_injection(应 block) / destructive(应 block),含 easy/medium/hard。 每条字段:id, category, difficulty, context, action, gold{decision,risk}, rationale, tags。 加载 from datasets import load_dataset ds = load_dataset("uninhibited-scholar/agent-safety-bench-zh", split="test")… See the full description on the dataset page: https://huggingface.co/datasets/uninhibited-scholar/agent-safety-bench-zh.text-classificationn<1K0 likes26 downloads3mo agoHugging Face15searchsim /agentsim-atc-multihop AgentSim Agent-Trace Corpus — Multi-hop A multi-hop sibling of the AgentSim Agent-Trace Corpus (agentsim-atc) with an evolved schema designed for student model distillation. 1 490 accepted SFT trajectories plus 2 980 step-level DPO preference pairs, generated over 5 multi-hop QA datasets through a 7-action agentic executor with an Always-Search Policy filter. This corpus accompanies a follow-up technical report to "AgentSim: A Platform for Verifiable Agent-Trace Simulation"… See the full description on the dataset page: https://huggingface.co/datasets/searchsim/agentsim-atc-multihop.textquestion-answering1K<n<10K0 likes18 downloads5mo agoHugging Face16Gene829 /gene-llm-agents-instruct llm-agents-instruct v116 Auto-built (demand): 1 open request(s) and 0 recent download(s) for 'llm-agents' with no dataset newer than 14 days Kind: synthetic Domain: llm-agents Records: 1000 Created: 2026-07-08T17:36:15+00:00 SHA-256: 281020e4a1db9e063ea6eaf359b69cfa40a89f13faeae521a4179cec586fc10c Pipeline: v2.0.0 Filters: {"min_quality": 0.55, "limit": 1000, "source": null, "backend": "llama", "min_judge": 0.7} Generated by: Qwen3-4B-Instruct-2507-Q4_K_M.gguf (backend:… See the full description on the dataset page: https://huggingface.co/datasets/Gene829/gene-llm-agents-instruct.text-generation1K<n<10K1 likes16 downloads3mo agoHugging Face17Gene829 /gene-llm-agents-corpus llm-agents-corpus v92 Auto-built (demand): 1 open request(s) and 0 recent download(s) for 'llm-agents' with no dataset newer than 14 days Kind: scraped Domain: llm-agents Records: 702 Created: 2026-07-08T17:36:14+00:00 SHA-256: 74c00af747e66d1d4abf248168263fb9d5f1e424182d7adcdef160533c32dae2 Pipeline: v2.0.0 Filters: {"min_quality": 0.55, "limit": 1000, "source": null, "backend": null, "min_judge": null} Sources huggingface: 301 papers: 197 arxiv: 145 github:… See the full description on the dataset page: https://huggingface.co/datasets/Gene829/gene-llm-agents-corpus.text-generationn<1K0 likes12 downloads3mo agoHugging Face18haritzpuerto /MetaQA_Agents_Predictions Dataset Card for MetaQA Agents' Predictions Dataset Summary This dataset contains the answer predictions of the QA agents for the QA datasets used in MetaQA paper. In particular, it contains the following QA agents' predictions: Span-Extraction Agents Agent: Span-BERT Large (Joshi et al.,2020) trained on SQuAD. Predictions for: SQuAD NewsQA HotpotQA SearchQA Natural Questions TriviaQA-web QAMR DuoRC DROP Agent: Span-BERT Large (Joshi et al.,2020) trained on… See the full description on the dataset page: https://huggingface.co/datasets/haritzpuerto/MetaQA_Agents_Predictions.question-answering1 likes10 downloads4y agoHugging Face19AgentsSci /scientific-agent-protocol-traces SciAgentTrace Matched cross-domain dataset of scientific-agent protocols. The central comparison contains the same 6,653 problems under two actor models and four protocols: 53,224 trajectories in 40 complete model--benchmark--protocol groups. The broader table-first package contains 68,892 trajectories. Begin with trajectories, outcomes, or matched_outcomes, then follow stable identifiers to messages and compressed raw traces. Repository:… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/scientific-agent-protocol-traces.tabulartext-generation1M<n<10M0 likes3 downloads2mo agoHugging Face20AgentsSci /EMNLP_Cost-Aware-Protocol-Routing Cost-Aware Protocol Routing: Matched Protocol Outcomes The short version. We ran the same 6,803 reasoning problems through four different LLM collaboration setups — from a single direct answer up to a four-agent deliberation — and recorded, for every problem, which ones got it right. Then we asked whether a model can look at a problem beforehand and predict which setup is worth paying for. It can predict whether it will fail. It cannot predict which collaboration protocol will… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/EMNLP_Cost-Aware-Protocol-Routing.tabulartabular-classification10K<n<100K0 likes4h agoHugging Face21AgentsSci /AAAI_Wrong-but-Useful Wrong but Useful — Trajectory Value Dataset Companion dataset to Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages (arXiv:2608.14375 · PDF) · Code · Project page READ BEFORE USE — licensing and do-not-train terms This dataset redistributes third-party benchmark content (question text and gold answers) alongside model-generated text and the trajectory-value measurements that are this paper's contribution. The benchmark content… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/AAAI_Wrong-but-Useful.tabularquestion-answering1M<n<10M0 likes4h agoHugging Face22AgentsSci /NeurIPS26_Precise-but-Uncoupled NeurIPS26_Precise-but-Uncoupled ⚠️ LICENCE — READ BEFORE USE. TWO RESTRICTIONS PASS DOWNSTREAM TO YOU. This dataset is license: other — a COMPOSITE. There is no single permissive licence. NONCOMMERCIAL. The MaScQA portion (642 problems) is CC-BY-NC-SA-4.0. commercial_use_policy = not_allowed. Commercial use of that portion is not permitted. SHAREALIKE + DO-NOT-TRAIN. The LAB-Bench portion (741 problems) is CC-BY-SA-4.0 and carries the upstream contamination… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/NeurIPS26_Precise-but-Uncoupled.tabularquestion-answering100K<n<1M0 likes5h agoHugging Face23AgentsSci /IEEE2026_BigData_MAS-4-Science-Matching SciAgentTrace An execution-layer trace resource for scientific-agent workload characterization. A protocol fixes who reasons, what each role can see, when feedback returns, and when a workflow stops. Those choices determine the sequence of model requests that produces an answer, so protocol design is also workload design. Two workflows that consume similar token totals can issue very different request sequences. SciAgentTrace records that difference. The matched core runs the… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/IEEE2026_BigData_MAS-4-Science-Matching.tabulartext-generation1M<n<10M0 likes5h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.