datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
data-agent-benchmarks
LongHorizon Full Data-Agent Benchmarks
Companion data artifacts for five complete evaluation tracks:
DataSciBench full55 / 167 metric entries
DABStep full450
DABStep-Research full100
DSBench Modeling full74
LongDS full68 / 2,225 turns
The companion GitHub repository contains processed manifests, evaluation code,
historical API ReAct baseline code, download/preparation tools, and the frozen
source lock. artifact_manifest.json records every uploaded object's size,
SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/noel7Y/data-agent-benchmarks.ecommerce-ai-data-analyst-agent-benchmark
E-commerce AI Data Analyst Agent Benchmark
A synthetic e-commerce dataset for evaluating AI data analyst agents on
realistic, multi-step business analysis, data-quality investigation, and
analytical reasoning.
This dataset is part of the
E-commerce AI Data Analyst Agent Benchmark.
Dataset summary
This dataset supports evaluation of AI data analyst agents on realistic,
multi-step e-commerce analysis.
It contains:
customers.csv
products.csv
orders.csv
returns.csv… See the full description on the dataset page: https://huggingface.co/datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark.agent-evaluation-benchmark
Agent Evaluation Benchmark
A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories.
Overview
This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more.
Categories
Category
Test Cases
Description
Weather
5
Forecasts, UV index, climate history
Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.industrial-agent-benchmark
Industrial Agent Benchmark
Industrial Agent Benchmark (IAB) is an open benchmark for evaluating Industrial AI systems, Manufacturing AI assistants, and Industrial Agents.
This Dataset Card describes the Hugging Face Dataset release for Industrial Agent Benchmark v2.2.0 Japanese Canonical Normalization.
Repository:
https://github.com/masahirosakae/industrial-agent-benchmark
Hugging Face Dataset Repository:
https://huggingface.co/datasets/MSakae/industrial-agent-benchmark… See the full description on the dataset page: https://huggingface.co/datasets/MSakae/industrial-agent-benchmark.sage-agent-benchmark
SAGE Agent Benchmark
Comprehensive benchmark for evaluating AI agent capabilities across three core competencies:
Tool Selection - Choosing appropriate tools for tasks
Task Planning - Decomposing complex tasks into step sequences
Timing Judgment - Deciding when to use tools vs. direct answers
Dataset Statistics
Total Samples: ~11,000
Tool Selection: ~6,000 samples
Task Planning: ~3,000 samples
Timing Judgment: ~2,000 samples
Splits: train, dev, test
Usage… See the full description on the dataset page: https://huggingface.co/datasets/intellistream/sage-agent-benchmark.agent-memory-benchmark
Agent Memory Compression & Evaluation Benchmark
This dataset is a controlled evaluation testbed designed to benchmark long-term memory architectures for conversational AI agents. It stress-tests how agents handle long conversations with complex fact dynamics.
Dataset Structure
1. conversation.json
A 100-turn synthetic conversation (50 user, 50 assistant turns) containing embedded facts categorized under:
Simple Facts: Baseline retrieval details.… See the full description on the dataset page: https://huggingface.co/datasets/kushalicious/agent-memory-benchmark.
