datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
data-agent
📈 Data Agent
Data-analysis tasks as a plain, load-and-go dataset — no runtime, no framework required. Each
row is one self-contained task: a real tabular dataset, a question about it, and a
deterministically-checkable gold answer. Load it, prompt any model however you like, and grade the
result with the bundled grader.
Where it comes from
Built from the jupyter-agent dataset
— real data-science notebooks over Kaggle datasets. Every question–answer pair was… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/data-agent.data-agent-benchmarks
LongHorizon Full Data-Agent Benchmarks
Companion data artifacts for five complete evaluation tracks:
DataSciBench full55 / 167 metric entries
DABStep full450
DABStep-Research full100
DSBench Modeling full74
LongDS full68 / 2,225 turns
The companion GitHub repository contains processed manifests, evaluation code,
historical API ReAct baseline code, download/preparation tools, and the frozen
source lock. artifact_manifest.json records every uploaded object's size,
SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/noel7Y/data-agent-benchmarks.sft-datatool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified
ToolACE - Tool-Use Agent Data Cleaned & Rectified
👥 Follow the Author
Aman Priyanshu
Overview
This dataset is a cleaned and restructured version of the Team-ACE/ToolACE dataset. ToolACE is a high-quality conversational tool-use dataset containing 11,300+ examples of natural language interactions requiring function calling across diverse domains. This version converts the original OpenAI function-call format into a standardized multi-turn tool-use… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified.data_agent
data_agent
Plain, Harbor-free version of the data-analysis agent tasks — usable directly via load_dataset.
Splits: train 5000, test 250, eval 144. Deterministic grading, no LLM judge.
Columns
task_id, source_row_id — ids
question — the question to answer
answer — gold answer; reward_mode (numeric/exact_short/exact_bool/list/list_csv/flexible), atol/rtol — how to grade
difficulty_level (1-5), difficulty_tier (easy/medium/hard)
kaggle_dataset — source Kaggle… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/data_agent.ecommerce-ai-data-analyst-agent-benchmark
E-commerce AI Data Analyst Agent Benchmark
A synthetic e-commerce dataset for evaluating AI data analyst agents on
realistic, multi-step business analysis, data-quality investigation, and
analytical reasoning.
This dataset is part of the
E-commerce AI Data Analyst Agent Benchmark.
Dataset summary
This dataset supports evaluation of AI data analyst agents on realistic,
multi-step e-commerce analysis.
It contains:
customers.csv
products.csv
orders.csv
returns.csv… See the full description on the dataset page: https://huggingface.co/datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark.dataagentbench-derived-official54-altimate-prompt-shell
DataAgentBench-Derived Official54 Altimate Prompt Shell
This public dataset contains 54 DataAgentBench-derived task prompt-shell rows compiled for the Altimate-centered DAB sandbox runtime. It is a derived compile/export artifact, not the official raw DataAgentBench release.
It is intended as a portable task/prompt/manifest source for downstream Altimate teacher rollouts, SFT construction, or RL data compilation. It is not a completed rollout dataset: the rows do not contain… See the full description on the dataset page: https://huggingface.co/datasets/forseasons/dataagentbench-derived-official54-altimate-prompt-shell.how-agentic-m1-research-data
How Agentic M1 Research Data
Training and validation data used for the M1-stage 500M-parameter
How Agentic research model.
Files
pretrain/m1_pretrain_5b_clean_train.jsonl.gz: cleaned pretraining split.
pretrain/m1_pretrain_5b_clean_val.jsonl.gz: pretraining validation split.
pretrain/m1_pretrain_5b_clean_report.json: corpus construction and quality report.
pretrain/m1_pretrain_5b_clean_rejected_sample.jsonl.gz: a small sample of rejected records for auditing.… See the full description on the dataset page: https://huggingface.co/datasets/jjyaoao/how-agentic-m1-research-data.agentic-publication-protocol-dev-data
APP compare-app benchmark
Paired reader conversations and blinded evaluations comparing an Agentic
Publication Protocol (APP) paper agent against a general repository-aware
agent, on 11 public quantum-physics papers. This is the public-paper
subset reported in the APP paper's compare-app table.
For each paper, a neutral reader asks the same scripted questions to both agents;
the two transcripts are anonymized and scored by a blinded evaluator on
accuracy, informativeness… See the full description on the dataset page: https://huggingface.co/datasets/LionSR/agentic-publication-protocol-dev-data.agentmemorybench-data
AgentMemoryBench Data
This dataset repository stores runtime benchmark data used by
AgentMemoryBench.
This repository is intended for public download by AgentMemoryBench users. Please keep upstream
benchmark licenses, citations, and redistribution notes up to date before broad distribution.
Repository
Dataset repo: BuptZZP/agentmemorybench-data
Generated at: 2026-06-17T08:22:11.782773+00:00
Source root name: data
Total data files: 1537
Total data bytes:… See the full description on the dataset page: https://huggingface.co/datasets/BuptZZP/agentmemorybench-data.toucan-agentic-thinking
Toucan Agentic with Thinking Dataset
This dataset contains agentic reasoning responses generated by MiniMax-M2.1 based on questions from Agent-Ark/Toucan-1.5M_SFT.
Dataset Description
For each user question, the model generates:
Thinking process: The model's reasoning wrapped in <think> tags
Response: A complete, helpful answer in natural language
The original tool definitions are preserved in the tools field for reference.
Statistics
Split
Examples… See the full description on the dataset page: https://huggingface.co/datasets/agent-data/toucan-agentic-thinking.
