data-agent
agents-last-exam-data
Agents Last Exam — Task Input Data
Input files (the materials each task hands to the agent at run start) for the
Agents Last Exam (ALE) benchmark. Browsable per-task directory layout.
The Agents Last Exam dataset family
ALE is published as three companion HuggingFace datasets:
Dataset
Contents
Access
Task Card Metadata
One row per task: titles, prompts, taxonomy, input-file descriptors
Open
Task Input Data
The input/ files each task hands the agent at… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam-data.agent-data-collection
Agent Data Collection
A comprehensive collection of agent interaction datasets for training and evaluating AI agents across diverse domains and tasks.
This dataset aggregates high-quality agent trajectories from various environments including web browsing, code generation, household tasks, knowledge base querying, and software engineering.
The dataset is collected through methods described in Agent Data Protocol.
Dataset Splits
Each dataset configuration provides up… See the full description on the dataset page: https://huggingface.co/datasets/neulab/agent-data-collection.data-agent-harbor-eval
🧪 Data Agent — Harbor (eval)
A small, difficulty-balanced validation split — 144 tasks — perfect for quick checkpoints
while you train. Same idea as the rest of the family: your agent gets a real dataset and a
question, explores and answers, and everything is graded deterministically, no LLM judge.
Packaged in Harbor format.
Where it comes from
Built from the jupyter-agent dataset
(real notebooks over Kaggle datasets). Every task was verified — a strong agent… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/data-agent-harbor-eval.Agriculture-Agent-RL-Training-Data
Agriculture Agent RL Training Data
A growing dataset of RL rollout trajectories for LLM agents on
natural/regenerative farming — the first RL/trajectory-shaped dataset in the
Copyleft Cultivars collection
(every prior dataset here is SFT/conversational Q&A). Agents call real tools
(primarily cultivars-mcp,
a plant-genomics MCP server) across 9 knowledge categories (plus a 10th,
organic_chemistry_soil_science, added 2026-08-11, and an 11th,
organic_chemistry_synthesis, added… See the full description on the dataset page: https://huggingface.co/datasets/CopyleftCultivars/Agriculture-Agent-RL-Training-Data.data-agent-harbor-train
📊 Data Agent — Harbor (train)
Teach an agent to actually do data science. This is a suite of 5,000 hands-on
data-analysis tasks: each one drops your agent into a sandbox with a real dataset and a
question, and asks it to explore the data, compute the answer, and write it down. Every answer is
checked deterministically — no LLM judge, no guesswork.
It's packaged in Harbor format, so it runs as a
ready-made agentic environment.
Where it comes from
Built from the… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/data-agent-harbor-train.agents-last-exam-data-archive
Agents Last Exam — Task Data Archive (input + reference)
⚠️ Gated dataset. This repo packages each task's input, software, and
reference (ground-truth) data into a single archive (ale-tasks-data.tar.gz)
for convenient one-shot download — in particular for running ALE locally with
the local Docker provider,
which fetches it and mounts each task's data at run time. Because it includes
the reference outputs used to score runs, access requires login, agreement to
the terms on the… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam-data-archive.
