datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-llm-traces-v2
Exgentic Agent LLM Traces v2 — Agent Chat Only
OpenTelemetry-shaped execution traces for 10,057 agent runs across 6 benchmarks (AppWorld, SWE-bench, BrowseCompPlus, τ²-bench Airline/Retail/Telecom), filtered to the agent under test's chat-only LLM calls. This is the dataset for replay testing, behavioral analysis, or any task where you care about what the benchmarked model actually did — not the eval scaffolding around it.
This v2 release expands upon Exgentic/agent-llm-traces… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces-v2.LLM-Agent-Harness-Survey
English | 中文
Agent Harness for Large Language Model Agents: A Survey
⭐ This repo is actively maintained. If you find it useful, please star the repo to stay updated and help others find it.
The agent execution harness — not the model — is the primary determinant of agent reliability at scale.This survey formalizes the harness as a first-class architectural object H = (E, T, C, S, L, V), surveys 110+ papers, blogs and reports across 23 systems, and maps 9 open… See the full description on the dataset page: https://huggingface.co/datasets/GloriaaaM/LLM-Agent-Harness-Survey.agent-llm-traces
Multi-Benchmark LLM Agent Traces
A comprehensive dataset of OpenTelemetry traces capturing LLM inference behavior across multiple agent frameworks, benchmarks, and model providers. This dataset enables research into LLM performance analysis, agent behavior patterns, and inference optimization.
Collected by Exgentic - A platform for LLM observability and performance optimization.
Dataset Overview
This dataset contains 1,781 execution traces capturing detailed agent… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces.automl_llm_agent_m2
AutoML-LLM Agent Module 2 Benchmark
This dataset contains the Module 2 benchmark for evaluating an assistant that converts a Module 1 recipe, a user request, and a processed tabular dataset into an auditable autogluon.cloud.TabularCloudPredictor configuration.
The repository is scoped to Module 2 only.
Tables
module2_cases: one row per Module 2 evaluation case.
module2_queries: user requests for each case.
module2_reference_configs: legacy reference JSON files… See the full description on the dataset page: https://huggingface.co/datasets/tecnologiactc/automl_llm_agent_m2.automl_llm_agent_m1
AutoML-LLM Agent Module 1 Benchmark
This dataset contains the Module 1 benchmark for evaluating an AutoML assistant that interprets user requests, selects tabular modeling settings, produces an auditable AutoGluon Tabular plan, and synthesizes the compact Module 1 recipe consumed by Module 2 through the mandatory final LLM writer used by all A-E variants.
The repository is scoped to Module 1 only.
Tables
cases: one row per Module 1 evaluation case.
queries: one… See the full description on the dataset page: https://huggingface.co/datasets/tecnologiactc/automl_llm_agent_m1.details_llm-agents__tora-70b-v1.0
Dataset Card for Evaluation run of llm-agents/tora-70b-v1.0
Dataset automatically created during the evaluation run of model llm-agents/tora-70b-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-70b-v1.0.A-Survey-for-LLM-Agent-Trajectory-Analysis
A Survey for LLM Agent Trajectory Analysis
This dataset repository hosts the survey paper A Survey for LLM Agent Trajectory Analysis: From Failure Attribution to Enhancement and a structured metadata snapshot of the companion paper collection from Awesome-LLM-Agent-Trajectory-Analysis.
The repository is intended for discovery, citation, and lightweight analysis of the literature around LLM agent trajectory analysis, including failure attribution, trajectory-based debugging… See the full description on the dataset page: https://huggingface.co/datasets/RobinChen2001/A-Survey-for-LLM-Agent-Trajectory-Analysis.CriticBench
Dataset Card for Dataset Name
CriticBench is a comprehensive benchmark designed to assess LLMs' abilities to generate, critique/discriminate and correct reasoning across a variety of tasks. CriticBench encompasses five reasoning domains: mathematical, commonsense, symbolic, coding, and algorithmic. It compiles 15 datasets and incorporates responses from three LLM families.
Dataset Details
Dataset Description
Curated by: THU
Funded by [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/llm-agents/CriticBench.agentic-llm-pretraining-1.7b
Agentic LLM Pretraining Dataset
A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.agent-llm-traces
Multi-Benchmark LLM Agent Traces
A comprehensive dataset of OpenTelemetry traces capturing LLM inference behavior across multiple agent frameworks, benchmarks, and model providers. This dataset enables research into LLM performance analysis, agent behavior patterns, and inference optimization.
Collected by Exgentic - A platform for LLM observability and performance optimization.
Dataset Overview
This dataset contains 1,781 execution traces capturing detailed agent… See the full description on the dataset page: https://huggingface.co/datasets/DiscoPosse/agent-llm-traces.aiwolf-nlp-agent-llm
AIWolfDial 2026 Power Play Evaluation
Public data release: 2026-09-14. This dataset is available at synonym/aiwolf-nlp-agent-llm, with the snapshot tag release-20260914. The matching code distribution is 1.0.0-rc.3, commit d427dc299bacf4eb4cb41c114c8af476b71ac7ed. The code distribution uses a single root commit; this dataset is separate and is not included in that repository. Paper publication identifiers are still pending. The dataset is distributed under the MIT license in… See the full description on the dataset page: https://huggingface.co/datasets/synonym/aiwolf-nlp-agent-llm.llm-prompt-collection
LLM Prompt Collection
Prompts from the first 100 000 rows of each dataset were collected then deduplicated and shuffled.
The prompts_k* configs are semantically clustered subsets of the all config for diversity and coverage.
The screened_prompts config is a subset of safe, high-quality prompts for the all config as classified using agentlans/bge-small-en-v1.5-prompt-screener
Source
Rows
agentlans/chatgpt all
100 000
agentlans/magpie all
99 982… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/llm-prompt-collection.llm-agentic-precomputed-v3details_llm-agents__tora-code-13b-v1.0
Dataset Card for Evaluation run of llm-agents/tora-code-13b-v1.0
Dataset automatically created during the evaluation run of model llm-agents/tora-code-13b-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-code-13b-v1.0.llm-agent-harness-reliability-next-prime
LLM Next Prime Harness Dataset
This dataset contains raw observations from an experiment studying how the agent harness affects reliability when an LLM has access to a deterministic tool.
The task is deliberately simple and objectively verifiable:
What is the smallest prime number that is strictly greater than n?
The deterministic tool computes the correct answer with a local Python next_prime(n) function. The experiment asks whether failures come from the model, the provider… See the full description on the dataset page: https://huggingface.co/datasets/hoololi/llm-agent-harness-reliability-next-prime.mm-llm-coder-agent-dataset
Coder Agent Dataset
Agent workflow dataset for training coding agents. Contains multi-step coding tasks with tool usage patterns, execution validation, and quality metrics.
Skill Type: Agent/ Skill
This dataset is part of the combined Myanmar LLM dataset collection:
chat-skill.md - [amkyawdev/ myanmar-llm-data](https://huggingface. co/datasets/ amkyawdev/ myanmar-llm-data)
agent-skill.md - Myanmar conversational data, translations, Q&A
code-skill.md- [amkyawdev/… See the full description on the dataset page: https://huggingface.co/datasets/amkyawdev/mm-llm-coder-agent-dataset.details_llm-agents__tora-7b-v1.0
Dataset Card for Evaluation run of llm-agents/tora-7b-v1.0
Dataset Summary
Dataset automatically created during the evaluation run of model llm-agents/tora-7b-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-7b-v1.0.details_llm-agents__tora-code-34b-v1.0
Dataset Card for Evaluation run of llm-agents/tora-code-34b-v1.0
Dataset automatically created during the evaluation run of model llm-agents/tora-code-34b-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-code-34b-v1.0.ECHO-Terminal-Agent-Prepared-Data
ECHO-style Terminal Agent Prepared Data for LFM RLVR
Prepared on 2026-06-09 for local no-Docker LFM terminal RLVR experiments.
This dataset converts public terminal-agent task archives into two formats:
echo_terminal_tasks_*.parquet: ECHO/SkyRL-style rows with prompt, path, and task_binary.
lfm_live_tasks_mixed.*: local LFM no-Docker trainer rows with prompt, task_id, source, task_binary_b64, and metadata.
Current manifest:
total rows: 1500
Endless Terminals: 772 rows… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/ECHO-Terminal-Agent-Prepared-Data.automl_llm_agent_m3
AutoML-LLM Agent Module 3 Benchmark
This dataset contains the Module 3 benchmark for evaluating an assistant that converts AutoGluon training evidence and a surrogate XAI report into a decision-oriented report for a domain expert without AI training.
The repository is scoped to Module 3 only. Its seven cases reuse the processed datasets and reference configurations established by Modules 1 and 2.
Tables
module3_cases: one row per evaluation case, with canonical… See the full description on the dataset page: https://huggingface.co/datasets/tecnologiactc/automl_llm_agent_m3.agentlans__Llama3.1-Daredevilish-Instruct-details
Dataset Card for Evaluation run of agentlans/Llama3.1-Daredevilish-Instruct
Dataset automatically created during the evaluation run of model agentlans/Llama3.1-Daredevilish-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/agentlans__Llama3.1-Daredevilish-Instruct-details.Josephgflowers__Tinyllama-STEM-Cinder-Agent-v1-details
Dataset Card for Evaluation run of Josephgflowers/Tinyllama-STEM-Cinder-Agent-v1
Dataset automatically created during the evaluation run of model Josephgflowers/Tinyllama-STEM-Cinder-Agent-v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Josephgflowers__Tinyllama-STEM-Cinder-Agent-v1-details.agentlans__Llama-3.2-1B-Instruct-CrashCourse12K-details
Dataset Card for Evaluation run of agentlans/Llama-3.2-1B-Instruct-CrashCourse12K
Dataset automatically created during the evaluation run of model agentlans/Llama-3.2-1B-Instruct-CrashCourse12K
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/agentlans__Llama-3.2-1B-Instruct-CrashCourse12K-details.llm-agentic-swiss-legal-checkpoints
LLM Agentic Legal Information Retrieval — Checkpoints
Public artifacts from competing in the Kaggle competition.
Best public LB
Submission
LB
v6 LightGBM baseline
0.0709
v9 DeepSeek paragraph injection
0.13167
v11 court sibling expansion
0.13665
v12 Qwen2.5-14B LoRA
0.13204
Structure
submissions/ — final submission CSVs per version
picks/ — per-query LLM output caches (V4-Pro picks, LoRA picks, profiles)
training/ — LEXam fine-tuning data… See the full description on the dataset page: https://huggingface.co/datasets/Dharun72/llm-agentic-swiss-legal-checkpoints.EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-COT-details.details_llm-agents__tora-13b-v1.0
Dataset Card for Evaluation run of llm-agents/tora-13b-v1.0
Dataset automatically created during the evaluation run of model llm-agents/tora-13b-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-13b-v1.0.EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.004-128K-code-ds-auto-details.details_Josephgflowers__TinyLlama-Cinder-Agent-RagLFM2.5-KO-Agentic-Fable-Grounded-LFMChat-Raw
LFM2.5-KO-Agentic-Fable-Grounded-LFMChat-Raw
Fable5/Helio Korean agentic traces and local grounded document/log examples converted to LFM chat JSONL.
This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow.
Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT
CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL
Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT
SFT GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-Agentic-Fable-Grounded-LFMChat-Raw.details_llm-agents__tora-code-7b-v1.0
Dataset Card for Evaluation run of llm-agents/tora-code-7b-v1.0
Dataset Summary
Dataset automatically created during the evaluation run of model llm-agents/tora-code-7b-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-code-7b-v1.0.
