datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lmcache-agentic-traces
LMCache Agentic Dataset Collection
A curated dataset collection of 787 multi-turn agentic LLM sessions (24,881 total LLM iterations) designed for benchmarking stateful LLM serving systems. Every session exhibits at least 5 turns with prefix growth and builds to at least 10K tokens of context — making it ideal for evaluating tiered KV Cache solutions like LMCache.
Motivation
Modern LLM agents (coding assistants, research agents, tool-calling systems) make dozens of… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/lmcache-agentic-traces.glm52-datagen-r11-100-agentic-function-calling-pivot-v2-tracesgrug-agentic-s3-step1903-repaired-eval-traces
Grug 67B repaired-export agentic evaluation traces
This dataset contains the final ATIF episode from 587 de-duplicated attempts in
a point-in-time snapshot of three active evaluations of
laion/grug-67b-a2b-sft-s3-agentic-step1903-repaired.
The snapshot was copied on 2026-07-29 at approximately 18:40 UTC.
Suite
Expected
Terminal attempts
Scored
Mean reward among scored
Exported trajectories
ID (dev_set_v2)
300
240
209
0.013963
237
SWE-bench Verified
300
126
81
0
126… See the full description on the dataset page: https://huggingface.co/datasets/laion/grug-agentic-s3-step1903-repaired-eval-traces.fable5-traces-agentic-cleanlmcache-agentic-traces
LMCache Agentic Dataset Collection
A curated dataset collection of 787 multi-turn agentic LLM sessions (24,881 total LLM iterations) designed for benchmarking stateful LLM serving systems. Every session exhibits at least 5 turns with prefix growth and builds to at least 10K tokens of context — making it ideal for evaluating tiered KV Cache solutions like LMCache.
Motivation
Modern LLM agents (coding assistants, research agents, tool-calling systems) make dozens of… See the full description on the dataset page: https://huggingface.co/datasets/DiscoPosse/lmcache-agentic-traces.lmcache-agentic-traces
LMCache Agentic Dataset Collection
A curated dataset collection of 787 multi-turn agentic LLM sessions (24,881 total LLM iterations) designed for benchmarking stateful LLM serving systems. Every session exhibits at least 5 turns with prefix growth and builds to at least 10K tokens of context — making it ideal for evaluating tiered KV Cache solutions like LMCache.
Motivation
Modern LLM agents (coding assistants, research agents, tool-calling systems) make dozens of… See the full description on the dataset page: https://huggingface.co/datasets/zeelHz/lmcache-agentic-traces.glm52-datagen-r11-101-agentic-indirect-prompt-injection-v2-tracesfable5-traces-agentic-clean-v2fable5-traces-agentic
fable5-traces-agentic
100K-max multi-domain coding + agentic SFT dataset.
Final rows
100,000
Targets
{
"agentic": 40000,
"coding": 20000,
"frontend": 11000,
"backend": 9000,
"reasoning": 15000,
"tools": 5000
}
Actual
{
"reasoning": 10321,
"frontend": 5384,
"tools": 5000,
"agentic": 50295,
"coding": 20000,
"backend": 9000
}
FABLE.5
Selected: 45,625
Target: 30,000
Minimum: 25,000
Build… See the full description on the dataset page: https://huggingface.co/datasets/usernamebetter/fable5-traces-agentic.grug-agentic-s3-step1903-agentic-evals-traces20260727-222003-grug-agentic-s3-step1903-grug-opencode-id-f947-tracescot-faithfulness-agentic-traces
Agentic traces: coding agents with planted test leaks
One row per trajectory of a Qwen3 coding agent working on an MBPP+ task inside a tiny repository
(spec, empty solution.py, three visible tests; hidden EvalPlus tests grade generality). Three
experiment families (experiment column):
experiment
what is planted
how a positive is certified
native_hacking
nothing; the agent may edit/skip tests or hardcode
programmatic: test/runner diff, skip/exit, hardcode + hidden-test… See the full description on the dataset page: https://huggingface.co/datasets/Narmeen07/cot-faithfulness-agentic-traces.grug-agentic-eval-v2-traces
