datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mimo-claude-code-traces-1k
MIMO Claude Code Traces
MIMO Claude Code Traces is a collection of coding-agent trajectories in a Claude Code-style environment. Each record contains a user coding task, the full multi-turn message trace, available tool schemas, assistant reasoning fields, tool calls, tool outputs, and metadata such as model name, category, duration, cost, token usage, and whether the trace used tools.
The traces were generated with mimo-v2.5-pro, MiMo's most capable model at the time of… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/mimo-claude-code-traces-1k.agent-llm-traces-v2
Exgentic Agent LLM Traces v2 — Agent Chat Only
OpenTelemetry-shaped execution traces for 10,057 agent runs across 6 benchmarks (AppWorld, SWE-bench, BrowseCompPlus, τ²-bench Airline/Retail/Telecom), filtered to the agent under test's chat-only LLM calls. This is the dataset for replay testing, behavioral analysis, or any task where you care about what the benchmarked model actually did — not the eval scaffolding around it.
This v2 release expands upon Exgentic/agent-llm-traces… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces-v2.bactrainus-hotpotqa-teacher-traces
Bactrainus HotpotQA Teacher Traces
SOURCE-LINKED v1.0.0
Archived Llama 3.1 rationale and question-decomposition supervision, paired with complete SFT conversations and stable HotpotQA identities.
198,660 ROWS
4 CONFIGURATIONS
SFT MESSAGES
8B + 70B LABELS
CC BY-SA 4.0
A focused release of recovered teacher-generated supervision for multi-hop question answering. Every row contains the normalized annotation, an ordered… See the full description on the dataset page: https://huggingface.co/datasets/bactrianus/bactrainus-hotpotqa-teacher-traces.figment-eval-traces
Figment Eval Traces
Synthetic and de-identified evaluation traces for Figment, a prototype protocol-navigation aid for trained rural-clinic and disaster-response field responders.
These records are intended for model and harness debugging. They are not clinical data, medical advice, diagnosis, treatment instructions, or a substitute for local protocol, clinician judgment, supervisor review, or trained responder judgment.
Dataset Summary
The dataset captures… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/figment-eval-traces.sample-fusion-intelligence-traces
Sample Fusion Intelligence Traces
Structured AI reasoning traces from dFusion's Fusion Intelligence system. Each record captures a complete agentic workflow: a real user query on a domain-specific topic, the full message chain including system prompts, tool calls, search results, intermediate reasoning steps, and a final synthesized answer — along with human feedback.
These are not synthetic benchmarks. They are traces from real queries submitted by real users on live financial… See the full description on the dataset page: https://huggingface.co/datasets/dFusionAILabs/sample-fusion-intelligence-traces.LIMO-TracesEvaluation of SOTA models on the LIMO dataset (v1).
LIMO is a small-scale dataset (817 Q&A) that was originally used for SFT LLMs to acquire reasoning capabilities.
This repo contains the answers provided by several recent models evaluated on the LIMO dataset, and the objective of it is to show that maybe, right now,
we are reaching a point where, nonreasoning models are performing well enough to make this dataset less appealing.
In the original LIMO paper, the LIMO dataset was used to… See the full description on the dataset page: https://huggingface.co/datasets/fedric95/LIMO-Traces.sangue-e-grafi-agent-traces
🩸 Sangue e Grafi — Agent Traces Dataset
100 recorded agent traces showing a KG-grounded agent solving adversarial Italian inheritance-law scenarios.
Dataset Description
This dataset contains 100 agent trace recordings from the Sangue e Grafi project. Each trace captures a complete reasoning episode: a small (4B) language model navigating a kinship knowledge graph via tool calls to answer adversarial inheritance-law questions in Italian.
These traces… See the full description on the dataset page: https://huggingface.co/datasets/cyberandy/sangue-e-grafi-agent-traces.scientific-agent-protocol-traces
SciAgentTrace
Matched cross-domain dataset of scientific-agent protocols. The central
comparison contains the same 6,653 problems under two actor models and four
protocols: 53,224 trajectories in 40 complete model--benchmark--protocol
groups. The broader table-first package contains 68,892
trajectories. Begin with trajectories, outcomes, or matched_outcomes, then
follow stable identifiers to messages and compressed raw traces.
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/scientific-agent-protocol-traces.PaperProf-traces
PaperProf Agent Trace
Step-by-step trace of PaperProf,
an AI study buddy that turns course PDFs into interactive quiz sessions.
What's in this dataset
Each row in paperprof_trace.jsonl is one LLM call. Fields:
Field
Description
session_id
Groups steps from the same session
step
Step index within the session (1–4)
type
question_generation / answer_evaluation / mcq_generation
topic
Domain of the source chunk
input
Exact input sent to the model… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/PaperProf-traces.fabella-traces
Fabella Anonymized Agent Traces
A public, anonymized log of the LangGraph ReAct loop inside Fabella, a small-model Gradio Space for parents who need help explaining hard things to their child in kid-appropriate language. The dataset exists for the Sharing is Caring merit badge in the Build Small Hackathon.
The first version of every explanation is drafted by google/gemma-4-E4B-it via a LangGraph ReAct loop with one tool (validate_explanation). A second small model —… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/fabella-traces.jev_turkish_mmlu_traces
JEV Turkish MMLU & MMLU-Pro Traces
Traces of Jev (jev-latest, TypeSafe System One) answering Turkish multiple-choice
questions from the turkish_mmlu and turkish-mmlu-pro-preview datasets.
Each source row becomes a choice question; rows are grouped by subject and sent as one
request per subject (the subject is the state). Every trace row records jev's chosen
option, confidence, probability distribution, and (when captured) the request id, token
usage, and latency.… See the full description on the dataset page: https://huggingface.co/datasets/aliarda/jev_turkish_mmlu_traces.
