datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
figment-eval-traces
Figment Eval Traces
Synthetic and de-identified evaluation traces for Figment, a prototype protocol-navigation aid for trained rural-clinic and disaster-response field responders.
These records are intended for model and harness debugging. They are not clinical data, medical advice, diagnosis, treatment instructions, or a substitute for local protocol, clinician judgment, supervisor review, or trained responder judgment.
Dataset Summary
The dataset captures… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/figment-eval-traces.PaperProf-traces
PaperProf Agent Trace
Step-by-step trace of PaperProf,
an AI study buddy that turns course PDFs into interactive quiz sessions.
What's in this dataset
Each row in paperprof_trace.jsonl is one LLM call. Fields:
Field
Description
session_id
Groups steps from the same session
step
Step index within the session (1–4)
type
question_generation / answer_evaluation / mcq_generation
topic
Domain of the source chunk
input
Exact input sent to the model… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/PaperProf-traces.packetcourt-golden-cases
PacketCourt Golden Cases
A small evidence-first evaluation set for auditing front-of-pack claims against
the text printed on the same Indian packaged-food label.
Each record contains:
front-label claim text
back-label evidence text
expected claims and conservative verdicts
expected persuasion-gap concepts
expected deterministic date or whole-packet calculations
The initial set is intentionally small and hand-audited. It is a regression and
demonstration asset, not a… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/packetcourt-golden-cases.fabella-traces
Fabella Anonymized Agent Traces
A public, anonymized log of the LangGraph ReAct loop inside Fabella, a small-model Gradio Space for parents who need help explaining hard things to their child in kid-appropriate language. The dataset exists for the Sharing is Caring merit badge in the Build Small Hackathon.
The first version of every explanation is drafted by google/gemma-4-E4B-it via a LangGraph ReAct loop with one tool (validate_explanation). A second small model —… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/fabella-traces.agent-eval-golden-dataset
Tech Interview Agent — Golden Eval Dataset
Stop guessing whether your AI interviewer is good. Start measuring it.
This dataset provides ground-truth benchmarks for evaluating AI agents that conduct tech job interviews. Each record is a structured test case: give it to your agent, collect the response, run it through the AI Agent Evaluation Pipeline, and get objective scores — no human review needed.
Generated by NVIDIA Nemotron-3-Nano-30B-A3B.
What's inside
40… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agent-eval-golden-dataset.
