agentic-evals
agentic-evals-artifacts
On Randomness in Agentic Evals — Results
This dataset contains the trajectory and evaluation results from the paper On Randomness in Agentic Evals. Agents are benchmarked on SWE-bench Verified across different scaffolds, models, and temperatures, with 10 independent runs per setting to enable pass@k and variance analysis.
Downloading the Data
Option 1 — HuggingFace CLI:
pip install huggingface-hub
huggingface-cli download ASSERT-KTH/agentic-evals-artifacts --repo-type… See the full description on the dataset page: https://huggingface.co/datasets/ASSERT-KTH/agentic-evals-artifacts.persona-and-other-evals
Qwen3.5-9B AMA adapters — persona evals
Inference code, the data it produced, and the tools that turn that data
into tables and an HTML viewer. The evals are Anthropic's persona set,
scored in three regimes: teacher-forced logprob of the answer literal,
greedy answer with the reasoning block pre-closed, and a full 16k-budget
reasoning trace.
Pinned models
base unsloth/Qwen3.5-9B @ 005429cee5cb648998cf2b70eebdd83175989c9a
util… See the full description on the dataset page: https://huggingface.co/datasets/agentic-moral-alignment/persona-and-other-evals.grug-agentic-s3-step1903-agentic-evals-traces
