CoolFace
Datasetpublic

violetxi/wmrl-v4-base9b-agentic-eval-20t-think

Base-model agentic eval transcripts, 20-turn budget (WM-RL v4) 1,000 complete agentic-evaluation transcripts of the untrained base model Qwen/Qwen3.5-9B (revision c202236235762e1c871ad0ccb60c8ee5ba337b9a), thinking enabled, on the 250 held-out firm-knowledge tasks of the WM-RL v4 study, at a 20-turn tool budget. This is the baseline every trained condition in the study is compared against; the transcripts are the raw rollouts, saved before grading. 250 tasks x 4 samples = 1,000… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/wmrl-v4-base9b-agentic-eval-20t-think.

sourceHugging Facemitupdated 12d agoView on Hugging Face
0likes46downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
violetxi/wmrl-v4-base9b-agentic-eval-20t-think · CoolFace