CoolFace
Datasetpublic

violetxi/wmrl-v4-scale70-3m-agentic-eval-20t-think

scale70-3m Agentic evaluation, 20 turns, thinking enabled 1,000 saved evaluation attempts: 250 C&H firm-knowledge tasks × four samples. The base model is Qwen/Qwen3.5-9B at revision c202236235762e1c871ad0ccb60c8ee5ba337b9a, trained for two epochs on a 70% notes / 30% recall mixture. This dataset has 2,997,925 total loss-bearing training-data tokens (2,098,505 notes + 899,420 recall). The token count describes the dataset before its two training exposures. Run:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/wmrl-v4-scale70-3m-agentic-eval-20t-think.

sourceHugging Facemitupdated 13d agoView on Hugging Face
0likes64downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
violetxi/wmrl-v4-scale70-3m-agentic-eval-20t-think · CoolFace