CoolFace
Datasetpublic

violetxi/wmrl-v4-scale70-3m-agentic-eval-20t-think

scale70-3m Agentic evaluation, 20 turns, thinking enabled 1,000 saved evaluation attempts: 250 C&H firm-knowledge tasks × four samples. The base model is Qwen/Qwen3.5-9B at revision c202236235762e1c871ad0ccb60c8ee5ba337b9a, trained for two epochs on a 70% notes / 30% recall mixture. This dataset has 2,997,925 total loss-bearing training-data tokens (2,098,505 notes + 899,420 recall). The token count describes the dataset before its two training exposures. Run:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/wmrl-v4-scale70-3m-agentic-eval-20t-think.

sourceHugging Facemitupdated 12d agoView on Hugging Face
0likes55downloads

violetxi/wmrl-v4-scale70-3m-agentic-eval-20t-think · main · files are served by the source, never re-hosted here