Nyandwi/qwen3.5-9b-data-agent-subset100-eval
Qwen3.5-2B on data_agent_rl_environment_train_subset_100 pass@1: 89/100 = 89% · model Qwen/Qwen3.5-9B · temperature 0.7 · 1 rollout/task · max 10 code turns Task environments ran as Modal sandboxes (one container per task, Kaggle slice pulled from the HF bucket into /home/user/input). Rewards come from each task's own tests/grader.py with the LLM-judge tier disabled, so grading is deterministic: exact string match, else numeric match within 1e-3. Where the… See the full description on the dataset page: https://huggingface.co/datasets/Nyandwi/qwen3.5-9b-data-agent-subset100-eval.
020
