CoolFace
Datasetpublic

Nyandwi/qwen3.5-9b-data-agent-subset100-eval

Qwen3.5-2B on data_agent_rl_environment_train_subset_100 pass@1: 89/100 = 89% · model Qwen/Qwen3.5-9B · temperature 0.7 · 1 rollout/task · max 10 code turns Task environments ran as Modal sandboxes (one container per task, Kaggle slice pulled from the HF bucket into /home/user/input). Rewards come from each task's own tests/grader.py with the LLM-judge tier disabled, so grading is deterministic: exact string match, else numeric match within 1e-3. Where the… See the full description on the dataset page: https://huggingface.co/datasets/Nyandwi/qwen3.5-9b-data-agent-subset100-eval.

sourceHugging Faceapache-2.0updated 4d agoView on Hugging Face
0likes20downloads
5 commits on main
25005894d ago

Upload trajectories.jsonl with huggingface_hub

Nyandwi
fa8aab94d ago

Upload config.json with huggingface_hub

Nyandwi
270864c4d ago

Upload results.parquet with huggingface_hub

Nyandwi
f232e384d ago

Upload README.md with huggingface_hub

Nyandwi
8bc914b4d ago

initial commit

Nyandwi