CoolFace
Datasetpublic

Nyandwi/qwen3.5-2b-data-agent-subset100-eval

Qwen3.5-2B on data_agent_rl_environment_train_subset_100 pass@1: 64/100 = 64% · model Qwen/Qwen3.5-2B · temperature 0.7 · 1 rollout/task · max 10 code turns Task environments ran as Modal sandboxes (one container per task, Kaggle slice pulled from the HF bucket into /home/user/input). Rewards come from each task's own tests/grader.py with the LLM-judge tier disabled, so grading is deterministic: exact string match, else numeric match within 1e-3. Where the… See the full description on the dataset page: https://huggingface.co/datasets/Nyandwi/qwen3.5-2b-data-agent-subset100-eval.

sourceHugging Faceapache-2.0updated 4d agoView on Hugging Face
0likes39downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Nyandwi/qwen3.5-2b-data-agent-subset100-eval · CoolFace