varadsrivastava/lm-playschool-qwen3.5-2b-sft-dpo-vllm
R2 (config variant for vLLM)
Part of a five-regime developmental sweep of post-training methods for dialogue-game competence (LM Playschool Challenge 2026, team DAIR).
Not a separate model. These are the same merged weights as `lm-playschool-qwen3.5-2b-sft-dpo` (R2), republished with a composite (vision-language) config.json so that vLLM will load them. At the time of our experiments, vLLM's Qwen3.5 integration expected the composite config while transformers writes a text-only one; neither could read the other's schema. Use this repo only if you need vLLM; use the R2 repo for transformers. Scores are those of R2.
All numbers are clemscore / statscore on the playpen validation split, measured in a single frozen environment (Python 3.11, clemcore pinned via playpen, clembench pinned requirements) with two upstream fixes applied: a division-by-zero guard in the privateshared Game Master and the punkt_tab NLTK resource for the IFEval scorer. Earlier revisions of this card reported numbers from an unpinned environment; see the paper for the environment-sensitivity analysis.
Checkpoint family (LM Playschool challenge, team DAIR)
Base model: Qwen3.5-2B (13.63 / 44.22 in the same environment). Paper: Raising a Small Language Model: From Imitation to Curiosity in Dialogue Games (LM Playschool Challenge 2026). <!-- TODO: add link -->
