CoolFace
Modelpublic

Lambent/ScenarioPlayTest2-100Steps-Qwen3-4B-adapter

sourceHugging Faceapache-2.0updated 9mo agoView on Hugging Face
0likes8downloads
Model Card

QLoRA adapter for Qwen3-4b playtesting the second draft of an RLVR environment of Mira's conceptualization.

Focus on one-shot roleplaying scenarios, even division of silly and serious, both narrative and problem-solving.

100 steps, cosine decay, batch size 4, learning rate 1e-5, rank 128, alpha 128.