CoolFace
Modelpublic

3l3ktr4/donorsim-qwen3-8b-abstract-step40

sourceHugging Faceupdated 29d agoView on Hugging Face
0likes616downloads
Model Card

donorsim-qwen3-8b-abstract-step40

Qwen3-8B fine-tuned with GRPO (verl 0.7.1, LoRA r16/alpha32 merged into bf16 weights) on the abstract / everyday-scene stage of the iterated Donor's Game: the model sees short naturalistic situations about named people in a small group (no payoff numbers, no "cooperate"/"defect" vocabulary) and answers CHOICE: 1 / CHOICE: 2, with the option order re-randomised every turn. Rewards are Term 1 (normalised payoff) + Term 2 (reciprocity); no group / CFE term. Partners rotate within a roster of n_players - 1 members with per-partner memory; re-encounter probability w and gossip probability q are phrased in words each turn.

40 abstract steps on top of donorsim-qwen3-8b-modeAB-step75 (75 steps of the structured group game). Run qwen3_8b_recipT2_abstract_20steps_20260829_033643, stage_abstract/global_step_40 (steps 1-20 job 3213330, 21-50 job 3216156, 2 nodes / 8 rollout replicas).

Merged full weights, loadable directly with transformers or vLLM (no adapter needed).