reimu996/OneReason-0.8B-Frontier-SFT372-RLOO-DAPO-Anchor-V2.2-W175
06
OneReason-0.8B Frontier SFT372 DAPO-Anchor V2.2 W175
This repository contains one LoRA policy adapter candidate for the Kuaishou LLM4Rec competition.
Training
- Start: Frontier SFT checkpoint-372
- Method: group-relative constrained policy optimization with DAPO-style asymmetric clipping and a calibrated auxiliary GT-set anchor
- Rollout: G=16, temperature=1.2, legal-SID prefix constraint
- Per logical window: 32 reward-nonflat RL groups, four RL optimizer updates, and one auxiliary anchor update
- LoRA rank / alpha / dropout: 64 / 64 / 0.0
- Completed logical windows: 175
- Source cursor: cycle_index=0, offset=15776 of 17016 source groups
Local selection evidence
These are selection-conditioned training diagnostics over the trailing 20 windows, not official evaluation scores.
- Mean reward over retained RL candidates: 0.04211719
- Exact candidate-slot rate: 0.498047%
- Groups containing at least one exact candidate: 3.906250%
- Mean unique SID count per 16 candidates: 14.37344
Files and evaluation status
Only README.md, adapter_config.json, and adapter_model.safetensors are published. No optimizer state, tokenizer, dataset, or recovery metadata is included. No official competition score or SOTA claim is made before formal evaluation.
