reimu996/OneReason-0.8B-Frontier-SFT372-RLOO-DAPO-Anchor-V2.2-W249
05
OneReason-0.8B Frontier SFT372 DAPO-Anchor V2.2 W249
This repository contains one LoRA policy adapter candidate for the Kuaishou LLM4Rec competition.
Training
- Start: Frontier SFT checkpoint-372
- Method: reference-free RLOO, DAPO asymmetric clipping, calibrated auxiliary GT-set anchor
- Rollout: G=16, temperature=1.2, legal-SID prefix constraint
- Per logical window: 32 RL groups, four RL updates, one auxiliary anchor update
- LoRA rank / alpha / dropout: 64 / 64 / 0.0
- Completed logical windows: 249
- Source cursor: cycle_index=1, offset=8480
Local selection evidence
Metrics below are selection-conditioned training diagnostics, not official evaluation scores.
- Trailing 20 windows: reward=0.06000488, exact-slot=1.191406%, exact-group=7.343750%, unique-SID/16=11.93594, source-to-RL=22.857143%
- Trailing 40 windows: exact-slot=1.083984%, exact-group=6.406250%, unique-SID/16=12.21328
Files and evaluation status
Only README.md, adapter_config.json, and adapter_model.safetensors are published. No optimizer state, tokenizer, dataset, or recovery metadata is included. No official competition score or SOTA claim is made before formal evaluation.
