reimu996/OneReason-0.8B-Frontier-SFT372-RLOO-DAPO-T12-K1-W532
08
OneReason-0.8B Frontier SFT372 RLOO-DAPO - Window 532
This repository contains the first effective-epoch LoRA adapter for the Kuaishou LLM4Rec competition.
Training
- Start: OneReason-0.8B competition base plus the Frontier SFT step-372 export
- Method: reference-free RLOO with DAPO-style dynamic group filtering
- Completed rollout windows: 532 / 1064
- Completed optimizer updates: 2128 / 4256
- Effective epoch: 1 / 2 (count-based, not one literal source-data traversal)
- Effective groups per window: 32
- Candidates per prompt: 16
- Sampling temperature: 1.2
- Asymmetric ratio clip: [0.8, 1.28]
- GT anchor / reference model / KL: disabled
- LoRA rank / alpha / dropout: 64 / 64 / 0.0
Files
adapter_model.safetensors: LoRA weightsadapter_config.json: the training checkpoint adapter configuration
The adapter weight SHA256 is 6d839e79e401639b98ef16073758d9c8111c48fcde1191e94c088086b98ac28c.
Evaluation status
This adapter has not yet received an official competition score. No performance improvement is claimed by this repository.
