reimu996/OneReason-0.8B-Frontier-SFT372-RLOO-DAPO-T12-K1-W150
06
OneReason-0.8B Frontier SFT372 RLOO-DAPO - Window 150
This repository contains one intermediate LoRA adapter for the Kuaishou LLM4Rec competition.
Training
- Start: OneReason-0.8B competition base plus the Frontier SFT step-372 export
- Method: reference-free RLOO with DAPO-style dynamic group filtering
- Completed rollout windows: 150 / 1064
- Completed optimizer updates: 600 / 4256
- Effective groups per window: 32
- Candidates per prompt: 16
- Sampling temperature: 1.2
- Asymmetric ratio clip: [0.8, 1.28]
- GT anchor / reference model / KL: disabled
- LoRA rank / alpha / dropout: 64 / 64 / 0.0
Files
adapter_model.safetensors: LoRA weightsadapter_config.json: the training checkpoint adapter configuration
The adapter weight SHA256 is 21f23234ac6c2b0a08352777cca4bd63d206d77dfefd8361501b6fab1400ef2c.
Evaluation status
This intermediate adapter has not yet received an official competition score. No performance improvement is claimed by this repository.
