CoolFace
Modelpublic

reimu996/OneReason-0.8B-Frontier-SFT372-RLOO-DAPO-T12-K1-W532

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes8downloads
Model Card

OneReason-0.8B Frontier SFT372 RLOO-DAPO - Window 532

This repository contains the first effective-epoch LoRA adapter for the Kuaishou LLM4Rec competition.

Training

  • —Start: OneReason-0.8B competition base plus the Frontier SFT step-372 export
  • —Method: reference-free RLOO with DAPO-style dynamic group filtering
  • —Completed rollout windows: 532 / 1064
  • —Completed optimizer updates: 2128 / 4256
  • —Effective epoch: 1 / 2 (count-based, not one literal source-data traversal)
  • —Effective groups per window: 32
  • —Candidates per prompt: 16
  • —Sampling temperature: 1.2
  • —Asymmetric ratio clip: [0.8, 1.28]
  • —GT anchor / reference model / KL: disabled
  • —LoRA rank / alpha / dropout: 64 / 64 / 0.0

Files

  • —adapter_model.safetensors: LoRA weights
  • —adapter_config.json: the training checkpoint adapter configuration

The adapter weight SHA256 is 6d839e79e401639b98ef16073758d9c8111c48fcde1191e94c088086b98ac28c.

Evaluation status

This adapter has not yet received an official competition score. No performance improvement is claimed by this repository.