CoolFace
Modelpublic

reimu996/OneReason-0.8B-Frontier-SFT372-RLOO-DAPO-T12-K1-W625

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes13downloads
Model Card

OneReason-0.8B Frontier SFT372 RLOO-DAPO - Window 625

This repository contains one intermediate LoRA adapter for the Kuaishou LLM4Rec competition.

Training

  • —Start: OneReason-0.8B competition base plus the Frontier SFT step-372 export
  • —Method: reference-free RLOO with DAPO-style dynamic group filtering
  • —Completed rollout windows: 625 / 1064
  • —Completed optimizer updates: 2500 / 4256
  • —Effective groups per window: 32
  • —Candidates per prompt: 16
  • —Sampling temperature: 1.2
  • —Asymmetric ratio clip: [0.8, 1.28]
  • —GT anchor / reference model / KL: disabled
  • —LoRA rank / alpha / dropout: 64 / 64 / 0.0

Files

  • —adapter_model.safetensors: LoRA weights
  • —adapter_config.json: the training checkpoint adapter configuration

The adapter weight SHA256 is 3f5ad34fe78039d3adf3722b55549d0f18d0c02721ad69431524800888357e46.

Evaluation status

This intermediate adapter has not yet received an official competition score. No performance improvement is claimed by this repository.