CoolFace
Modelpublic

reimu996/OneReason-0.8B-Frontier-SFT372-RLOO-DAPO-Anchor-V2.2-W175

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes6downloads
Model Card

OneReason-0.8B Frontier SFT372 DAPO-Anchor V2.2 W175

This repository contains one LoRA policy adapter candidate for the Kuaishou LLM4Rec competition.

Training

  • —Start: Frontier SFT checkpoint-372
  • —Method: group-relative constrained policy optimization with DAPO-style asymmetric clipping and a calibrated auxiliary GT-set anchor
  • —Rollout: G=16, temperature=1.2, legal-SID prefix constraint
  • —Per logical window: 32 reward-nonflat RL groups, four RL optimizer updates, and one auxiliary anchor update
  • —LoRA rank / alpha / dropout: 64 / 64 / 0.0
  • —Completed logical windows: 175
  • —Source cursor: cycle_index=0, offset=15776 of 17016 source groups

Local selection evidence

These are selection-conditioned training diagnostics over the trailing 20 windows, not official evaluation scores.

  • —Mean reward over retained RL candidates: 0.04211719
  • —Exact candidate-slot rate: 0.498047%
  • —Groups containing at least one exact candidate: 3.906250%
  • —Mean unique SID count per 16 candidates: 14.37344

Files and evaluation status

Only README.md, adapter_config.json, and adapter_model.safetensors are published. No optimizer state, tokenizer, dataset, or recovery metadata is included. No official competition score or SOTA claim is made before formal evaluation.