CoolFace
Modelpublic

reimu996/OneReason-0.8B-Frontier-SFT372-RLOO-DAPO-Anchor-V2.2-W249

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes5downloads
Model Card

OneReason-0.8B Frontier SFT372 DAPO-Anchor V2.2 W249

This repository contains one LoRA policy adapter candidate for the Kuaishou LLM4Rec competition.

Training

  • —Start: Frontier SFT checkpoint-372
  • —Method: reference-free RLOO, DAPO asymmetric clipping, calibrated auxiliary GT-set anchor
  • —Rollout: G=16, temperature=1.2, legal-SID prefix constraint
  • —Per logical window: 32 RL groups, four RL updates, one auxiliary anchor update
  • —LoRA rank / alpha / dropout: 64 / 64 / 0.0
  • —Completed logical windows: 249
  • —Source cursor: cycle_index=1, offset=8480

Local selection evidence

Metrics below are selection-conditioned training diagnostics, not official evaluation scores.

  • —Trailing 20 windows: reward=0.06000488, exact-slot=1.191406%, exact-group=7.343750%, unique-SID/16=11.93594, source-to-RL=22.857143%
  • —Trailing 40 windows: exact-slot=1.083984%, exact-group=6.406250%, unique-SID/16=12.21328

Files and evaluation status

Only README.md, adapter_config.json, and adapter_model.safetensors are published. No optimizer state, tokenizer, dataset, or recovery metadata is included. No official competition score or SOTA claim is made before formal evaluation.