yujackein/onereason-8b-lora-item32k-user75-rec50-worldclean1601-all1-lr2e4-r32a32-step646
07
OneReason-8B LoRA: Item32K / User75 / Rec50 / WorldClean1601 / LR2e-4 / Step646
This is an experimental LoRA adapter for `OpenOneRec/OneReason-8B-pretrain-competition`, trained for the OneReason recommendation competition.
This repository contains the final checkpoint at step 646 from a two-pass training schedule. The matching intermediate one-pass state is published separately as the step-323 adapter.
Data recipe
All assistant loss weights were normalized to 1.0. The 32,768-token packing plan covers 99.1727% of rendered tokens in one safe pass.
Training
- Base model: OneReason-8B pretrain competition release
- Method: LoRA (
r=32,alpha=32, dropout0.05) - Target modules:
q_proj,k_proj,v_proj,o_proj,gate_proj,up_proj,down_proj - Precision: BF16
- Context length: 32,768
- Global batch size: 16 on 4 x NVIDIA A800 80GB
- Learning rate:
2e-4, cosine schedule, 3% warmup - Training length: 646 optimizer steps (two packing passes)
- Final recorded training loss:
0.9556 - Seed: 42
Usage
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base_id = "OpenOneRec/OneReason-8B-pretrain-competition"
adapter_id = "yujackein/onereason-8b-lora-item32k-user75-rec50-worldclean1601-all1-lr2e4-r32a32-step646"
tokenizer = AutoTokenizer.from_pretrained(adapter_id, trust_remote_code=True)
base_model = AutoModelForCausalLM.from_pretrained(
base_id,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True,
)
model = PeftModel.from_pretrained(base_model, adapter_id)Status
Training completed successfully on 2026-08-12. Formal evaluation results have not yet been added.
