CoolFace
Modelpublic

shareit/Supervisor-FRPT-Phi-4-reasoning-BestSeed

sourceHugging Facemitupdated 3mo agoView on Hugging Face
0likes5downloads
Model Card

Supervisor-FRPT — Supervisor-FRPT-Phi-4-reasoning-BestSeed

This model is a LoRA-fine-tuned supervisor (CS quality evaluator) for electronics customer-support chatbot conversations. It was trained on 20260331_HumanFeedBack_selfdist.jsonl (3,771 human-labelled dialogues) with the FRPT ("Fact-Reasoning Process Training") research training methodology applied to a lora_sequential LoRA recipe.

The job of this model: given (category, multi-turn user/assistant transcript, retrieved reference document), produce a Korean <think>...</think> rubric chain and a JSON verdict {"label": "correct|incorrect", "reason": "..."}.

Test metrics (held-out 199 dialogues)

MetricValue
Accuracy0.709
Macro-F10.666
F1 (correct)0.547
F1 (incorrect)0.785
Precision (incorrect)0.862
Recall (incorrect)0.721
F₀.₅ (incorrect)0.829
Precision (correct)0.461
Recall (correct)0.673
Unparsed0/199

Why these metrics

The deployment goal for this supervisor is catching incorrect chatbot responses with high reliability, so the operationally critical metric is the chance that when this model says "incorrect", the chatbot really did answer incorrectly — i.e. precision(incorrect). The model was selected from a multi-method, multi-seed grid by F₀.₅(incorrect) = (1 + 0.25)·P·R / (0.25·P + R), which weighs precision twice as much as recall on the incorrect class while still penalising excessive misses.

Training methodology — research highlights

The training methodology bundles two layers:

  1. 1.Base LoRA recipe — lora_sequential with rank 16, alpha 32, dropout 0.05, target modules qkv_proj, o_proj, down_proj, gate_up_proj (Phi-3 family) or the q/k/v/o/MLP equivalents for Gemma-4. Optimizer AdamW, cosine schedule, warmup ratio 0.05, grad clip 1.0, BF16, SDPA attention.
  1. 1.FRPT-aware data shaping (Fact-grounded Reasoning Process Training):
  2. 2.Process-supervision view — the assistant turn already exposes a 3-axis rubric (Query-Document Alignment, Response-Document Consistency, Response Completeness) inside <think>...</think>. We train the entire assistant response, so the model learns the reasoning process, not just the verdict.
  3. 3.Fact-grounded SFT — loss is masked on user/system tokens; only the assistant span (think + JSON) contributes to gradient. This forces the model to learn how to evaluate, not what the user said.
  4. 4.Class-imbalance aware — incorrect : correct = 2616 : 1155 (~2.3:1) in train. We monitor F1-correct (the minority class) as the primary model-selection signal.
  5. 5.(Sequential variant) — lora_sequential groups the 33 product categories into 5 buckets (DRW, TV, SBS, REF_AUD_MNT, OTHERS) and trains them in order, exposing the model to per-category structure while sharing one adapter across the curriculum.

Hyperparameters of the final run

FieldValue
Base modelmicrosoft/Phi-4-reasoning
Methodlora_sequential
LoRA rank16
LoRA alpha32
LoRA dropout0.05
Learning rate0.0005
Epochs7
Seed42
Train samples3,771
Test samples199
Max sequence length4096

Quick inference

python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

mid = "shareit/Supervisor-FRPT-Phi-4-reasoning-BestSeed"
tok = AutoTokenizer.from_pretrained(mid, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(mid, dtype=torch.bfloat16,
                                            device_map="auto",
                                            trust_remote_code=True)

system = "당신은 전자제품 CS 챗봇의 품질을 평가하는 수퍼바이저입니다."
user = ("[Category] PC\n\n[Conversation Transcript]\n"
        "Turn 1 - User: ...\nTurn 1 - Assistant: ...\n\n"
        "[Retrieved Document]\n(title) ...\n(content) ...")

msgs = [{"role": "system", "content": system},
        {"role": "user", "content": user}]
inp = tok.apply_chat_template(msgs, tokenize=True, add_generation_prompt=True,
                              return_tensors="pt").to(model.device)
out = model.generate(inp, max_new_tokens=900, do_sample=False)
print(tok.decode(out[0, inp.shape[1]:], skip_special_tokens=True))

The generated text follows:

<think>
[Query-Document Alignment] ...
[Response-Document Consistency] ...
[Response Completeness] ...
</think>
{"label": "correct", "reason": "..."}

Citation / theory

This model embodies the FRPT (Fact-Reasoning Process Training) research program. Key references that inform the methodology:

  • —Gekhman et al. 2024 — fine-tuning new facts can encourage hallucination.
  • —Lightman et al. 2023 — Let's Verify Step by Step (process supervision).
  • —Hu et al. 2021 — LoRA.
  • —Dettmers et al. 2023 — QLoRA.
  • —LoRA Learns Less and Forgets Less (Biderman et al.) — PEFT/FullFT tradeoffs.

For the merge-before-forget continual-learning theory that motivated the sequential variant, see the internal Session 1~4 reports.