CoolFace
Modelpublic

yusifnuri/phi-4-mini-instruct_classification

sourceHugging Facemitupdated 11d agoView on Hugging Face
0likes15downloads
Model Card

phi-4-mini-instruct — Topic classification adapter

A LoRA adapter that specialises microsoft/Phi-4-mini-instruct (3.80 B parameters) for a single enterprise task: it assigns a news item to one of four topics (World, Sports, Business, Sci/Tech).

It was produced for the MSc thesis Fine-Tune or Pay Per Token? An Enterprise Benchmark of Small Language Models (SRH University Hamburg), which measures fine-tuned small models against frontier provider APIs on accuracy, latency, cost, privacy exposure and return-on-investment breakeven volume. The adapter is released so that the benchmark can be independently verified.

Measured performance

MetricValue
Accuracy0.92
Mean latency, batch 1162 ms
Cost per 1M generated tokensUSD 5.60

Measured on a single NVIDIA H200 (141 GB) at batch size one and full utilisation, priced at an imputed USD 3.99 per GPU-hour. Latency excludes network transit. Scores are not comparable across tasks — each task carries its own metric. Evaluation ran on 5 July 2026; the complete matrix is at `results/benchmark_matrix.csv`.

Training

MethodLoRA
DatasetAG News (fancyzhx/ag_news)
Dataset licenceCustom, research use
Training examples5,000 (500 held out for checkpoint selection)
Rank / alpha / dropout16 / 32 / 0.05
Target modulesq_proj, k_proj, v_proj, o_proj
Learning rate2e-4, cosine schedule, 3% warmup
Epochs3
Effective batch size16 (4 x 4 gradient accumulation)
Max sequence length512 tokens
OptimiserAdamW
Seed42

Hyperparameters were held constant across every model and task rather than tuned per cell, so these figures are a conservative lower bound on attainable performance.

Usage

python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = AutoModelForCausalLM.from_pretrained("microsoft/Phi-4-mini-instruct", device_map="auto")
model = PeftModel.from_pretrained(base, "<your-hf-username>/phi-4-mini-instruct_classification")
tokenizer = AutoTokenizer.from_pretrained("microsoft/Phi-4-mini-instruct")

The adapter was trained on this prompt format and expects it at inference:

text
Classify the following news text into exactly one category (World / Sports / Business / Technology):
{text}
Category:

Limitations

  • Trained once, with a single seed. Reported differences confound model quality with initialisation variance.
  • Specialised to one task on one public corpus. It is not a general-purpose assistant and should not be treated as one.
  • The evaluation corpora are long-standing public benchmarks and are plausibly present in the base model's pretraining data, which inflates absolute scores.
  • Evaluation used 200 held-out instances (all 164 problems for code generation), so detectable effect sizes are bounded at roughly ten percentage points.

Links

Citation

bibtex
@mastersthesis{nuri2026finetune,
  title  = {Fine-Tune or Pay Per Token? An Enterprise Benchmark of Small Language Models},
  author = {Nuri, Yusif},
  school = {SRH University Hamburg},
  year   = {2026}
}