CoolFace
Modelpublic

yusifnuri/Mistral-7B-v0.3_financial_sentiment

sourceHugging Faceapache-2.0updated 16d agoView on Hugging Face
0likes20downloads
Model Card

Mistral-7B-v0.3 — Financial sentiment adapter

A QLoRA (4-bit NF4, double quantisation) adapter that specialises mistralai/Mistral-7B-v0.3 (7.25 B parameters) for a single enterprise task: it classifies a financial sentence as negative, neutral or positive.

It was produced for the MSc thesis Fine-Tune or Pay Per Token? An Enterprise Benchmark of Small Language Models (SRH University Hamburg), which measures fine-tuned small models against frontier provider APIs on accuracy, latency, cost, privacy exposure and return-on-investment breakeven volume. The adapter is released so that the benchmark can be independently verified.

Read this before using the adapter

  • —The training corpus (Financial PhraseBank) is licensed CC BY-NC-SA 3.0. This adapter is a research artefact and is not commercially deployable; commercial use would require a licensed corpus or your own annotations.
  • —Adapted with QLoRA from the base release, not an instruction-tuned one. Any deficit is jointly attributable to the model and to 4-bit adaptation; the two cannot be separated within this design.

Measured performance

MetricValue
Accuracy0.845
Mean latency, batch 1562 ms
Cost per 1M generated tokensUSD 19.47

Measured on a single NVIDIA H200 (141 GB) at batch size one and full utilisation, priced at an imputed USD 3.99 per GPU-hour. Latency excludes network transit. Scores are not comparable across tasks — each task carries its own metric. Evaluation ran on 5 July 2026; the complete matrix is at `results/benchmark_matrix.csv`.

Training

MethodQLoRA (4-bit NF4, double quantisation)
DatasetFinancial PhraseBank (AllAgree) ((not redistributable))
Dataset licenceCC BY-NC-SA 3.0 — non-commercial
Training examples5,000 (500 held out for checkpoint selection)
Rank / alpha / dropout16 / 32 / 0.05
Target modulesq_proj, k_proj, v_proj, o_proj
Learning rate2e-4, cosine schedule, 3% warmup
Epochs3
Effective batch size16 (2 x 8 gradient accumulation)
Max sequence length512 tokens
OptimiserAdamW
Seed42

Hyperparameters were held constant across every model and task rather than tuned per cell, so these figures are a conservative lower bound on attainable performance.

Usage

python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = AutoModelForCausalLM.from_pretrained("mistralai/Mistral-7B-v0.3", device_map="auto")
model = PeftModel.from_pretrained(base, "<your-hf-username>/Mistral-7B-v0.3_financial_sentiment")
tokenizer = AutoTokenizer.from_pretrained("mistralai/Mistral-7B-v0.3")

The adapter was trained on this prompt format and expects it at inference:

text
Classify the sentiment of this financial sentence (negative / neutral / positive):
{text}
Sentiment:

Limitations

  • —Trained once, with a single seed. Reported differences confound model quality with initialisation variance.
  • —Specialised to one task on one public corpus. It is not a general-purpose assistant and should not be treated as one.
  • —The evaluation corpora are long-standing public benchmarks and are plausibly present in the base model's pretraining data, which inflates absolute scores.
  • —Evaluation used 200 held-out instances (all 164 problems for code generation), so detectable effect sizes are bounded at roughly ten percentage points.

Links

Citation

bibtex
@mastersthesis{nuri2026finetune,
  title  = {Fine-Tune or Pay Per Token? An Enterprise Benchmark of Small Language Models},
  author = {Nuri, Yusif},
  school = {SRH University Hamburg},
  year   = {2026}
}