CoolFace
Modelpublic

manisankardhabal/slm-125m-grounded-sft

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes14downloads
Model Card

slm-125m-grounded-sft

This is the primary epoch-3 supervised fine-tuned checkpoint derived from `thesreedath/slm-125m-base`. It was trained to answer from supplied context and to refuse when the answer is not present in that context.

Provenance

  • —Base model: thesreedath/slm-125m-base
  • —Published checkpoint: primary epoch-3 final model
  • —Source path before upload: /data/finetune/sft_10k_v2/full/final_model
  • —Train data: /data/finetune/sft_10k_v2/train.tokenized.jsonl
  • —Validation data: /data/finetune/sft_10k_v2/validation.tokenized.jsonl
  • —Test data: /data/finetune/sft_10k_v2/test.tokenized.jsonl

The tokenizer is the original base-model tokenizer. No vocabulary remapping or new tokenizer training was performed.

Training Summary

  • —Epochs: 3
  • —Optimizer steps: 828
  • —Initial validation loss: 3.202
  • —Final validation loss: 1.356
  • —Initial test loss: 3.328
  • —Final test loss: 1.209

Validation and test loss both improved during the primary epoch-3 SFT run. Downstream users should still evaluate this small experimental model on their own task mix before relying on it.

Deterministic Evaluation Snapshot

Evaluation report: /data/finetune/sft_10k_v2/eval/comparison_report.json

ModelTest recordsGrounded answer accuracyRefusal accuracyUnsupported claim rateRepetition rate
Base8150.0010.00.9960.913
SFT epoch 38150.1190.7620.6650.037
Published epoch 38150.1190.7620.6650.037

Recommended checkpoint in the local evaluation report: final_epoch3.

Intended Use

Use this checkpoint for small-scale experiments with context-grounded question answering, answer-not-present refusal behavior, extraction, summarization, classification, rewriting, and simple reasoning over supplied context.

This model is small and experimental. It can hallucinate, miss answers, or produce unsupported claims. Do not use it for legal, financial, medical, or other high-stakes decisions without independent verification.

Loading

python
from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "manisankardhabal/slm-125m-grounded-sft"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(repo_id)