CoolFace
Modelpublic

Nishef/SmolLM2-360M-Full_KNOWLEDGE_RETAINING_ENHANCED_KTO_20251227_151509-merged

sourceHugging Faceapache-2.0updated 9mo agoView on Hugging Face
0likes32downloads
Model Card

SmolLM2-360M - Knowledge-Retaining-Enhanced-KTO (Merged)

This is the fully merged standalone version. No adapter loading required!

<div align="center"> <img src="thesisplots/prospecttheory_visualization.png" alt="Prospect Theory" width="600"/> </div>

๐ŸŽฏ Method Overview

This model was fine-tuned using Knowledge-Retaining-Enhanced-KTO combining:

  1. 1.Kahneman-Tversky Prospect Theory - Asymmetric value functions
  2. 2.KL Divergence Preservation - Maintains base model knowledge
  3. 3.Binary Feedback Optimization - Simple desirable/undesirable labels

๐Ÿ“Š Training Results

MetricValue
Improvement18.7%
Training Steps416
Base ModelHuggingFaceTB/SmolLM2-360M

<div align="center"> <img src="thesisplots/trainingloss_curve.png" alt="Training Loss" width="700"/> </div>

๐Ÿš€ Quick Start

python
from transformers import AutoModelForCausalLM, AutoTokenizer

# Direct loading - no PEFT needed!
model = AutoModelForCausalLM.from_pretrained("Nishef/SmolLM2-360M-Full_KNOWLEDGE_RETAINING_ENHANCED_KTO_20251227_151509-merged")
tokenizer = AutoTokenizer.from_pretrained("Nishef/SmolLM2-360M-Full_KNOWLEDGE_RETAINING_ENHANCED_KTO_20251227_151509-merged")

prompt = "What is machine learning?"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=100)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

๐Ÿ“ˆ Method Comparison

<div align="center"> <img src="thesisplots/methodcomparison.png" alt="Method Comparison" width="700"/> </div>

๐Ÿ“‹ Also Available

๐Ÿ“Š Benchmark Results

Performance Comparison

MethodHellaSwagTruthfulQAMMLUAverage
DPO0.5500.3610.2640.392
ORPO0.5260.3730.2490.383
Enhanced KTO0.4960.3900.2890.392
Standard KTO0.3940.4740.2540.374
Knowledge-Retaining-Enhanced-KTO0.3920.4500.2440.362

Key Findings

๐ŸŽฏ TruthfulQA Excellence: Our method achieves 0.450 accuracy on TruthfulQA, significantly outperforming DPO (0.361) and ORPO (0.373). This demonstrates the effectiveness of Prospect Theory's loss aversion in promoting truthful outputs.

๐Ÿ“ˆ Comparison with Standard KTO: Knowledge-Retaining-Enhanced-KTO maintains similar TruthfulQA performance (0.450 vs 0.474) while providing more stable training dynamics.

<div align="center"> <img src="thesisplots/benchmarkcomparison.png" alt="Benchmark Comparison" width="700"/> </div>

Radar Chart Comparison

<div align="center"> <img src="thesisplots/benchmarkradar.png" alt="Radar Chart" width="500"/> </div>

TruthfulQA Performance

<div align="center"> <img src="thesisplots/truthfulqacomparison.png" alt="TruthfulQA" width="600"/> </div>


<div align="center"> <b>๐ŸŽ“ Part of MSc Thesis on LLM Alignment Methods</b> </div>