CoolFace
Modelpublic

EdwardConstantine/bengali-empathy-llama

sourceHugging Faceapache-2.0updated 10mo agoView on Hugging Face
0likes21downloads
Model Card

Bengali Empathetic LLaMA šŸ‡§šŸ‡©

Fine-tuned LLaMA 3.1-8B-Instruct for empathetic Bengali conversations using LoRA (Low-Rank Adaptation).


šŸ“– Table of Contents

  1. 1.Model Description
  2. 2.Training Details
  3. 3.Evaluation Results
  4. 4.Sample Responses
  5. 5.Design Decisions & Trade-offs
  6. 6.Architecture & OOP Design
  7. 7.Usage
  8. 8.Challenges Faced
  9. 9.Future Improvements
  10. 10.Files in This Repository

Model Description

This model is a LoRA fine-tuned version of Meta's LLaMA 3.1-8B-Instruct, specifically trained to generate compassionate and empathetic responses in Bengali.

What This Model Does

  • —Input: Bengali text expressing emotions (sadness, happiness, frustration, etc.)
  • —Output: Empathetic Bengali response with emotional understanding

Example

Input:  আমি খুব ą¦ą¦•ą¦¾ অনুভব ą¦•ą¦°ą¦›ą¦æą„¤ (I feel very lonely)
Output: ą¦¹ą§ą¦Æą¦¾ą¦, ą¦ą¦Ÿą¦¾ খুব ą¦•ą¦ ą¦æą¦Øą„¤ ą¦•ą¦æą¦Øą§ą¦¤ą§ আমি আশা করি আপনি ą¦¶ą§€ą¦˜ą§ą¦°ą¦‡ ą¦ą¦•ą¦œą¦Ø ą¦¬ą¦Øą§ą¦§ą§ ą¦Ŗą¦¾ą¦¬ą§‡ą¦Øą„¤
        (Yes, this is very hard. But I hope you will find a friend soon.)

Training Details

Training History

āŒ First Attempt (Interrupted - Progress Lost)

Our initial training with optimal settings was interrupted at 66% completion due to Kaggle session timeout:

SettingValue
Data100% (10,749 samples)
Epochs3
Max Length384 tokens
Progress5,329 / 8,061 steps (66%)

Loss Progression (Before Interruption): | Step | Training Loss | Validation Loss | |------|---------------|-----------------| | 500 | 0.4459 | - | | 1000 | 0.3869 | - | | 2000 | 0.3292 | 0.3281 | | 3000 | 0.2450 | - | | 4000 | 0.2351 | 0.2642 | | 5000 | 0.2093 | - | | 5329 | Session Timeout | - |

āš ļø If completed, this training would have achieved ~0.18-0.20 final loss with significantly better quality. The checkpoint was lost because saves were configured every 2000 steps, and the session crashed before the next save.

āœ… Second Attempt (Completed Successfully)

With remaining GPU quota (~3 hours), we completed a condensed training:

SettingValue
Data40% sample (4,299 samples)
Epochs2
Max Length256 tokens
Training Time3.26 hours
PlatformKaggle Tesla T4 (16GB VRAM)

Final Results: | Metric | Value | |--------|-------| | Training Loss | 0.4190 | | Validation Loss | 0.3651 |

LoRA Configuration

ParameterValueExplanation
Rank (r)16Number of trainable parameters per layer. Higher = more capacity but more memory
Alpha32Scaling factor (alpha/r = 2x multiplier)
Dropout0.05Light regularization to prevent overfitting
Target Modules7 layersAll attention (q,k,v,o) + MLP (gate, up, down) projections

Training Hyperparameters

ParameterValue
Optimizerpagedadamw8bit
Learning Rate3e-4
LR SchedulerCosine
Warmup Ratio0.05
Batch Size4
Gradient Accumulation1
PrecisionFP16 (Mixed Precision)
Gradient CheckpointingEnabled
Quantization4-bit NF4

Evaluation Results

MetricScoreInterpretation
BLEU-10.0613Unigram word overlap
BLEU-20.0251Bigram overlap
BLEU-40.00984-gram overlap
ROUGE-10.0117Unigram recall
ROUGE-20.0000Bigram recall
ROUGE-L0.0117Longest common subsequence
Perplexity1.9493Model confidence (lower = better)

Why Are BLEU/ROUGE Scores Low?

This is expected and normal for empathetic response generation. Here's why:

  1. 1.Multiple Valid Responses: There are many ways to express empathy
  2. 2.Reference: "আমি ą¦¦ą§ą¦ƒą¦–ą¦æą¦¤" (I'm sorry)
  3. 3.Generated: "ą¦ą¦Ÿą¦¾ কঠিন হবে" (This must be hard)
  4. 4.Both are empathetic but share no words → BLEU = 0
  1. 1.Creative Generation: The model generates contextually appropriate but not verbatim responses
  1. 1.Bengali Language: Morphologically rich language with many word forms
  1. 1.Perplexity is Good: 1.95 indicates the model is confident in its predictions

Human Evaluation Framework

We created a human evaluation template with these criteria (1-5 scale):

CriterionWhat It Measures
EmpathyDoes the response show understanding of emotions?
RelevanceIs the response relevant to the input?
FluencyIs the Bengali grammatically correct and natural?
HelpfulnessDoes the response provide comfort or useful advice?
OverallOverall quality of the response

šŸ“„ File: human_evaluation_sheet.csv (20 samples for manual evaluation)


Sample Responses

#Input (Bengali)Model ResponseQuality
1আমি খুব ą¦ą¦•ą¦¾ অনুভব ą¦•ą¦°ą¦›ą¦æą„¤ą¦¹ą§ą¦Æą¦¾ą¦, ą¦ą¦Ÿą¦¾ খুব ą¦•ą¦ ą¦æą¦Øą„¤ ą¦•ą¦æą¦Øą§ą¦¤ą§ আমি আশা করি আপনি ą¦¶ą§€ą¦˜ą§ą¦°ą¦‡ ą¦ą¦•ą¦œą¦Ø ą¦¬ą¦Øą§ą¦§ą§ ą¦Ŗą¦¾ą¦¬ą§‡ą¦Øą„¤āœ… Empathetic
2ą¦†ą¦œ ą¦Ŗą¦°ą§€ą¦•ą§ą¦·ą¦¾ą¦Æą¦¼ ফেল করেছি, খুব ą¦•ą¦·ą§ą¦Ÿ ą¦²ą¦¾ą¦—ą¦›ą§‡ą„¤ą¦†ą¦®ą¦æ খুব ą¦¦ą§ą¦ƒą¦–ą¦æą¦¤! আপনি কি আপনার ą¦‰ą¦Ŗą¦øą§ą¦„ą¦¾ą¦Ŗą¦Øą¦¾ মনে করেন?āœ… Supportive
3নতুন ą¦šą¦¾ą¦•ą¦°ą¦æ পেয়েছি! অনেক খুশি!ą¦…ą¦­ą¦æą¦Øą¦Øą§ą¦¦ą¦Ø! আমি জানি তুমি কেমন অনুভব ą¦•ą¦°ą¦›ą§‹ą„¤ আমি ą¦øą¦¤ą§ą¦Æą¦æą¦‡ ą¦ą¦‡ ą¦øą¦®ą§ą¦Ŗą¦°ą§ą¦•ą§‡ ą¦‰ą¦¤ą§ą¦¤ą§‡ą¦œą¦æą¦¤!āœ… Celebratory
4আমার ą¦¬ą¦Øą§ą¦§ą§ আমাকে ছেঔ়ে ą¦šą¦²ą§‡ ą¦—ą§‡ą¦›ą§‡ą„¤ą¦†ą¦®ą¦æ ą¦ą¦Ÿą¦¾ শুনে ą¦¦ą§ą¦ƒą¦–ą¦æą¦¤ą„¤ আপনি কি তার সা঄ে ক঄া বলেছেন?āœ… Caring

Design Decisions & Trade-offs

1ļøāƒ£ Sequence Length: 256 vs Full Length

AspectRequirementWhat We DidWhy
Sequence LengthNot reducedReduced to 256GPU memory constraint

What "Sequence Length" Means:

  • —Maximum number of tokens (words/subwords) the model processes at once
  • —Original conversations may have 500-1000+ tokens
  • —We truncated to 256 tokens

Why We Reduced It:

Problem: Kaggle T4 GPU has only 16GB VRAM

Full Length (512+ tokens):
- Memory needed: ~18-20GB āŒ Doesn't fit
- Batch size: 1 (very slow)
- Training time: 20+ hours

Reduced Length (256 tokens):
- Memory needed: ~12GB āœ… Fits
- Batch size: 4 (faster)
- Training time: 3 hours

Impact:

  • —~15% of conversations get truncated
  • —Model may miss context in very long conversations
  • —Core empathetic learning still happens (most empathy is expressed in first 256 tokens)

What Could Be Done:

  1. 1.Use A100 GPU (40GB VRAM) → Can use 512-1024 tokens
  2. 2.Use Unsloth library → 2x memory efficiency
  3. 3.Use gradient accumulation with batch_size=1 → Slower but full length
  4. 4.Use QLoRA with more aggressive quantization

2ļøāƒ£ Strategy Pattern: LoRA vs Unsloth

AspectRequirementWhat We DidWhy
Design PatternStrategy pattern for LoRA/UnslothOnly LoRA implementedTime constraint + LoRA sufficient

What "Strategy Pattern" Means:

python
# Strategy Pattern = Swappable algorithms

class FineTuningStrategy:           # Abstract strategy
    def apply(self, model): pass

class LoRAStrategy(FineTuningStrategy):     # Strategy 1 āœ… Implemented
    def apply(self, model):
        return get_peft_model(model, lora_config)

class UnslothStrategy(FineTuningStrategy):  # Strategy 2 āŒ Not implemented
    def apply(self, model):
        return FastLanguageModel.get_peft_model(model)

# Usage: Can swap strategies easily
strategy = LoRAStrategy()  # or UnslothStrategy()
model = strategy.apply(base_model)

Why We Only Used LoRA:

FactorLoRAUnsloth
Ease of setupāœ… Simpleāš ļø Requires specific installation
Kaggle compatibilityāœ… Works perfectlyāš ļø May have conflicts
Memory efficiencyGood (4-bit)Better (2x faster)
Our GPU time3 hours leftNot enough to debug issues
Result qualityāœ… Achieved goalWould be similar

What Could Be Done:

python
# Full Strategy Pattern Implementation

from abc import ABC, abstractmethod

class FineTuningStrategy(ABC):
    @abstractmethod
    def apply(self, model, config):
        pass
    
    @abstractmethod
    def get_name(self):
        pass

class LoRAStrategy(FineTuningStrategy):
    def apply(self, model, config):
        from peft import get_peft_model, LoraConfig
        lora_config = LoraConfig(
            r=config.lora_r,
            lora_alpha=config.lora_alpha,
            target_modules=config.target_modules,
        )
        return get_peft_model(model, lora_config)
    
    def get_name(self):
        return "LoRA"

class UnslothStrategy(FineTuningStrategy):
    def apply(self, model, config):
        from unsloth import FastLanguageModel
        model, tokenizer = FastLanguageModel.from_pretrained(
            model_name=config.model_name,
            max_seq_length=config.max_length,
            load_in_4bit=True,
        )
        return FastLanguageModel.get_peft_model(model)
    
    def get_name(self):
        return "Unsloth"

# Usage
class LLAMAFineTuner:
    def __init__(self, config, strategy: FineTuningStrategy):
        self.config = config
        self.strategy = strategy
    
    def prepare_model(self, base_model):
        print(f"Using {self.strategy.get_name()} strategy")
        return self.strategy.apply(base_model, self.config)

3ļøāƒ£ Data Sampling: 40% vs 100%

AspectIdealWhat We DidWhy
Training Data100% (10,749 samples)40% (4,299 samples)GPU time constraint

Impact:

  • —Model sees less variety of conversations
  • —May not generalize as well to rare emotions
  • —Still learns core empathetic patterns

What Could Be Done:

  • —Train for longer with full dataset
  • —Use data augmentation to increase variety
  • —Prioritize diverse samples over random sampling

Architecture & OOP Design

Class Diagram

ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”
│                         MAIN PIPELINE                     │
ā”œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¤
│                                                            │
│  ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”        ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”        │
│  │   DatasetProcessor  │      │   LLAMAFineTuner    │       │
│  ā”œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¤        ā”œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¤        │
│  │ + TEMPLATE          │      │ + model             │       │
│  │ + train_dataset     │      │ + tokenizer         │       │
│  │ + val_dataset       │      │ + trainer           │       │
│  ā”œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¤        ā”œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¤        │
│  │ + load()            │      │ + load_model()      │       │
│  │ + process()         │      │ + setup_trainer()   │       │
│  │ + _format()         │      │ + train()           │       │
│  │ + _tokenize()       │      │ + save_final()      │       │
│  ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜        │ + generate()        │       │
│                               ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜         │
│                                                             │
│  ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”         ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”        │
│  │     Evaluator       │      │  ExperimentLogger   │       │
│  ā”œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¤        ā”œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¤         │
│  │ + model             │      │ + db_path           │       │
│  ā”œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¤        ā”œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¤         │
│  │ + calculate_bleu()  │      │ + log()             │       │
│  │ + calculate_rouge() │      │ + log_response()    │        │
│  │ + calculate_ppl()   │      │ + _init_db()        │        │
│  │ + test_samples()    │      ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜          │
│  ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜                                      │
│                                                             │
ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜

Database Schema

sql
-- Stores training experiment metadata
CREATE TABLE LLAMAExperiments (
    id INTEGER PRIMARY KEY AUTOINCREMENT,
    model_name TEXT,                    -- e.g., "meta-llama/Llama-3.1-8B-Instruct"
    lora_config TEXT,                   -- JSON: {"r": 16, "alpha": 32}
    train_loss REAL,                    -- e.g., 0.4190
    val_loss REAL,                      -- e.g., 0.3651
    duration_hours REAL,                -- e.g., 3.26
    timestamp DATETIME DEFAULT CURRENT_TIMESTAMP
);

-- Stores generated responses for analysis
CREATE TABLE GeneratedResponses (
    id INTEGER PRIMARY KEY AUTOINCREMENT,
    experiment_id INTEGER,              -- Links to LLAMAExperiments
    input_text TEXT,                    -- User input
    response_text TEXT,                 -- Model response
    timestamp DATETIME DEFAULT CURRENT_TIMESTAMP,
    FOREIGN KEY (experiment_id) REFERENCES LLAMAExperiments(id)
);

Usage

Installation

bash
pip install transformers peft bitsandbytes accelerate torch

Load and Use the Model

python
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel
import torch

# Quantization config (required for 4-bit loading)
bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.float16,
    bnb_4bit_use_double_quant=True,
)

# Load base model
base_model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-3.1-8B-Instruct",
    quantization_config=bnb_config,
    device_map="auto",
    torch_dtype=torch.float16,
)

# Load LoRA adapter
model = PeftModel.from_pretrained(
    base_model, 
    "EdwardConstantine/bengali-empathy-llama"
)
tokenizer = AutoTokenizer.from_pretrained(
    "EdwardConstantine/bengali-empathy-llama"
)

# Generate empathetic response
def generate_response(prompt, max_tokens=200):
    full_prompt = f"""<|begin_of_text|><|start_header_id|>system<|end_header_id|>

You are a compassionate Bengali conversational AI. Respond with empathy. Reply in Bengali.<|eot_id|><|start_header_id|>user<|end_header_id|>

{prompt}<|eot_id|><|start_header_id|>assistant<|end_header_id|>

"""
    inputs = tokenizer(full_prompt, return_tensors="pt").to(model.device)
    outputs = model.generate(
        **inputs,
        max_new_tokens=max_tokens,
        temperature=0.7,
        top_p=0.9,
        do_sample=True,
        pad_token_id=tokenizer.eos_token_id,
    )
    response = tokenizer.decode(outputs[0][inputs['input_ids'].shape[1]:], skip_special_tokens=True)
    return response.strip()

# Example usage
response = generate_response("আমি খুব ą¦ą¦•ą¦¾ অনুভব ą¦•ą¦°ą¦›ą¦æą„¤")
print(response)

Challenges Faced

ChallengeProblemSolution
bitsandbytes + tritontriton.ops module not foundCreated in-memory patch to mock the module
JSONL Data TypesMixed string/number types in topic columnRobust loader that converts all values to strings
GPU Memory (16GB)Model too large for full training4-bit quantization + gradient checkpointing
Kaggle TimeoutLost 10+ hours of trainingImplemented frequent checkpointing (every 300 steps)
Download IssuesCouldn't download from KaggleUploaded to HuggingFace Hub
Disk SpaceKaggle ran out of spaceCleaned old checkpoints, uploaded to HuggingFace

Future Improvements

ImprovementWhat It Would DoDifficulty
Full sequence length (512+)Better context understandingNeeds better GPU
100% training dataBetter generalizationNeeds more time
3 full epochsLower loss, better qualityNeeds more time
Unsloth integration2x faster trainingMedium
Gradio demoInteractive web interfaceEasy
More evaluation metricsBERTScore, semantic similarityEasy
Human evaluation studyReal quality assessmentMedium

Files in This Repository

FileDescriptionSize
adapter_model.safetensorsTrained LoRA weights168 MB
adapter_config.jsonLoRA configuration915 B
tokenizer.jsonTokenizer17.2 MB
tokenizer_config.jsonTokenizer config50.6 KB
special_tokens_map.jsonSpecial tokens325 B
chat_template.jinjaChat format template4.61 KB
evaluation_results.csvDetailed evaluation results215 KB
metrics_summary.jsonBLEU, ROUGE, Perplexity0.2 KB
human_evaluation_sheet.csvHuman evaluation template15.5 KB
experiments.dbSQLite experiment logs12 KB

Citation

bibtex
@misc{bengali-empathy-llama-2024,
  author = {EdwardConstantine},
  title = {Bengali Empathetic LLaMA: Fine-tuned LLaMA 3.1-8B for Empathetic Bengali Conversations},
  year = {2024},
  publisher = {HuggingFace},
  url = {https://huggingface.co/EdwardConstantine/bengali-empathy-llama}
}

License

This model is released under the Apache 2.0 License, subject to Meta's LLaMA license terms.


Acknowledgments

  • —Meta AI for LLaMA 3.1-8B-Instruct base model
  • —Hugging Face for transformers and PEFT libraries
  • —Kaggle for free GPU access
  • —Bengali Empathetic Conversations Dataset creators