Madhesh4124/metaphor-detection-backend
Crosslingual Metaphor Detection & Interpretation System
A production-grade NLP architecture for detecting and explaining metaphors across Hindi, Tamil, Telugu, and Kannada using fine-tuned transformer models, explainable AI (XAI) feature attributions, and 5-layer LLM cognitive interpretations.
๐ Live Demo: https://crosslingual-metaphor-detection-interpretation-kovgohbol.vercel.app/
๐๏ธ System Architecture Overview
The system combines dedicated fine-tuned transformer encoders with gradient saliency attribution and generative LLM interpretation:
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ REACT FRONTEND (Vite) โ
โ - On-screen script keyboards (เคนเคฟเคเคฆเฅ / เฎคเฎฎเฎฟเฎดเฏ / เฐคเฑเฐฒเฑเฐเฑ / เฒเฒจเณเฒจเฒก) โ
โ - Web Speech API real-time microphone input โ
โ - Multi-lingual target output selector (EN / HI / TA / TE / KN) โ
โ - Interactive XAI Saliency Badges & 5-Layer Cognitive Interpretations โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ HTTP POST /predict
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ FASTAPI BACKEND ORCHESTRATION โ
โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ 1. Script & Language Auto-Detector (Unicode Range + LangDetect) โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ 2. Indic Sentence Tokenizer & Punctuation Segmenter (เฅค . ? !) โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ 3. On-Demand Lazy Model Loader (Per-Language PyTorch Checkpoints) โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โผ โผ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ 4. Two-Pass Context Engine โ โ 5. PyTorch XAI Gradient โ โ
โ โ & Temperature Scaler โ โ Embedding Saliency Norm โ โ
โ โโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโ โ
โ โ โ โ
โ โโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโ โ
โ โผ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ 6. Gemma 4 26B (gemma-4-26b-a4b-it via Google GenAI) โ โ
โ โ - 5-Layer Semantic & Cultural Interpretation (Target Language) โ โ
โ โ - Secondary Classification Verification โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ 7. Async MongoDB Atlas Persistence (Predictions & History Stats) โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ๐ฌ Core Algorithms & Computational Pipeline
1. Language & Script Detection
Input text is parsed through a hybrid detection hierarchy:
- Direct Unicode Range Mapping (Highest reliability for Indic scripts):
- Devanagari (Hindi):
\u0900to\u097F - Tamil:
\u0B80to\u0BFF - Telugu:
\u0C00to\u0C7F - Kannada:
\u0C80to\u0CFF - LangDetect Engine: Serves as a fallback for mixed or Romanized inputs.
2. Fine-Tuned Transformer Models
Each language runs on an optimized transformer architecture fine-tuned specifically for metaphor binary classification (metaphor vs. normal):
3. Temperature Scaling & Calibration
Raw neural network logits often output overconfident probabilities for RoBERTa architectures and underconfident spreads for BERT/MuRIL on narrow margins. Calibrated confidence is computed via temperature-scaled softmax:
$$\hat{P}(y = k \mid x) = \frac{\exp(zk / T)}{\sum{j} \exp(z_j / T)}$$
- Hindi / Tamil (XLM-RoBERTa): $T = 2.2$ (Smoothes extreme $>99.9\%$ overconfidence).
- Telugu / Kannada (MuRIL / BERT): $T = 0.35$ (Amplifies narrow $52-60\%$ margins to true certainty).
4. Two-Pass Context-Aware Chaining
In multi-sentence paragraphs, metaphors frequently span across sentence boundaries where subsequent sentences appear literal in isolation:
Sentence 1: "เคเคธเคเคพ เคฆเคฟเคฒ เคชเคคเฅเคฅเคฐ เคนเฅเฅค" (His heart is stone.) -> [Metaphor] -> Set Anchor = S1
Sentence 2: "เคเฅเค เคฌเคพเคค เค
เคธเคฐ เคจเคนเฅเค เคเคฐเคคเฅเฅค" (Nothing affects him.) -> Pass 1: [Literal]
Pass 2: Context Evaluation [S1 + S2] -> [Metaphor]- Pass 1 (Isolated Evaluation): Evaluates $S_i$ independently.
- If $Si = \text{Metaphor}$, it updates the active anchor: $\text{Anchor} \leftarrow Si$.
- Pass 2 (Context Evaluation): If $S_i = \text{Literal}$ and an active anchor exists:
- Concatenates: $S{\text{context}} = [\text{Anchor}] \oplus [Si]$.
- If the concatenated forward pass classifies as $\text{Metaphor}$, the label for $S_i$ is updated to Metaphor.
- If both passes return Literal, the context chain is broken and $\text{Anchor} \leftarrow \emptyset$.
5. Explainable AI (XAI) Embedding Gradient Saliency
To explain why the neural network made a classification decision, token-level feature attribution is calculated via backpropagation through the model's word embeddings layer:
- Let $E \in \mathbb{R}^{L \times D}$ be the input embedding tensor for sequence tokens $t1, t2, \dots, t_L$.
- Compute the gradient of the target class logit $z_{\text{target}}$ with respect to the input embeddings:
$$Gi = \nabla{Ei} z{\text{target}}$$
- The saliency score for each token $i$ is calculated using the $L_2$ Euclidean norm of its gradient vector:
$$Si = \|Gi\|_2$$
- Saliency values are normalized into percentage weights across the sequence:
$$Pi = \frac{Si}{\sum{j=1}^{L} Sj} \times 100\%$$
- Dynamic Key Trigger Threshold: A token is marked as a Key Trigger if:
$$Pi \ge (\muS + 0.5 \cdot \sigmaS) \quad \text{OR} \quad (\max(P) - Pi) \le 1.5\%$$
This dual criterion flags both statistically significant tokens above the standard deviation and tightly clustered co-triggers (e.g., "เคเคธเคเคพ" and "เคเคพเคเคฆ" in "เคเคธเคเคพ เคเฅเคนเคฐเคพ เคเคพเคเคฆ เคนเฅ").
6. 5-Layer Cognitive Interpretation (Gemma 4 26B)
For every detected metaphor, the backend queries Gemma 4 26B (gemma-4-26b-a4b-it) using structured few-shot prompting to output five perspectives in the selected language:
- Translation: Direct idiomatic cross-lingual translation.
- Literal: Word-for-word grammatical translation.
- Emotional: Mood, affect, and emotional resonance conveyed.
- Philosophical: Abstract life insight or metaphysical meaning.
- Cultural: Indic cultural context, folklore, and metaphorical traditions.
โก Memory & Performance Optimization
- On-Demand Lazy Loading: PyTorch model weights are only loaded into RAM/VRAM when a request in that specific language is received, ensuring instant server startups and minimal idle memory footprints.
- In-Memory Hash Caching: MD5-hashed cache for identical predictions and LLM calls with a 1-hour TTL ($3600\text{s}$).
- Fast-Fail Database Connectors: MongoDB Atlas client initialized with
serverSelectionTimeoutMS=2000to prevent request stalling when cloud network drops occur.
๐ Project Structure
โโโ backend/
โ โโโ main.py # FastAPI server, endpoints, XAI & inference engine
โ โโโ database.py # Motor async MongoDB connector & history operations
โ โโโ requirements.txt # Backend dependencies
โโโ frontend/
โ โโโ src/
โ โ โโโ App.jsx # Main application component
โ โ โโโ App.css # Core design system & theme variables
โ โ โโโ History.jsx # Historical analytics & record viewer
โ โ โโโ History.css # History modal styling
โ โ โโโ VirtualKeyboard.jsx # On-screen native script keyboards
โ โ โโโ VirtualKeyboard.css # Virtual keyboard styling
โ โโโ package.json # Frontend dependencies & Vite scripts
โโโ models/ # Local model cache (downloaded from Hugging Face)
โโโ datasets/ # Evaluation & benchmark datasets
โโโ training_code/ # Training scripts for XLM-R, MuRIL & BERT
โโโ app.py # Hugging Face Spaces deployment entrypoint
โโโ README.md # System documentation & architectural reference๐ Getting Started
1. Environment Setup
Create a .env file in the project root:
GEMINI_API_KEY=your_google_genai_api_key_here
MONGODB_URL=mongodb+srv://<username>:<password>@cluster0.mongodb.net/?retryWrites=true&w=majority
MONGODB_DB_NAME=metaphor_detector2. Backend Execution
Activate your Python environment and run:
uvicorn backend.main:app --host 127.0.0.1 --port 8000 --reload3. Frontend Execution
In a separate terminal:
cd frontend
npm install
npm run dev -- --port 5173Navigate to http://localhost:5173.
๐ API Reference
POST /predict
Request:
{
"text": "เคเคธเคเคพ เคเฅเคนเคฐเคพ เคเคพเคเคฆ เคนเฅ",
"interpretation_language": "english"
}Response:
{
"language": "hindi",
"label": "metaphor",
"confidence": 0.9421,
"text": "เคเคธเคเคพ เคเฅเคนเคฐเคพ เคเคพเคเคฆ เคนเฅ",
"is_paragraph": false,
"sentences": [
{
"sentence": "เคเคธเคเคพ เคเฅเคนเคฐเคพ เคเคพเคเคฆ เคนเฅ",
"label": "metaphor",
"confidence": 0.9421,
"interpretations": {
"translation": "Her face is radiant like the moon.",
"literal": "Her face moon is.",
"emotional": "Conveys deep romantic admiration and aesthetic wonder.",
"philosophical": "Reflects how human beauty mirrors celestial luminosity.",
"cultural": "In classical Indian poetics (Kavya), comparing a beloved's face to the moon is a standard archetype."
},
"word_attributions": [
{"word": "เคเคธเคเคพ", "score": 15.47, "is_key_trigger": true},
{"word": "เคเฅเคนเคฐเคพ", "score": 13.65, "is_key_trigger": false},
{"word": "เคเคพเคเคฆ", "score": 15.26, "is_key_trigger": true},
{"word": "เคนเฅ", "score": 14.70, "is_key_trigger": false}
],
"decision_reasoning": "The model identified 'เคเคธเคเคพ', 'เคเคพเคเคฆ' as key attribution trigger(s) driving the METAPHOR decision.",
"is_verified": true,
"verification_status": "Verified by Gemini"
}
]
}๐ License
Released for educational, academic, and research purposes.
