CoolFace
Modelpublic

NUHASHROXME/bangla-fake-news-interpretable

sourceHugging Facemitupdated 1y agoView on Hugging Face
2likes9downloads
Model Card

language:

  • bn license: mit library_name: transformers tags:
  • fake-news-detection
  • bangla
  • bert
  • interpretable-ai
  • psycholinguistics
  • discourse-analysis
  • text-classification datasets:
  • BanFakeNews-2.0 metrics:
  • f1
  • accuracy base_model: sagorsarker/bangla-bert-base model-index:
  • name: bangla-fake-news-interpretable results:
  • task: type: text-classification name: Fake News Detection dataset: type: BanFakeNews-2.0 name: Bangla Fake News Dataset 2.0 split: test metrics:
  • type: f1 value: 0.8437 name: Test F1 Score
  • type: accuracy value: 0.8573 name: Test Accuracy ---

Discourse-Aware Psycholinguistic Bangla Fake News Detector

This model combines BERT embeddings with psycholinguistic and discourse features for interpretable Bangla fake news detection.

Model Details

  • Architecture: BERT + 17 psycholinguistic features + 5 discourse features
  • Base Model: sagorsarker/bangla-bert-base
  • Language: Bengali (Bangla)
  • Task: 4-class fake news detection
  • Training Data: 42,380 samples from BanFakeNews-2.0 dataset

Performance

  • Test F1-Score: 84.37%
  • Test Accuracy: 85.73%
  • Validation F1: 84.65%
  • Classes: 4 categories (0-3)

Key Innovation

This is the first systematic integration of psycholinguistic theory with deep learning for Bangla fake news detection, enabling explainable predictions while maintaining state-of-the-art performance.

Features Extracted

Psycholinguistic Features (17):

  • Emotional intensity markers (fear, anger, positive/negative sentiment)
  • Uncertainty and hedging patterns
  • Cognitive load indicators (repetition, disfluency)
  • Deception-specific linguistic patterns (self-reference, present tense usage)

Discourse Features (5):

  • Text coherence scores across paragraphs
  • Argumentative structure analysis (claims vs evidence)
  • Topic progression and transition patterns

Model Architecture

The model integrates:

  1. 1.BERT contextual embeddings (768 dimensions)
  2. 2.Psycholinguistic features (17 dimensions)
  3. 3.Discourse features (5 dimensions)
  4. 4.Feature fusion layer
  5. 5.Final classification head

Usage

python
# Note: This model requires custom feature extraction
from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("NUHASHROXME/bangla-fake-news-interpretable")
# Custom model loading code required for interpretable features