CoolFace
Modelpublic

DrGM/DrGM-ConvNeXt-V2L-Facial-Emotion-Recognition

sourceHugging Facecc-by-nc-4.0updated 10mo agoView on Hugging Face
2likes49downloads
Model Card

DrGM-ConvNeXt-V2-Large-FER

Model Description

This is a State-of-the-Art (SOTA) Facial Emotion Recognition (FER) model based on the ConvNeXt V2 Large architecture. It has been fine-tuned to recognize 7 distinct facial emotions with high accuracy.

Model Architecture

  • —Base Model: facebook/convnextv2-large-22k-224
  • —Parameters: ~198M
  • —Fine-tuning: Optimized with BF16 mixed precision, Label Smoothing (0.1), and RandAugment on an A100 GPU.

⚠️ License & Usage

This model is released under the CC-BY-NC-4.0 license.

  • —Personal Use: ✅ Allowed. You can use this for personal projects, research, and education.
  • —Commercial Use: ❌ Forbidden without prior permission.
  • —Commissions: If you wish to use this model for commercial applications or commissions, please contact the author for licensing.

Training History

The model was trained for 15 epochs on an A100 GPU. Below is the detailed progression of loss and metrics:

EpochTraining LossValidation LossAccuracyF1 (Weighted)
11.01260.948876.31%0.7629
20.86000.839481.91%0.8177
30.71610.793285.02%0.8495
40.63400.755287.52%0.8748
50.59560.740588.34%0.8829
60.55680.724788.94%0.8893
70.52590.725189.03%0.8902
80.51490.720889.16%0.8913
90.50710.717289.66%0.8964
100.49840.715689.66%0.8963
110.49330.710189.88%0.8989
120.48570.707189.92%0.8991
130.48030.703890.25%0.9025
140.47180.703190.43%0.9042
150.47300.701390.40%0.9039

Final Training Metrics

  • —Total Training Time: ~49 minutes (2933.82 seconds)
  • —Global Steps: 11,805
  • —Final Training Loss: 0.5977
  • —Throughput: 257.37 samples/second

Performance

The model achieves exceptional performance on the Facial Emotion Expressions dataset.

Final Evaluation Results (Test Set)

After training, the model was evaluated on the unseen test set:

MetricValue
Accuracy90.43%
F1 Score (Weighted)0.9042
Validation Loss0.7031
Inference Time (Batch)23.63s (Total)
Throughput532.68 samples/sec

(Note: These metrics are from the held-out test split, confirming the model generalizes well and is not just memorizing data.)

Classification Report (Full Dataset Evaluation)

ClassPrecisionRecallF1-ScoreSupport
Angry0.970.970.978989
Disgust1.001.001.008989
Fear0.970.960.978989
Happy0.980.980.988989
Neutral0.960.970.978989
Sad0.960.960.968989
Surprise0.990.990.998989
Accuracy0.9862923

(Note: Full dataset evaluation includes both training and validation samples, indicating high model capacity and learning)

Confusion Matrix

[image]

📊 Advanced Model Statistics

Global Accuracy Metrics:

  • —Top-1 Accuracy: 97.78%
  • —Top-2 Accuracy: 99.05% (Correct emotion is in the top 2 predictions)
  • —Top-3 Accuracy: 99.41%

Per-Emotion Performance Breakdown

EmotionAccuracyAvg ConfidenceSamples
angry97.49%89.99%8989
disgust100.00%91.34%8989
fear96.07%89.66%8989
happy98.16%90.72%8989
neutral97.35%90.39%8989
sad96.27%89.55%8989
surprise99.13%90.81%8989

Inference Speed Benchmark

Tested on an NVIDIA A100 GPU with a batch size of 1 (simulating real-time usage):

  • —Average Latency: 20.75 ms per image
  • —Frame Rate: 48.19 FPS

This performance indicates the model may be suitable for real-time video processing applications.

Usage

python
from transformers import AutoImageProcessor, AutoModelForImageClassification
import torch
from PIL import Image

# Load Model
repo_name = "DrGM/DrGM-ConvNeXt-V2L-Facial-Emotion-Recognition" 
processor = AutoImageProcessor.from_pretrained(repo_name)
model = AutoModelForImageClassification.from_pretrained(repo_name)

# Predict
image = Image.open("path/to/your/image.jpg").convert("RGB")
inputs = processor(image, return_tensors="pt")

with torch.no_grad():
    logits = model(**inputs).logits
    predicted_label = logits.argmax(-1).item()
    print(model.config.id2label[predicted_label])