CoolFace
Modelpublic

Zorryy/NewsBERT-germ-210m

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
0likes11downloads
Model Card

NewsBERT-germ-210m

German news article classifier based on EuroBERT-210m, fine-tuned on hand-annotated news articles about the 2025 German federal election (Bundestagswahl).

Model Description

  • —Base Model: EuroBERT/EuroBERT-210m (210M parameters)
  • —Task: Multi-class text classification (13 categories)
  • —Language: German
  • —Training Data: Hand-annotated German news articles (1580 train, 395 validation, 385 test)
  • —Dataset: Zorryy/news_articles_2025_elections_germany

Performance (Test Set)

MetricScore
F1 Macro0.8208
F1 Weighted0.8211
Precision Macro0.8395
Recall Macro0.8169
Accuracy0.8182

Per-Class Performance

CategoryPrecisionRecallF1Support
Klima / Energie0.78790.86670.825430
Zuwanderung0.83870.86670.852530
Renten0.85710.80000.827630
Soziales Gefälle0.65520.63330.644130
AfD/Rechte0.79310.76670.779730
Arbeitslosigkeit1.00000.70000.823530
Wirtschaftslage0.56000.93330.700030
Politikverdruss0.90000.72000.800025
Gesundheitswesen, Pflege0.90320.93330.918030
Kosten/Löhne/Preise0.92000.76670.836430
Ukraine/Krieg/Russland0.93550.96670.950830
Bundeswehr/Verteidigung0.96300.86670.912330
Andere0.80000.80000.800030

Hyperparameters

Optimized via Optuna TPE sampler with 3-fold stratified cross-validation (Best Trial: 1, CV F1 Macro: 0.8369).

ParameterValue
Learning Rate3.13e-05
LR Schedulerlinear
Epochs15
Batch Size (per device)8
Effective Batch Size16
Warmup Ratio0.1455
Weight Decay0.0021
Label Smoothing0.0832
Max Sequence Length2048
Gradient Clipping0.5
Optimizeradamwtorchfused
Early Stoppingpatience=3
Mixed PrecisionBF16=True, FP16=False

Training Data

The model was trained on hand-annotated German news articles related to the 2025 German federal election. Articles were manually labeled into 13 topic categories by human annotators. The dataset is available at Zorryy/news_articles_2025_elections_germany.

Important: All labels are hand-annotated (not machine-generated), ensuring high label quality.

Usage

python
from transformers import pipeline

classifier = pipeline(
    "text-classification",
    model="Zorryy/NewsBERT-germ-210m",
    trust_remote_code=True,
)

text = "Die Bundesregierung plant neue Massnahmen zur Reduzierung der CO2-Emissionen."
result = classifier(text)
print(result)
# [{"label": "Klima / Energie", "score": 0.95}]

Categories

The model classifies articles into 13 categories:

  1. 1.Klima / Energie
  2. 2.Zuwanderung
  3. 3.Renten
  4. 4.Soziales Gefälle
  5. 5.AfD/Rechte
  6. 6.Arbeitslosigkeit
  7. 7.Wirtschaftslage
  8. 8.Politikverdruss
  9. 9.Gesundheitswesen, Pflege
  10. 10.Kosten/Löhne/Preise
  11. 11.Ukraine/Krieg/Russland
  12. 12.Bundeswehr/Verteidigung
  13. 13.Andere

Limitations

  • —Trained specifically on German news articles about the 2025 Bundestagswahl
  • —May not generalize well to other domains, time periods, or languages
  • —Performance varies by category (see per-class metrics above)

Citation

If you use this model, please cite the underlying dataset and base model.