CoolFace
Modelpublic

HassanB4/halluscoring-marbert-nli

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes26downloads
Model Card

halluscoring-marbert-nli

MARBERTv2 (UBC-NLP/MARBERTv2) fine-tuned on HalluScoring 2026 Task 1.1 using the NLI framing ([CLS] gold_answer [SEP] model_answer [SEP]). Internally this is run S03 — confirms the NLI-framing gain first found on CAMeLBERT (`halluscoring-camelbert-nli`) is architecture-independent; scores essentially tie (clean-dev AUC-ROC 0.9266 vs. 0.9272, −0.06pp).

Not submitted to the competition — kept as an internal experiment and as a component of the S23/S24v/S25v ensembles (dropped from the final best ensemble, S25v2, whose predictions turned out to be redundant with the ARBERT model since both are UBC-NLP architectures). See `SYSTEM_WRITEUP.md` for the officially-submitted models.

How to Use

python
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

model_id = "HassanB4/halluscoring-marbert-nli"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
model.eval()

gold_answer = "..."
model_answer = "..."

inputs = tokenizer(gold_answer, model_answer, truncation=True, max_length=512, return_tensors="pt")
with torch.no_grad():
    logits = model(**inputs).logits
    prob_hallucinated = torch.softmax(logits, dim=-1)[0, 1].item()

print(f"hallucinated={int(prob_hallucinated > 0.5)}, score={prob_hallucinated:.4f}")

Training

ParameterValue
Base modelUBC-NLP/MARBERTv2
Input formatnli (gold_answer + model_answer)
Max sequence length512
Batch size16
Epochs5
Learning rate2e-5
Warmup ratio0.1
Weight decay0.01
Losscross-entropy
Seed42

Evaluation

SplitAUC-ROCF1-Macro
Dev (official, n=1300)0.95620.8989
Dev (clean, unseen-question subset, n=800)0.9266

Limitations

Not evaluated on the hidden test set — internal dev-only experiment. Predictions were found to be redundant with halluscoring-arbert-nli (both UBC-NLP architectures), which is why it was excluded from the final S25v2 ensemble in favor of a more architecturally diverse model.

Citation

bibtex
@inproceedings{namaa2026halluscoring,
    title={{NAMAA at HalluScoring 2026: NLI-Framed BERT Classifiers and Ensembling for Model-Agnostic Arabic Hallucination Detection}},
    author={[AUTHOR NAMES TBD]},
    year={2026},
    booktitle={Proceedings of ArabicNLP 2026},
    note={HalluScoring 2026 Shared Task, Track 1}
}