CoolFace
Modelpublic

zaher-m/stanceeval2026

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes
Model Card

StanceEval-2026 · Arabic Stance Detection

Trained models, generated data, and code for our StanceEval-2026 system on Mawqif-v2. The task is to label a (target, tweet) pair as Favor, Against or None, both for targets that appear in training and for targets held out completely.

The system uses three kinds of probability source: fine-tuned Arabic encoders, a retrieval-augmented few-shot LLM, and LoRA-adapted generative models. A per-class linear head combines their outputs, and a threshold rule derived from the evaluation metric turns the scores into labels.

PathContents
models/encoders/14 fine-tuned encoders, each with its best.json training config
models/lora/7 LoRA adapters, each with its adapter_config.json
stacking_ensemble/the stacking head and its probability bank
data/the training sets we generated
code/training and prediction pipeline
submissions/126 label files, one per configuration we ran
assets/figures

Task and data

Mawqif-v2 has two settings. One tests a new target inside a familiar campaign, the other tests topics the model has never seen. The metric is F<sub>avg2</sub>, the mean of the Favor and Against F1 scores, so None never contributes to it directly.

Track 1, seenTrack 2, unseen
TrainWomen Empowerment, Covid Vaccine, Digital Transformation (3,502)Covid Vaccine, Digital Transformation (2,721)
Devin-domain splitWomen Empowerment, held out (1,400)
TestWomen Driving (352)Ecars (332), Trimester (312)

data/ holds only the augmented and synthesized sets, described in `data/README.md`.

Experiments

The system was not designed up front. Each experiment answered a question the previous one left open, and Figure 1 shows that order.

[image]

1. Which encoder carries the task?

Four Arabic encoder families, same recipe for all of them (max_len 128, batch 16, seed 42, 12 epochs), with the loss and the preprocessing mode ablated on top.

[image]

MARBERTv2 came out ahead on both settings, but the more useful signal was the shape of the gap. It grows from 0.1 points over AraBERT-twitter on seen targets to 3.6 points on unseen ones. Focal loss and class weighting both lost accuracy against plain cross-entropy (0.779 and 0.809 against 0.820).

EncoderSeen (dev)Unseen (dev)
MARBERTv20.83600.8195
AraBERT-twitter0.83510.7837
CAMeLBERT-mix0.79040.7506
AraModernBERT0.78970.7308

A gap that only widens when the target changes looks like a transfer problem rather than a capacity problem. The encoders were probably leaning on lexical cues tied to the targets they trained on, and neither a different loss nor a different family fixes that. So we tried a model that is told the task instead of trained on the target.

2. Does prompting transfer better than fine-tuning?

We put the same task to a served instruction model as a few-shot classifier and compared it against the encoder ensemble. It won on both settings, and by more on the unseen one. If demonstrations are what carry the transfer, then which demonstrations end up in the prompt should matter more than anything else, and that is easy to test.

3. Does shot selection matter?

We replaced randomly drawn demonstrations with MARBERTv2 dense retrieval, kept the classes balanced, and swept the shot count and the number of self-consistency votes.

[image]

Retrieval beat random selection on both settings, 0.8811 to 0.8964 and 0.8879 to 0.8942 on dev, with a fairly sharp optimum around six shots. Selection matters, as expected, but the sweep exposed a bigger problem: the pool only held training-target tweets, so a new target retrieved nothing on-target at all. The bottleneck was the pool, not the retriever.

4. What if we generate the shots the pool is missing?

We generated on-target tweets for each test target, kept the ones an independent labelling pass agreed with, and added them to the pool. This was the largest single gain in the study: 0.815 to 0.844 on seen targets and 0.851 to 0.879 on unseen ones. On the hardest unseen target the jump was larger still, 0.841 to 0.904.

Two things could explain it. Either the new shots sit closer to the target, or there are simply more of them to draw on. Those predict opposite results under scaling, so one experiment separates them.

5. Proximity or volume?

We doubled the generated pool and doubled the shots per query. Both settings got worse, 0.844 to 0.836 and 0.879 to 0.867. So it is proximity: extra shots push the closest ones out of the retrieved window. The next step was to make shots closer in a stronger sense than topic.

6. Does surface form count as proximity?

We re-anchored generation on the real unlabelled test tweets, using them as style templates and seeding the campaign's own hashtags. The clearest evidence is in retrieval reach rather than the score: the share of test tweets whose nearest same-class demonstration was a real on-target one went from 62% to 81% for Favor and 72% to 87% for Against, and the score moved 0.8445 to 0.8530.

The same treatment lost accuracy on targets with no distinctive campaign surface (0.8786 to 0.8698). That is consistent rather than awkward, since what helps is matching a surface, and where there is no surface to match there is nothing to gain. It also gives a prediction we can check: if surface form carries stance signal, normalizing it away should hurt.

7. What else could explain the results?

Eight alternatives, including the normalization test experiment 6 predicts. All eight lost accuracy.

[image]

AlternativeΔ *F*<sub>avg2</sub>What it rules out
Add a decorrelated member from another base−0.0373decorrelation on its own; this member shared the base's directional bias
Reason step by step before answering−0.0298explicit reasoning, which reads implicit criticism as opposition
Normalize hashtags and elongation−0.0166as predicted, the surface was carrying signal
Self-train on the system's own labels−0.0159pseudo-labelling; the problem is self-confirmation, not label noise
Swap in a stronger embedding retriever−0.0095retriever quality, already saturated around k=6
Surface-ground a target with no campaign−0.0088surface grounding as a general recipe
Train on real out-of-campaign labelled data−0.0047that real labels beat proximal generated ones
Retrain a member to be more accurate−0.0027member accuracy as the ensemble objective

Eight independent attempts to add model capability all failed, which reads as the single-model signal being saturated. What we had not looked at was anything downstream of the probabilities: how they become labels, and how the members are combined.

8. Is argmax the right decision rule?

F<sub>avg2</sub> ignores None, but argmax implicitly treats all three classes as equally worth claiming. Swapping it for the plug-in rule for that metric, claim class c only when P(c) exceeds F<sub>c</sub>/2 and fall back to None when neither class clears its threshold, improved the same probabilities without touching any model. The rule is standard F-measure thresholding (Zhao et al. 2013; Lipton et al. 2014; Koyejo et al. 2014).

If the layer that reads the probabilities had that much unclaimed value in it, the layer that produces them was worth the same look.

9. How much does weighting whole members cost?

Every combiner up to here gave each member one scalar weight. This experiment asks whether that form is expressive enough. We fitted weighted-mixture families and a per-class linear head to the same 40-source bank, with a random-label control to show what fitting alone can do.

[image]

CombinerParamsAgreement5-fold CVReproduces all 644 labels
Best single source00.8660.8665no
Uniform average00.8630.8634no
Convex mixture39no, certified (γ\* = −0.469)
Signed mixture39no, certified (γ\* = −0.341)
Log mixture39no, certified (γ\* = −2.124)
Per-class linear head3631.0000.8680yes
Random-label control3630.6090.486no

The mixture rows are not failed searches. Each family admits a Farkas certificate with a strictly negative margin γ\*, which proves no weight vector in that family reproduces the labeling. The per-class head does, using 363 parameters over 120 features. Fitting the identical head to shuffled labels reaches only 0.609 in-sample and 0.486 cross-validated, so it is picking up real structure in where the members disagree rather than absorbing arbitrary assignments.

The reason is easy to state. A member that is reliable on Against and unreliable on None gets one number under any mixture, while the per-class head can keep the first and discount the second. That asymmetry is what 39 parameters cannot express.

Results

[image]

Things worth carrying to a similar task:

  1. 1.Making the nearest retrievable demonstration look like the query, down to surface form, was worth more than any model change we tried.
  2. 2.A member is useful to an ensemble for the direction of its errors, not its accuracy. Making one more accurate removed what made it useful.
  3. 3.Validation sets from a different distribution were worse than having no validation intuition.

The system

[image]

Each member emits an (N, 3) probability array. The bank holds 40 sources on the unseen setting and 38 on the seen one: 24 few-shot runs varying pool, shot count, prompt, retriever and served model, 10 encoders, and 6 LoRA members. stacking_ensemble/predict.py applies the head to that bank and regenerates the released label vectors exactly.

Models

21 checkpoints, with per-model detail in `MODELS.md`.

GroupCountBasesConfig
Encoders14MARBERTv2, AraBERTv02-twitter, AraModernBERT, CAMeLBERT-mixmax_len 128, batch 16, seed 42, 12 epochs
LoRA7ALLaM-7B-Instruct (6), Qwen2.5-7B-Instruct (1)r=16, α=32, all 7 projections

Encoders use label order ["Against","Favor","None"] (id2label 0=Against, 1=Favor, 2=None). LoRA members are scored by label log-probability. Base LLMs load from their own repositories rather than being re-hosted here.

Usage

bash
pip install -r requirements.txt

Encoder member:

python
from transformers import AutoModelForSequenceClassification, AutoTokenizer

REPO = "zaher-m/stanceeval2026"
SUB  = "models/encoders/t2_marbert"          # any subfolder under models/encoders/
tok  = AutoTokenizer.from_pretrained(REPO, subfolder=SUB)
clf  = AutoModelForSequenceClassification.from_pretrained(REPO, subfolder=SUB)

x = tok("Covid Vaccine", "التطعيم ضروري لحماية الجميع",
        return_tensors="pt", truncation=True, max_length=128)
print({0: "Against", 1: "Favor", 2: "None"}[clf(**x).logits.argmax(-1).item()])   # Favor

LoRA member, with the base pulled from its own repo:

python
from transformers import AutoModelForCausalLM
from peft import PeftModel

base  = AutoModelForCausalLM.from_pretrained("ALLaM-AI/ALLaM-7B-Instruct-preview", device_map="auto")
model = PeftModel.from_pretrained(base, "zaher-m/stanceeval2026",
                                  subfolder="models/lora/allam_t2")

Reproducing the released predictions needs numpy and no GPU:

python
from huggingface_hub import snapshot_download
import subprocess, sys

d = snapshot_download("zaher-m/stanceeval2026", allow_patterns="stacking_ensemble/*")
subprocess.run([sys.executable, f"{d}/stacking_ensemble/predict.py", "--track", "t2"])
# [t2] 644 labels -> t2_perclass_stack_out.txt | matches bundled predictions: True

The full pipeline and the training entry points are in `code/README.md`.

Limitations

These are research artifacts for Arabic stance detection on dialectal tweets. Individual members are ensemble components rather than standalone classifiers, and several are kept because they were useful counterweights or because they are documented failures.

  • —The decode threshold optimizes F<sub>avg2</sub>, which ignores None, so the system under-predicts None on purpose. Retune the thresholds if you care about that class.
  • —Campaign hashtags and meme surface carry signal in this data. We have not measured performance on MSA news text or other domains.
  • —Four auxiliary encoders (AraELECTRA, XLM-R-large, ARBERTv2, AraBERT-large) were used as probability sources but never saved. Their outputs are in the bank and code/ can rebuild them.
  • —Stance labels are contested and target-relative. Do not read the output as a fact about anyone's views.

Base models

ModelRole
`ALLaM-AI/ALLaM-7B-Instruct-preview`base for the ALLaM adapters
`Qwen/Qwen2.5-7B-Instruct`base for the Qwen adapter
`UBC-NLP/MARBERTv2`main encoder and retrieval embedder
`aubmindlab/bert-base-arabertv02-twitter`encoder member
`NAMAA-Space/AraModernBert-Base-V1.0`encoder ablation
`CAMeL-Lab/bert-base-arabic-camelbert-mix`encoder ablation
LilaRest/gemma-4-31B-it-NVFP4-turbo, nvidia/Qwen3.6-35B-A3B-NVFP4served models for the few-shot classifier

References

  • —Alturayeif, Luqman & Ahmed (2022). Mawqif: A Multi-label Arabic Dataset for Target-specific Stance Detection. WANLP.
  • —Abdul-Mageed, Elmadany & Nagoudi (2021). ARBERT & MARBERT: Deep Bidirectional Transformers for Arabic. ACL.
  • —Antoun, Baly & Hajj (2020). AraBERT: Transformer-based Model for Arabic Language Understanding. OSACT/LREC.
  • —Inoue, Alhafni, Baimukan, Bouamor & Habash (2021). The Interplay of Variant, Size, and Task Type in Arabic Pre-trained Language Models. WANLP.
  • —Hu et al. (2021). LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685.
  • —Wang et al. (2022). Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv:2203.11171.
  • —Wolpert (1992). Stacked Generalization. Neural Networks 5(2).
  • —Lin et al. (2017). Focal Loss for Dense Object Detection. ICCV.
  • —Zhao et al. (2013); Lipton, Elkan & Naryanaswamy (2014), ECML; Koyejo et al. (2014), NeurIPS. Plug-in thresholding for F-measures.