CoolFace
Modelpublic

BaoNhan/wikibert-ViFactCheck-FC

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes12downloads
Model Card

wikibert-ViFactCheck-FC

This model is TurkuNLP/wikibert-base-vi-cased fine-tuned for VFC-FC on ViFactCheck using the claim paired with full article context.

Evaluation protocol

  • —Dataset size: 7,232 examples.
  • —Shared fixed stratified splits for FC and GE: 5,785 train / 723 development / 724 test.
  • —Labels: Supported, Refuted, and Not Enough Information.
  • —Fine-tuning seeds: [42, 22, 202].
  • —Training: 3 epoch(s), AdamW, learning rate 2e-05, weight decay 0.01, warmup ratio 0.1.
  • —Effective train batch size: 8 (hard-validated against every published run).
  • —Maximum sequence length: 256.
  • —Input mode: raw Vietnamese claim and passage.
  • —The claim is always preserved; only the second sequence (full article context) is truncated when the pair exceeds the encoder limit.
  • —Topic, author, outlet, URL and other source metadata are excluded from model inputs.
  • —No class weighting, resampling, retrieval model, sentence ranking, test-time model selection or external evidence is used.
  • —Checkpoints are selected by development Macro-F1. The representative published checkpoint is seed 42, selected only by development Macro-F1.

Results

Test metrics are reported as mean ± sample standard deviation over seeds [42, 22, 202].

MetricMean ± std
Test Macro-F10.6229 ± 0.0014
Test accuracy0.6225 ± 0.0021
Test macro precision0.6278 ± 0.0018
Test macro recall0.6217 ± 0.0022
Development Macro-F10.6125 ± 0.0085

Per-seed results

seeddev_macro_f1test_macro_f1test_accuracymicro_batch_sizegradient_accumulation_steps
22.0000000.6079210.6240510.6229288.0000001.000000
42.0000000.6223520.6213260.6201668.0000001.000000
202.0000000.6073210.6232030.6243098.0000001.000000

Label mapping

json
{
  "0": "supported",
  "1": "refuted",
  "2": "not_enough_information"
}

Usage

python
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer

model_id = "BaoNhan/wikibert-ViFactCheck-FC"
tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=False)
model = AutoModelForSequenceClassification.from_pretrained(model_id)

claim = "Thông tin này đã được cơ quan chức năng xác nhận."
context = "Bài báo cung cấp bằng chứng liên quan đến phát biểu trên."
inputs = tokenizer(
    claim,
    context,
    return_tensors="pt",
    truncation="only_second",
    max_length=256,
)
with torch.no_grad():
    probabilities = model(**inputs).logits.softmax(dim=-1)[0]
predicted_id = int(probabilities.argmax())
print(model.config.id2label[predicted_id], probabilities.tolist())

Files

  • —aggregate_metrics.json: aggregate metrics and training manifest.
  • —artifacts/per_seed_results.csv: one row per fine-tuning seed.
  • —artifacts/seed_*_confusion_matrix.csv: confusion matrix for each seed.
  • —artifacts/seed_*_classification_report.json: per-class metrics.
  • —artifacts/seed_*_test_predictions.csv: IDs, gold/predicted labels and probabilities; raw claims and passages are excluded.

Limitations

ViFactCheck supplies the correct source article and therefore does not evaluate open-web evidence retrieval. VFC-FC can truncate relevant information in long articles and jointly measures verification plus robustness to irrelevant context. VFC-GE uses oracle gold evidence and must not be presented as a realistic end-to-end deployment setting. This model is a research classifier, not an automated arbiter of truth, and may produce confidently incorrect predictions.

Dataset citation

bibtex
@inproceedings{hoa2025vifactcheck,
  title={ViFactCheck: A New Benchmark Dataset and Methods for Multi-domain News Fact-Checking in Vietnamese},
  author={Hoa, Tran Thai and Duy, Tran Quang and Tran, Khanh Quoc and Nguyen, Kiet Van},
  booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
  volume={39},
  number={1},
  pages={308--316},
  year={2025},
  doi={10.1609/aaai.v39i1.32008}
}