ai4data/gliner_datause
02.2k
glinerdatauseextended
Fine-tune of urchade/gliner_large-v2.1 for data-use mention extraction (dataset / survey / census / registry mentions in economics research papers).
Labels
NAMED_DATA— a proper name, title, or acronym of a specific data sourceDESCRIPTIVE_DATA— a source described in words but not namedVAGUE_DATA— generic data wording with no identifiable source
Training
- base model:
urchade/gliner_large-v2.1 - dataset:
rafmacalaba/data-use-mentions-extended(gliner config) - epochs: 5
- learning rate: 5e-06
- batch size: 16
- precision: bf16
Evaluation (holdout)
Best F0.5: 0.8437 (thr=0.6) Best F1: 0.8493 (thr=0.5)
<!-- NERCOMPARISONSTART -->
NER holdout comparison
device: NVIDIA H100 NVL
rafmacalaba/data-use-mentions-extended (n=9249)
F0.5 by threshold (sweet spots side-by-side):
<!-- NERCOMPARISONEND -->
