ai4data/gliner2_datause
022
gliner2datauseextended
Fine-tune of fastino/gliner2-large-v1 (GLiNER2) for data-use mention extraction (dataset / survey / census / registry mentions in economics research papers).
Labels
NAMED_DATA— a proper name, title, or acronym of a specific data sourceDESCRIPTIVE_DATA— a source described in words but not namedVAGUE_DATA— generic data wording with no identifiable source
Training
- base model:
fastino/gliner2-large-v1 - dataset:
rafmacalaba/data-use-mentions-extended(gliner2 config) - epochs: 5
- encoder LR: 1e-05
- task LR: 0.0005
- batch size: 8
- precision: bf16
Evaluation (holdout, label-agnostic)
Best F0.5: 0.8645 (thr=0.7) Best F1: 0.8632 (thr=0.6)
<!-- NERCOMPARISONSTART -->
NER holdout comparison
device: NVIDIA H100 NVL
rafmacalaba/data-use-mentions-extended (n=9249)
F0.5 by threshold (sweet spots side-by-side):
<!-- NERCOMPARISONEND -->
