CoolFace
Modelpublic

ai4data/lfm2.5-Encoder-350M-datause

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes23downloads
Model Card

lfm2.5-Encoder-350M-datause-extended

Fine-tune of LiquidAI/LFM2.5-Encoder-350M for BIO data-use mention tagging (dataset / survey / census / registry mentions in economics research papers).

Labels

  • NAMED_DATA — a proper name, title, or acronym of a specific data source
  • DESCRIPTIVE_DATA — a source described in words but not named
  • VAGUE_DATA — generic data wording with no identifiable source

Training

  • base model: LiquidAI/LFM2.5-Encoder-350M
  • dataset: rafmacalaba/data-use-mentions-extended (bio config)
  • epochs: 5
  • learning rate: 2e-05
  • batch size: 16
  • precision: bf16

Evaluation (holdout, label-agnostic)

thrtpfpfnprecisionrecallf0.5f1
0.107803171547610.81980.62110.77050.7067
0.207803171547610.81980.62110.77050.7067
0.307803171547610.81980.62110.77050.7067
0.407787170247770.82060.61980.77070.7062
0.507658162549060.82490.60950.77050.7011
0.607083130454810.84450.56380.76800.6761
0.706560105760040.86120.52210.76220.6501

Best F0.5: 0.7707 (thr=0.4) Best F1: 0.7067 (thr=0.1)