CoolFace
Modelpublic

cometadata/funding-parsing-lora-Llama_3.1_8B-instruct-ep2-r64-a32-grpo

sourceHugging Facellama3.1updated 6mo agoView on Hugging Face
0likes11downloads
Model Card

Funding Parsing LoRA — Llama 3.1 8B Instruct + GRPO

A LoRA adapter for extracting structured funding information from funding statements in scholarly works.

Model Details

Training Pipeline

Stage 1: Supervised Fine-Tuning

  • —Data: cometadata/funding-extraction-sft-data — 1,316 real examples (train.jsonl) + 2,531 synthetic examples (synthetic.jsonl) upsampled 2x = 6,378 total
  • —Epochs: 2
  • —LoRA rank: 64
  • —LoRA alpha: 32
  • —Learning rate: ~2.86e-4
  • —Batch size: 128
  • —Max sequence length: 4,096 tokens
  • —LR schedule: Linear decay
  • —Renderer: llama3 (Llama 3.1 Instruct chat template)
  • —Train on: Last assistant message only

Stage 2: Reinforcement Learning (GRPO)

  • —Algorithm: Group Relative Policy Optimization (GRPO)
  • —Starting checkpoint: SFT final weights
  • —Data: 3,462 train / 385 eval examples
  • —Learning rate: 3e-5
  • —Temperature: 0.8
  • —Batch size: 16, Group size: 8
  • —KL penalty: 0.03 (against SFT reference policy)
  • —Best checkpoint: Step 130 / 217 (selected by eval reward)
  • —Eval reward at best step: 0.961

Reward Function

See https://github.com/cometadata/funding-metadata-enrichment/tree/main/train for the full training code

Gated, hierarchical matching on funder using the Hungarian algorithm for limiting subordinate fields 1:1 funder pairing:

  • —Funder name - Fuzzy matching using a token-sort with acronym and containment boosts, F0.5 score, weight 0.50
  • —Award IDs - Normalized exact matching, F0.5 score, weight 0.50
  • —Funding scheme - Not weighted
  • —Award title - Not weighted

Usage

python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base_model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.1-8B-Instruct")
model = PeftModel.from_pretrained(base_model, "cometadata/funding-parsing-lora-Llama_3.1_8B-instruct-ep2-r64-a32-grpo")
tokenizer = AutoTokenizer.from_pretrained("cometadata/funding-parsing-lora-Llama_3.1_8B-instruct-ep2-r64-a32-grpo")

messages = [
    {"role": "system", "content": "Extract funding information from the text. Return a JSON array of funders."},
    {"role": "user", "content": "Extract funding information from the following statement:\n\nThis work was supported by the National Science Foundation (Grant No. 2045678) and the European Research Council (ERC-2021-StG-101039567)."}
]

inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True)
outputs = model.generate(inputs, max_new_tokens=512, temperature=0.1)
print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))

Expected output:

json
[
  {
    "funder_name": "National Science Foundation",
    "awards": [
      {
        "award_ids": ["2045678"],
        "funding_scheme": [],
        "award_title": []
      }
    ]
  },
  {
    "funder_name": "European Research Council",
    "awards": [
      {
        "award_ids": ["ERC-2021-StG-101039567"],
        "funding_scheme": [],
        "award_title": []
      }
    ]
  }
]

Training Infrastructure

Trained on Tinker by Thinking Machines Lab

Eval Results

StepEval RewardFunder F0.5Award F0.5Format ValidKL
00.9440.9610.92699.7%0.0005
100.9370.9540.92099.7%0.0015
200.9500.9690.932100%0.0020
300.9540.9660.942100%0.0025
400.9510.9710.931100%0.0013
500.9380.9560.919100%0.0051
600.9490.9670.931100%0.0047
700.9540.9680.939100%0.0025
800.9510.9620.940100%0.0021
900.9450.9590.931100%0.0026
1000.9430.9630.92399.7%0.0016
1100.9450.9610.92999.5%0.0036
1200.9500.9640.93699.5%0.0028
1300.9610.9740.948100%0.0026
1400.9550.9730.938100%0.0020
1500.9570.9720.942100%0.0012
1600.9470.9630.93199.7%0.0034
1700.9510.9570.944100%0.0023
1800.9440.9600.928100%0.0013
1900.9330.9560.91099.5%0.0004
2000.9420.9610.92299.7%0.0017
2100.9570.9670.947100%0.0014