cometadata/funding-parsing-lora-Llama_3.1_8B-instruct-ep2-r64-a32-grpo
011
Funding Parsing LoRA — Llama 3.1 8B Instruct + GRPO
A LoRA adapter for extracting structured funding information from funding statements in scholarly works.
Model Details
- Base model: meta-llama/Llama-3.1-8B-Instruct
- Method: Supervised fine-tuning (SFT) followed by reinforcement learning (GRPO)
- Task: Given a funding statement, extract a structured JSON array of funders with their award IDs
- Training data: cometadata/funding-extraction-sft-data
Training Pipeline
Stage 1: Supervised Fine-Tuning
- Data: cometadata/funding-extraction-sft-data — 1,316 real examples (
train.jsonl) + 2,531 synthetic examples (synthetic.jsonl) upsampled 2x = 6,378 total - Epochs: 2
- LoRA rank: 64
- LoRA alpha: 32
- Learning rate: ~2.86e-4
- Batch size: 128
- Max sequence length: 4,096 tokens
- LR schedule: Linear decay
- Renderer: llama3 (Llama 3.1 Instruct chat template)
- Train on: Last assistant message only
Stage 2: Reinforcement Learning (GRPO)
- Algorithm: Group Relative Policy Optimization (GRPO)
- Starting checkpoint: SFT final weights
- Data: 3,462 train / 385 eval examples
- Learning rate: 3e-5
- Temperature: 0.8
- Batch size: 16, Group size: 8
- KL penalty: 0.03 (against SFT reference policy)
- Best checkpoint: Step 130 / 217 (selected by eval reward)
- Eval reward at best step: 0.961
Reward Function
See https://github.com/cometadata/funding-metadata-enrichment/tree/main/train for the full training code
Gated, hierarchical matching on funder using the Hungarian algorithm for limiting subordinate fields 1:1 funder pairing:
- Funder name - Fuzzy matching using a token-sort with acronym and containment boosts, F0.5 score, weight 0.50
- Award IDs - Normalized exact matching, F0.5 score, weight 0.50
- Funding scheme - Not weighted
- Award title - Not weighted
Usage
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base_model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.1-8B-Instruct")
model = PeftModel.from_pretrained(base_model, "cometadata/funding-parsing-lora-Llama_3.1_8B-instruct-ep2-r64-a32-grpo")
tokenizer = AutoTokenizer.from_pretrained("cometadata/funding-parsing-lora-Llama_3.1_8B-instruct-ep2-r64-a32-grpo")
messages = [
{"role": "system", "content": "Extract funding information from the text. Return a JSON array of funders."},
{"role": "user", "content": "Extract funding information from the following statement:\n\nThis work was supported by the National Science Foundation (Grant No. 2045678) and the European Research Council (ERC-2021-StG-101039567)."}
]
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True)
outputs = model.generate(inputs, max_new_tokens=512, temperature=0.1)
print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))Expected output:
[
{
"funder_name": "National Science Foundation",
"awards": [
{
"award_ids": ["2045678"],
"funding_scheme": [],
"award_title": []
}
]
},
{
"funder_name": "European Research Council",
"awards": [
{
"award_ids": ["ERC-2021-StG-101039567"],
"funding_scheme": [],
"award_title": []
}
]
}
]Training Infrastructure
Trained on Tinker by Thinking Machines Lab
