CoolFace
Modelpublic

loris3/Llama-3.2-1B_tulu-3-sft-olmo-2-mixture-0225_lr0.0001_seed42

sourceHugging Facecc-by-4.0updated 9mo agoView on Hugging Face
0likes12downloads
Model Card

Model Card for Llama-3.2-1Btulu-3-sft-olmo-2-mixture-0225lr0.0001_seed42

This model is a instruction fine-tuned version of meta-llama/Llama-3.2-1B trained using a LoRA adapter on Tülu3 for one epoch via TRL.

Uses

This model was created for training data influence estimation experiments using DataInf and LESS. See our paper and repo for details.

Quick start

python
from huggingface_hub import hf_hub_download
import json
from transformers import AutoTokenizer,AutoModelForCausalLM
from peft import PeftModel
from transformers import pipeline

repo_id = "loris3/Llama-3.2-1B_tulu-3-sft-olmo-2-mixture-0225_lr0.0001_seed42"

adapter_path = hf_hub_download(repo_id=repo_id, filename="adapter_config.json")
adapter_config = json.load(open(adapter_path))
    
base_model_name_or_path = adapter_config["base_model_name_or_path"]
chat_template = open(hf_hub_download(repo_id=repo_id, filename="chat_template.jinja")).read()


tokenizer = AutoTokenizer.from_pretrained(base_model_name_or_path)
tokenizer.chat_template = chat_template
if tokenizer.pad_token is None:
    tokenizer.pad_token = tokenizer.eos_token 
    
model = AutoModelForCausalLM.from_pretrained(base_model_name_or_path)
model.resize_token_embeddings(len(tokenizer))
model = PeftModel.from_pretrained(model, repo_id, is_trainable=False)



question = "Could you give us some of your political beliefs?"
generator = pipeline("text-generation", model=model, tokenizer=tokenizer, device=0)
output = generator([{"role": "user", "content": question}], max_new_tokens=128, do_sample=False, temperature=1.0, top_p=1.0, return_full_text=False)[0]
print(output["generated_text"])

Training procedure

<img src="https://raw.githubusercontent.com/wandb/assets/main/wandb-github-badge-28.svg" alt="Visualize in Weights & Biases" width="150" height="24"/>

Training Hyperparameters

ParameterValue
Precisionbfloat16
OptimizerAdamW (torch fused)
Learning rate1×10⁻⁴
LR schedulerLinear
Weight decay0.0
Max grad norm1.0
LoRA rank (r)16
LoRA alpha32
LoRA dropout0.1
LoRA biasnone
Target modulesqproj, kproj, vproj, oproj
Trainable paramsLoRA only
Train batch size / device4
Gradient accumulation8
Effective batch size32
Training epochs1
Max sequence length1024
Gradient checkpointingFalse
Seed42

Framework versions

  • —PEFT 0.17.1
  • —TRL: 0.23.0
  • —Transformers: 4.56.2
  • —Pytorch: 2.8.0+cu126
  • —Datasets: 4.0.0
  • —Tokenizers: 0.22.1

Evaluation

We evaluate with OLMES

Task suites: core_9mcqa::olmes, mmlu:mc::olmes, olmo_2_generative::olmes, olmo_2_heldout::olmes | Task | Score| |------|---------| | AGIEval | 0.24 | | ARCC | 0.38 | | ARCE | 0.60 | | BBH | 0.32 | | BoolQ | 0.67 | | CSQA | 0.50 | | CoQA | 0.65 | | DROP | 0.25 | | GSM8K | 0.08 | | HSwag | 0.53 | | JPRDY | 0.53 | | MMLU | 0.30 | | MMLU-Pro | 0.16 | | NatQs | 0.16 | | OBQA | 0.39 | | PIQA | 0.67 | | SIQA | 0.47 | | SQuAD | 0.73 | | TriviaQA | 0.48 | | WinoG | 0.58 |