nagohachi/Llama-3.2-tiny-lm-japanese-500m-base-v2
Llama-3.2-tiny-lm-japanese-500m-base-v2
Built with Llama. This model is a derivative of Meta's Llama-3.2-3B, produced by structured (width) pruning followed by continued pre-training in Japanese.
A ~476M-parameter Japanese base model created by width-pruning Llama-3.2-3B down to a 500M-class model, swapping in the Japanese llm-jp/llm-jp-3-440m tokenizer, and recovery-pretraining on ~10B tokens of educational Japanese web text. This is the base model — no instruction tuning, no chat template. See Llama-3.2-tiny-lm-japanese-500m-sft-v2 and Llama-3.2-tiny-lm-japanese-500m-dpo-v2 for the instruction-tuned / preference-aligned variants.
It is the pruning-based counterpart to the from-scratch tiny-lm-japanese-500m-base-v1; at an equal ~10B-token budget it reaches the from-scratch model's final quality in roughly half the tokens and scores higher on downstream tasks (see Evaluation).
The model uses a custom Llama-style implementation shipped with the repo (trust_remote_code=True required).
Model details
Training
- Data: hotchpotch/fineweb-2-edu-japanese (
sample_10BT) - Schedule: width-prune Llama-3.2-3B → recovery pre-training on ~10B tokens
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "nagohachi/Llama-3.2-tiny-lm-japanese-500m-base-v2"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True, dtype="bfloat16")
inputs = tok("日本の首都は", return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=64, do_sample=True, temperature=0.7, top_p=0.9)
print(tok.decode(out[0], skip_special_tokens=True))Evaluation
llm-jp-eval (v2.1.5)
We evaluated the models using 100 examples from the test split with greedy decoding. The llm-jp baselines were re-evaluated with the same harness. Base models use the 4-shot setting; SFT/DPO models and the instruct baseline are prompted through their chat template. The -v1 rows are the from-scratch counterparts for reference.
License
This model is a derivative of Meta Llama 3.2 and is distributed under the Llama 3.2 Community License (see also the Acceptable Use Policy). Attribution is provided in the NOTICE file. Built with Llama.
