axiom-of-choice/gemma-3-4b-es-reasoning-peft
gemma-3-4b-es-reasoning-peft
A transformers/peft-compatible port of `axiom-of-choice/gemma-3-4b-es-reasoning-qlora`, a rank-8 LoRA adapter trained in MLX that makes `google/gemma-3-4b-it-qat-q4_0-unquantized` reason in Spanish. Same adapter weights, converted losslessly into the shapes and names peft expects -- every LoRA update matrix (B @ A) verified equal to the MLX original at rtol=1e-5. If you use mlx-lm, get the original repo instead; this one is for everyone else (transformers, vLLM with --enable-lora, anything built on peft).
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained("google/gemma-3-4b-it-qat-q4_0-unquantized", dtype="bfloat16")
model = PeftModel.from_pretrained(base, "axiom-of-choice/gemma-3-4b-es-reasoning-peft")
tokenizer = AutoTokenizer.from_pretrained("google/gemma-3-4b-it-qat-q4_0-unquantized")
messages = [{"role": "user", "content": "¿Cuánto es 17 por 24?"}]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt"
)
out = model.generate(inputs, max_new_tokens=1024)
print(tokenizer.decode(out[0], skip_special_tokens=True))Verified against the original (phase 2 parity)
Format conversion (key names, tensor orientation, scale→lora_alpha) is checked with no model load and proves the files are self-consistent, but not that this repo behaves like the original. This table is a real generation run: 20 items from the held-out Spanish GSM8K eval, this repo (google/gemma-3-4b-it-qat-q4_0-unquantized, bf16) paired item-by-item against the original (mlx-community/gemma-3-4b-it-qat-4bit, MLX -- quantized MLX build of google/gemma-3-4b-it-qat-q4_0-unquantized).
Paired counts: 2 correct only here, 2 correct only in the original, out of 20. Clean parity: the 2 vs 2 disagreement has no direction and is within what greedy decoding's own numerical noise produces between MLX and PyTorch.
What "converted" is not
Only the adapter tensors are provably identical to the original. The base model is not always the same weights: MLX may run a quantized base (dequantized on the fly) where this repo runs bf16 natively, and MLX and PyTorch are different numerical engines even on identical weights. The table above is the actual measurement, not an assumption that identical adapter math implies identical behavior.
Training
Same recipe, dataset and MLX-side evaluation as the original -- see `axiom-of-choice/gemma-3-4b-es-reasoning-qlora` for the full training card.
En español
Puerto a transformers/peft de `axiom-of-choice/gemma-3-4b-es-reasoning-qlora`, un adaptador LoRA de rango 8 entrenado en MLX que hace que `google/gemma-3-4b-it-qat-q4_0-unquantized` razone en español. Mismos pesos del adaptador, convertidos sin pérdida al formato que espera peft -- cada matriz de actualización LoRA verificada igual al original MLX con rtol=1e-5.
Verificado contra el original (paridad de fase 2)
Conteos pareados: 2 correctos solo aquí, 2 correctos solo en el original, de 20. Paridad limpia: el desacuerdo de 2 contra 2 no tiene dirección y cae dentro del ruido numérico propio de la decodificación greedy entre MLX y PyTorch.
Misma receta de entrenamiento y evaluación que el original -- ver `axiom-of-choice/gemma-3-4b-es-reasoning-qlora` para la ficha completa.
