CoolFace
Modelpublic

axiom-of-choice/gemma-3-4b-es-reasoning-peft

sourceHugging Facegemmaupdated 2mo agoView on Hugging Face
1likes10downloads
Model Card

gemma-3-4b-es-reasoning-peft

A transformers/peft-compatible port of `axiom-of-choice/gemma-3-4b-es-reasoning-qlora`, a rank-8 LoRA adapter trained in MLX that makes `google/gemma-3-4b-it-qat-q4_0-unquantized` reason in Spanish. Same adapter weights, converted losslessly into the shapes and names peft expects -- every LoRA update matrix (B @ A) verified equal to the MLX original at rtol=1e-5. If you use mlx-lm, get the original repo instead; this one is for everyone else (transformers, vLLM with --enable-lora, anything built on peft).

Usage

python
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base = AutoModelForCausalLM.from_pretrained("google/gemma-3-4b-it-qat-q4_0-unquantized", dtype="bfloat16")
model = PeftModel.from_pretrained(base, "axiom-of-choice/gemma-3-4b-es-reasoning-peft")
tokenizer = AutoTokenizer.from_pretrained("google/gemma-3-4b-it-qat-q4_0-unquantized")

messages = [{"role": "user", "content": "¿Cuánto es 17 por 24?"}]
inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt"
)
out = model.generate(inputs, max_new_tokens=1024)
print(tokenizer.decode(out[0], skip_special_tokens=True))

Verified against the original (phase 2 parity)

Format conversion (key names, tensor orientation, scale→lora_alpha) is checked with no model load and proves the files are self-consistent, but not that this repo behaves like the original. This table is a real generation run: 20 items from the held-out Spanish GSM8K eval, this repo (google/gemma-3-4b-it-qat-q4_0-unquantized, bf16) paired item-by-item against the original (mlx-community/gemma-3-4b-it-qat-4bit, MLX -- quantized MLX build of google/gemma-3-4b-it-qat-q4_0-unquantized).

accuracyagreement
this repo75.0%80.0%
original (mlx-community/gemma-3-4b-it-qat-4bit, MLX)75.0%

Paired counts: 2 correct only here, 2 correct only in the original, out of 20. Clean parity: the 2 vs 2 disagreement has no direction and is within what greedy decoding's own numerical noise produces between MLX and PyTorch.

What "converted" is not

Only the adapter tensors are provably identical to the original. The base model is not always the same weights: MLX may run a quantized base (dequantized on the fly) where this repo runs bf16 natively, and MLX and PyTorch are different numerical engines even on identical weights. The table above is the actual measurement, not an assumption that identical adapter math implies identical behavior.

Training

Same recipe, dataset and MLX-side evaluation as the original -- see `axiom-of-choice/gemma-3-4b-es-reasoning-qlora` for the full training card.


En español

Puerto a transformers/peft de `axiom-of-choice/gemma-3-4b-es-reasoning-qlora`, un adaptador LoRA de rango 8 entrenado en MLX que hace que `google/gemma-3-4b-it-qat-q4_0-unquantized` razone en español. Mismos pesos del adaptador, convertidos sin pérdida al formato que espera peft -- cada matriz de actualización LoRA verificada igual al original MLX con rtol=1e-5.

Verificado contra el original (paridad de fase 2)

precisiónacuerdo
este repo75.0%80.0%
original (mlx-community/gemma-3-4b-it-qat-4bit, MLX)75.0%

Conteos pareados: 2 correctos solo aquí, 2 correctos solo en el original, de 20. Paridad limpia: el desacuerdo de 2 contra 2 no tiene dirección y cae dentro del ruido numérico propio de la decodificación greedy entre MLX y PyTorch.

Misma receta de entrenamiento y evaluación que el original -- ver `axiom-of-choice/gemma-3-4b-es-reasoning-qlora` para la ficha completa.