axiom-of-choice/gemma-3-4b-es-reasoning-qlora
Gemma-3-4B Spanish-reasoning QLoRA
Also published for `transformers`/`peft`: `axiom-of-choice/gemma-3-4b-es-reasoning-peft` -- same adapter weights, verified against this one (phase-2 parity numbers on the linked repo's card). Use this repo withmlx-lm; use the linked one fortransformers,vLLM, or anything else.
A LoRA adapter trained on the 4-bit (QAT) base — i.e. QLoRA — that teaches mlx-community/gemma-3-4b-it-qat-4bit to produce a chain-of-thought in Spanish, a capability the base model does not have at all. On a held-out Spanish GSM8K split it raises accuracy from 41.0% to 68.9% (+27.9 points) while adding the Spanish reasoning. Both numbers are measured on the full 283-row split below, not recalled.
Unlike a model with a native thinking mode, Gemma-3 treats <think> as plain text, so the base model simply answers directly and never reasons. This adapter makes it open a <think>...</think> block and reason through the problem in Spanish before answering.
Published at `scale 5.6` (0.7x of the 8 it was trained at). scale multiplies the LoRA update directly in mlx-lm, so lowering it at inference interpolates between the base model and the fine-tune using the same weights. A sweep (below) found 0.7x is the accuracy peak: strong enough to reason, not so strong it fails to terminate.
What it does, measured
All 283 rows of the held-out Spanish GSM8K split. Identical sampling for both models: temp 0.0 (greedy, deterministic), max_tokens 2048, repetition_penalty 1.1, batch 16, no \boxed{} hint.
The base model reasons on 0% of prompts (it has no thinking mode) and scores 41.0%. The adapter reasons on 90.1%, in Spanish (mean confidence 0.9966), and scores 68.9%.
Paired McNemar exact test on the 119 rows where the two disagree: p < 0.0001 (20 base-only wins vs 99 adapter-only). The gain is a decisive paired difference, not sampling noise.
Spanish is intact (and the two flagged rows are not real failures)
langdetect flagged 2 of the 255 reasoned traces below 0.90 Spanish confidence. Both were read by hand and neither is a genuine failure: one is pure Spanish misscored by the known Spanish/Portuguese langdetect confusion; the other reasons in Spanish to the correct answer, then appends an English symbolic re-derivation of the same result. The effective real Spanish-failure rate is 0 of 255. The crude flagged count is reported here rather than hidden.
The remaining errors are termination, not math
When the adapter fails, it is almost always because it entered a loop and ran to the token ceiling without closing </think> (8.1% of all rows), not because it reasoned to a wrong answer — reasoned-only accuracy is 74.5%. Raising repetition_penalty above the 1.1 used here, or raising max_tokens, may recover some of those rows for your use case.
Usage (mlx-lm)
from mlx_lm import load, generate
model, tokenizer = load(
"mlx-community/gemma-3-4b-it-qat-4bit",
adapter_path="<downloaded adapter dir>", # this repo
)
messages = [
{"role": "system", "content":
"Eres un asistente experto que resuelve problemas pensando y "
"explicando completamente en espanol."},
{"role": "user", "content": "¿Cuánto es 27 x 14? Piensa paso a paso."},
]
prompt = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
print(generate(model, tokenizer, prompt, max_tokens=2048))The adapter ships adapter_config.json with scale 5.6 already set, so mlx_lm.load(..., adapter_path=...) applies the published scale automatically.
Training
Distillation. The reasoning traces come from bespoke-stratos-es: each source question was re-solved from scratch by DeepSeek V4 Flash, generating the chain-of-thought natively in Spanish (not machine-translated from the English traces). So this is a distillation — Gemma-3-4B learns to reason in Spanish from a DeepSeek V4 Flash teacher, not from its own base behaviour (which does not reason at all).
License and use
This is a derivative of Google's Gemma-3, distributed under the Gemma license. Use of this adapter with the base model is subject to Google's Gemma Terms of Use and Prohibited Use Policy. It is not an Apache-2.0 pass-through.
Provenance
Every figure above is computed from measured artifacts by scripts/upload_adapter_gemma.py; the full experiment log, including the scale sweep and the manual review of the flagged rows, is in EXPERIMENTS-GEMMA3-4B.md in the source repository.
Gemma-3-4B QLoRA de razonamiento en español
También publicado para `transformers`/`peft`: `axiom-of-choice/gemma-3-4b-es-reasoning-peft` -- mismos pesos del adaptador, verificados contra este (números de paridad de fase 2 en la ficha del repo enlazado). Usa este repo conmlx-lm; usa el enlazado paratransformers,vLLM, o cualquier otra cosa.
Un adaptador LoRA entrenado sobre el modelo base en 4 bits (QAT) — es decir, QLoRA — que enseña a mlx-community/gemma-3-4b-it-qat-4bit a producir su cadena de razonamiento en español, una capacidad que el modelo base no tiene en absoluto. Sobre un split de GSM8K en español fuera del entrenamiento, sube la precisión de 41.0% a 68.9% (+27.9 puntos) a la vez que añade el razonamiento en español. Ambos números están medidos sobre las 283 filas completas del split, no recordados.
A diferencia de un modelo con modo de pensamiento nativo, Gemma-3 trata <think> como texto plano, así que el modelo base simplemente responde directo y nunca razona. Este adaptador hace que abra un bloque <think>...</think> y razone el problema en español antes de responder.
Publicado con `scale 5.6` (0.7x del 8 con el que se entrenó). En mlx-lm scale multiplica directamente la actualización de bajo rango, así que bajarlo en inferencia interpola entre el modelo base y el fine-tune usando los mismos pesos. Un barrido (abajo) encontró que 0.7x es el pico de precisión: suficientemente fuerte para razonar, no tanto como para no terminar.
Qué hace, medido
Las 283 filas del split de GSM8K en español, fuera del entrenamiento. Muestreo idéntico para ambos modelos: temp 0.0 (greedy, determinista), max_tokens 2048, repetition_penalty 1.1, batch 16, sin pista de \boxed{}.
El modelo base razona en el 0% de los prompts (no tiene modo de pensamiento) y puntúa 41.0%. El adaptador razona en el 90.1%, en español (confianza media 0.9966), y puntúa 68.9%.
Test emparejado de McNemar exacto sobre las 119 filas en las que ambos discrepan: p < 0.0001 (20 aciertos exclusivos del base contra 99 del adaptador). La mejora es una diferencia emparejada decisiva, no ruido de muestreo.
El español está intacto (y las dos filas marcadas no son fallos reales)
langdetect marcó 2 de las 255 trazas razonadas por debajo de 0.90 de confianza en español. Ambas se revisaron a mano y ninguna es un fallo genuino: una es español puro mal puntuado por la conocida confusión español/portugués de langdetect; la otra razona en español hasta la respuesta correcta y luego añade una re-derivación simbólica del mismo resultado en inglés. La tasa real de fallo de español es 0 de 255. Se reporta aquí el conteo crudo en vez de esconderlo.
Los errores restantes son de terminación, no de matemáticas
Cuando el adaptador falla, casi siempre es porque entró en un bucle y llegó al techo de tokens sin cerrar </think> (8.1% de todas las filas), no porque razonara hacia una respuesta equivocada — la precisión solo-razonadas es 74.5%. Subir repetition_penalty por encima del 1.1 usado aquí, o subir max_tokens, puede recuperar algunas de esas filas para tu caso.
Destilación
Las trazas de razonamiento vienen de bespoke-stratos-es: cada pregunta de origen fue re-resuelta desde cero por DeepSeek V4 Flash, generando la cadena de pensamiento nativamente en español (no traducida a máquina de las trazas en inglés). Así que esto es una destilación: Gemma-3-4B aprende a razonar en español de un teacher DeepSeek V4 Flash, no de su comportamiento base (que no razona en absoluto).
Licencia y uso
Esto es un derivado de Gemma-3 de Google, distribuido bajo la licencia Gemma. El uso de este adaptador con el modelo base está sujeto a los Términos de Uso de Gemma de Google y a su Política de Uso Prohibido. No es un traspaso limpio de Apache-2.0.
