DFveloper/gemma-4-26B-A4B-it-qat-q4_0-Claude-Opus-Distill-v2-GGUF
π Lumen 3.1 Pulsar QAT Γ Claude Opus 4.6 Distill
This model combines the Quantization-Aware Trained Lumen 3.1 Pulsar checkpoint with the fine-tuning delta extracted from:
`TeichAI/gemma-4-26B-A4B-it-Claude-Opus-Distill-v2`
Rather than directly fine-tuning or quantizing the TeichAI model, its learned weight changes were extracted relative to its original base model and approximated as a LoRA adapter.
That adapter was then merged into the Lumen 3.1 Pulsar QAT checkpoint.
In short:
Original Gemma 4 Base
β
β Fine-tuning
βΌ
TeichAI Claude Opus Distill
β
β Weight Diff
βΌ
Extracted LoRA
β
β Merge
βΌ
Lumen 3.1 Pulsar QAT
β
βΌ
Final ModelThe resulting model therefore retains Lumen 3.1 Pulsar as its primary checkpoint while incorporating part of the learned behavior introduced by the Claude Opus reasoning distillation.
𧬠Model Construction
The model was produced using a delta-to-LoRA transfer pipeline.
Stage 1 β Obtain the Fine-Tuning Delta
Let:
W_base = weights of the original Gemma 4 base model
W_ft = weights of the TeichAI Claude Opus distilled modelThe fine-tuning delta is:
ΞW = W_ft - W_baseThis isolates the changes introduced by the Claude Opus reasoning fine-tuning from the underlying Gemma 4 weights.
Stage 2 β Approximate the Delta with LoRA
Instead of directly merging the full dense weight difference, the delta matrices were decomposed into low-rank representations:
ΞW β B Β· Awhere A and B form a LoRA adapter.
Conceptually:
TeichAI Fine-Tuned Model
-
Original Base Model
β
βΌ
Dense Fine-Tuning Delta
β
βΌ
Low-Rank Decomposition
β
βΌ
LoRA AdapterThis allows the learned fine-tuning direction to be transferred independently of the original full checkpoint.
Stage 3 β Merge into Lumen 3.1 Pulsar QAT
The extracted LoRA was then applied to the Lumen 3.1 Pulsar QAT checkpoint:
W_final = W_lumen_qat + ΞW_lorawhere:
W_lumen_qat = Lumen 3.1 Pulsar QAT weights
ΞW_lora = reconstructed Claude Opus distillation deltaThis is fundamentally different from:
TeichAI Model
β
QAT
β
Quantized TeichAI ModelNo such pipeline was used.
Instead:
TeichAI Distill
β
βΌ
Extract learned delta
β
βΌ
Compress delta into LoRA
β
βΌ
Merge into Lumen 3.1 Pulsar QATβ‘ Why This Approach?
Directly replacing Lumen 3.1 Pulsar with the TeichAI checkpoint would discard the training and QAT characteristics already present in Lumen.
The delta-transfer approach attempts to preserve the existing Lumen checkpoint while selectively importing useful behavior learned during a separate fine-tuning run.
This has several advantages.
Preserve the Lumen base
The final checkpoint remains fundamentally based on:
Lumen 3.1 Pulsar QATrather than the TeichAI model.
This preserves the majority of Lumen's existing:
- model behavior,
- training history,
- reasoning characteristics,
- QAT adaptation,
- quantization robustness.
Transfer fine-tuning behavior
The extracted LoRA represents only the difference introduced by fine-tuning.
As a result, it can transfer portions of the Claude Opus distillation without copying the entire source model.
Maintain QAT characteristics
Because the destination checkpoint is already Quantization-Aware Trained, the merge is performed directly on top of the QAT-trained Lumen weights.
The final model should therefore be considered:
Lumen 3.1 Pulsar QAT with an externally extracted reasoning LoRA merged into it
rather than:
a QAT version of the TeichAI Claude Opus model
π§ Source of the Reasoning Delta
The transferred LoRA originates from:
TeichAI/gemma-4-26B-A4B-it-Claude-Opus-Distill-v2
That model was fine-tuned using high-effort Claude reasoning data, including:
These datasets were not necessarily used to directly train this checkpoint.
Instead, their influence is inherited indirectly through the extracted fine-tuning delta.
This distinction is important:
Datasets
β
TeichAI Fine-Tune
β
Weight Changes
β
LoRA Extraction
β
Lumen 3.1 Pulsar QAT㪠Technical Overview
The complete pipeline can be represented as:
ββββββββββββββββββββββββββββββββββ
β Original Gemma 4 Checkpointβ β
ββββββββββββββββ¬ββββββββββββββββββ
β
β
β
β
β
β SFT
β
β
βΌ
ββββββββββββββββββββββββββββββββββ
β TeichAI Claude Opus Distillβ β
ββββββββββββββββ¬ββββββββββββββββββ
β
β
β
β
β
Compute parameter β differences
β
β
βΌ
β
β
ΞW = W_ft - W_base
β
β
β
β
β
β Low-rank approximation
β
β
βΌ
ββββββββββββββββββββ
β Extracted LoRAβ
β
βββββββββ¬βββββββββββ
β
β
β
β Merge
β
βΌ
βββββββββββββββββββββββββββ
β Lumen 3.1 Pulsar QAT β
ββββββββββββββ¬βββββββββββββ
β
βΌ
ββββββββββββββββββββββββββββ
β Final Model β
ββββββββββββββββββββββββββββπΉ Delta-LoRA Extraction
For a matrix-shaped parameter:
D = W_ft - W_basethe dense delta can be approximated using a truncated low-rank decomposition such as:
D β U_r Ξ£_r V_rα΅which can then be represented in LoRA form as:
B = U_r Ξ£_r
A = V_rα΅giving:
BA β DOnly the portion of the fine-tuning delta captured by the selected rank is transferred.
Therefore, this method is intentionally an approximation of the source model's training delta, not an exact reconstruction of the entire fine-tuned model.
β οΈ Important Notes
This is not a conventional LoRA fine-tune
The LoRA adapter used here was not trained directly on Lumen 3.1 Pulsar.
It was reconstructed from the parameter difference between another model and its original base.
Therefore, the resulting behavior can differ from both:
- the original Lumen checkpoint, and
- the original TeichAI Claude Opus distilled model.
LoRA transfer is not lossless
Because the dense fine-tuning delta is projected into a limited rank:
ΞW β Low-Rank ΞWsome information is discarded.
The amount retained depends on factors such as:
- LoRA rank,
- singular-value distribution,
- target modules,
- architecture compatibility,
- merge scaling,
- differences between source and destination checkpoints.
Architecture compatibility matters
Delta transfer assumes that corresponding tensors between the source base model, source fine-tuned model, and destination model remain structurally compatible.
The technique is most meaningful when:
shape(source tensor)
=
shape(base tensor)
=
shape(destination tensor)for the transferred modules.
QAT status
The QAT component comes from Lumen 3.1 Pulsar QAT itself.
The extracted Claude Opus LoRA was derived from a separate fine-tuned model and subsequently merged into the QAT checkpoint.
Therefore the correct lineage is:
Lumen 3.1 Pulsar
β
βΌ
QAT
β
βΌ
Lumen 3.1 Pulsar QAT
β
β + Extracted Claude Opus LoRA
βΌ
Final Modelπ Intended Capabilities
The final model aims to combine characteristics from both sources.
From Lumen 3.1 Pulsar QAT:
- general reasoning,
- instruction following,
- coding,
- efficient inference,
- quantization robustness,
- Lumen-specific model behavior.
From the extracted Claude Opus distillation delta:
- structured reasoning,
- complex task decomposition,
- coding-oriented reasoning,
- analytical responses,
- high-effort problem-solving behavior.
The exact degree to which each capability transfers depends on the extracted LoRA rank and the interaction between the source delta and the Lumen checkpoint.
π Evaluation
Benchmarks are currently being evaluated.
Particularly useful comparisons include:
Lumen 3.1 Pulsar QAT
vs.
Lumen 3.1 Pulsar QAT + Claude Opus Delta-LoRA
vs.
Original TeichAI Claude Opus DistillRecommended metrics include:
Because this is a cross-checkpoint delta transfer, direct evaluation against the unmodified Lumen checkpoint is particularly important.
π οΈ Usage
Use the same inference stack and chat template as the corresponding Lumen 3.1 Pulsar checkpoint.
Example with Transformers:
from transformers import AutoProcessor, AutoModelForCausalLM
MODEL_ID = "YOUR_MODEL_ID"
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID,
dtype="auto",
device_map="auto",
)
messages = [
{
"role": "user",
"content": "Explain why low-rank adaptation can approximate a fine-tuning delta."
}
]
text = processor.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = processor(
text=text,
return_tensors="pt",
).to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=1024,
)
response = processor.decode(
outputs[0][inputs["input_ids"].shape[-1]:],
skip_special_tokens=False,
)
print(response)π Acknowledgements
- Google β for the Gemma architecture.
- TeichAI β for releasing the Claude Opus reasoning-distilled Gemma checkpoint used as the source of the fine-tuning delta.
- Crownelius β for the Opus reasoning dataset used in the source model.
- Unsloth β for efficient Gemma fine-tuning tooling and ecosystem support.
π Model Lineage
SOURCE BRANCH
unsloth/gemma-4-26B-A4B-it
β
βΌ
Claude Opus Reasoning Fine-Tuning
β
βΌ
TeichAI/gemma-4-26B-A4B-it-Claude-Opus-Distill-v2
β
β subtract original checkpoint
βΌ
Fine-Tuning Delta
β
β low-rank decomposition
βΌ
Extracted LoRA
β
β
ββββββββββββββββββββββ
β
DESTINATION BRANCH
β
Lumen 3.1 Pulsar
β
βΌ
QAT
β
βΌ
Lumen 3.1 Pulsar QAT
β
β + LoRA Merge
βΌ
Final ModelCitation
If you use this model, please also credit the upstream checkpoints and datasets from which the transferred training delta was derived.
