CoolFace
Modelpublic

DFveloper/gemma-4-26B-A4B-it-qat-q4_0-Claude-Opus-Distill-v2-GGUF

sourceHugging Faceapache-2.0updated 16d agoView on Hugging Face
5likes1.1kdownloads
Model Card

🌟 Lumen 3.1 Pulsar QAT Γ— Claude Opus 4.6 Distill

This model combines the Quantization-Aware Trained Lumen 3.1 Pulsar checkpoint with the fine-tuning delta extracted from:

`TeichAI/gemma-4-26B-A4B-it-Claude-Opus-Distill-v2`

Rather than directly fine-tuning or quantizing the TeichAI model, its learned weight changes were extracted relative to its original base model and approximated as a LoRA adapter.

That adapter was then merged into the Lumen 3.1 Pulsar QAT checkpoint.

In short:

text
Original Gemma 4 Base
        β”‚
        β”‚ Fine-tuning
        β–Ό
TeichAI Claude Opus Distill
        β”‚
        β”‚ Weight Diff
        β–Ό
Extracted LoRA
        β”‚
        β”‚ Merge
        β–Ό
Lumen 3.1 Pulsar QAT
        β”‚
        β–Ό
Final Model

The resulting model therefore retains Lumen 3.1 Pulsar as its primary checkpoint while incorporating part of the learned behavior introduced by the Claude Opus reasoning distillation.


🧬 Model Construction

The model was produced using a delta-to-LoRA transfer pipeline.

Stage 1 β€” Obtain the Fine-Tuning Delta

Let:

text
W_base = weights of the original Gemma 4 base model
W_ft   = weights of the TeichAI Claude Opus distilled model

The fine-tuning delta is:

text
Ξ”W = W_ft - W_base

This isolates the changes introduced by the Claude Opus reasoning fine-tuning from the underlying Gemma 4 weights.


Stage 2 β€” Approximate the Delta with LoRA

Instead of directly merging the full dense weight difference, the delta matrices were decomposed into low-rank representations:

text
Ξ”W β‰ˆ B Β· A

where A and B form a LoRA adapter.

Conceptually:

text
TeichAI Fine-Tuned Model
        -
Original Base Model
        β”‚
        β–Ό
Dense Fine-Tuning Delta
        β”‚
        β–Ό
Low-Rank Decomposition
        β”‚
        β–Ό
LoRA Adapter

This allows the learned fine-tuning direction to be transferred independently of the original full checkpoint.


Stage 3 β€” Merge into Lumen 3.1 Pulsar QAT

The extracted LoRA was then applied to the Lumen 3.1 Pulsar QAT checkpoint:

text
W_final = W_lumen_qat + Ξ”W_lora

where:

text
W_lumen_qat = Lumen 3.1 Pulsar QAT weights
Ξ”W_lora     = reconstructed Claude Opus distillation delta

This is fundamentally different from:

text
TeichAI Model
    ↓
QAT
    ↓
Quantized TeichAI Model

No such pipeline was used.

Instead:

text
TeichAI Distill
      β”‚
      β–Ό
Extract learned delta
      β”‚
      β–Ό
Compress delta into LoRA
      β”‚
      β–Ό
Merge into Lumen 3.1 Pulsar QAT

⚑ Why This Approach?

Directly replacing Lumen 3.1 Pulsar with the TeichAI checkpoint would discard the training and QAT characteristics already present in Lumen.

The delta-transfer approach attempts to preserve the existing Lumen checkpoint while selectively importing useful behavior learned during a separate fine-tuning run.

This has several advantages.

Preserve the Lumen base

The final checkpoint remains fundamentally based on:

text
Lumen 3.1 Pulsar QAT

rather than the TeichAI model.

This preserves the majority of Lumen's existing:

  • β€”model behavior,
  • β€”training history,
  • β€”reasoning characteristics,
  • β€”QAT adaptation,
  • β€”quantization robustness.

Transfer fine-tuning behavior

The extracted LoRA represents only the difference introduced by fine-tuning.

As a result, it can transfer portions of the Claude Opus distillation without copying the entire source model.

Maintain QAT characteristics

Because the destination checkpoint is already Quantization-Aware Trained, the merge is performed directly on top of the QAT-trained Lumen weights.

The final model should therefore be considered:

Lumen 3.1 Pulsar QAT with an externally extracted reasoning LoRA merged into it

rather than:

a QAT version of the TeichAI Claude Opus model

🧠 Source of the Reasoning Delta

The transferred LoRA originates from:

TeichAI/gemma-4-26B-A4B-it-Claude-Opus-Distill-v2

That model was fine-tuned using high-effort Claude reasoning data, including:

DatasetPurpose
TeichAI/Claude-Opus-4.6-Reasoning-887xClaude Opus 4.6 reasoning trajectories
TeichAI/claude-4.5-opus-high-reasoning-250xClaude Opus high-effort reasoning
Crownelius/Opus-4.6-Reasoning-2100x-formattedAdditional formatted Opus reasoning trajectories

These datasets were not necessarily used to directly train this checkpoint.

Instead, their influence is inherited indirectly through the extracted fine-tuning delta.

This distinction is important:

text
Datasets
   ↓
TeichAI Fine-Tune
   ↓
Weight Changes
   ↓
LoRA Extraction
   ↓
Lumen 3.1 Pulsar QAT

πŸ”¬ Technical Overview

The complete pipeline can be represented as:

text
             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
             β”‚ Original Gemma 4 Checkpoint  β”‚
             β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β€…β€…                          β”‚
β€…β€…                          β”‚ SFT
β€…β€…                          β–Ό
             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
             β”‚ TeichAI Claude Opus Distill  β”‚
             β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β€…β€…                          β”‚
β€…β€…        Compute parameter β”‚ differences
β€…β€…                          β–Ό
β€…β€…                Ξ”W = W_ft - W_base
β€…β€…                          β”‚
β€…β€…                          β”‚ Low-rank approximation
β€…β€…                          β–Ό
                   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                   β”‚ Extracted LoRAβ€… β”‚
                   β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
 β€…                         β”‚
 β€…                         β”‚ Merge
 β€…                         β–Ό
              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
              β”‚ Lumen 3.1 Pulsar QAT  β”‚
              β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                          β”‚
                          β–Ό
              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
              β”‚      Final Model       β”‚
              β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

πŸ”Ή Delta-LoRA Extraction

For a matrix-shaped parameter:

text
D = W_ft - W_base

the dense delta can be approximated using a truncated low-rank decomposition such as:

text
D β‰ˆ U_r Ξ£_r V_rα΅€

which can then be represented in LoRA form as:

text
B = U_r Ξ£_r
A = V_rα΅€

giving:

text
BA β‰ˆ D

Only the portion of the fine-tuning delta captured by the selected rank is transferred.

Therefore, this method is intentionally an approximation of the source model's training delta, not an exact reconstruction of the entire fine-tuned model.


⚠️ Important Notes

This is not a conventional LoRA fine-tune

The LoRA adapter used here was not trained directly on Lumen 3.1 Pulsar.

It was reconstructed from the parameter difference between another model and its original base.

Therefore, the resulting behavior can differ from both:

  • β€”the original Lumen checkpoint, and
  • β€”the original TeichAI Claude Opus distilled model.

LoRA transfer is not lossless

Because the dense fine-tuning delta is projected into a limited rank:

text
Ξ”W β†’ Low-Rank Ξ”W

some information is discarded.

The amount retained depends on factors such as:

  • β€”LoRA rank,
  • β€”singular-value distribution,
  • β€”target modules,
  • β€”architecture compatibility,
  • β€”merge scaling,
  • β€”differences between source and destination checkpoints.

Architecture compatibility matters

Delta transfer assumes that corresponding tensors between the source base model, source fine-tuned model, and destination model remain structurally compatible.

The technique is most meaningful when:

text
shape(source tensor)
=
shape(base tensor)
=
shape(destination tensor)

for the transferred modules.


QAT status

The QAT component comes from Lumen 3.1 Pulsar QAT itself.

The extracted Claude Opus LoRA was derived from a separate fine-tuned model and subsequently merged into the QAT checkpoint.

Therefore the correct lineage is:

text
Lumen 3.1 Pulsar
      β”‚
      β–Ό
     QAT
      β”‚
      β–Ό
Lumen 3.1 Pulsar QAT
      β”‚
      β”‚ + Extracted Claude Opus LoRA
      β–Ό
Final Model

🌟 Intended Capabilities

The final model aims to combine characteristics from both sources.

From Lumen 3.1 Pulsar QAT:

  • β€”general reasoning,
  • β€”instruction following,
  • β€”coding,
  • β€”efficient inference,
  • β€”quantization robustness,
  • β€”Lumen-specific model behavior.

From the extracted Claude Opus distillation delta:

  • β€”structured reasoning,
  • β€”complex task decomposition,
  • β€”coding-oriented reasoning,
  • β€”analytical responses,
  • β€”high-effort problem-solving behavior.

The exact degree to which each capability transfers depends on the extracted LoRA rank and the interaction between the source delta and the Lumen checkpoint.


πŸ“Š Evaluation

Benchmarks are currently being evaluated.

Particularly useful comparisons include:

text
Lumen 3.1 Pulsar QAT
vs.
Lumen 3.1 Pulsar QAT + Claude Opus Delta-LoRA
vs.
Original TeichAI Claude Opus Distill

Recommended metrics include:

MetricPurpose
PerplexityDetect general degradation
Top-1 token agreementMeasure behavioral changes
KL divergenceCompare output distributions
MMLU-ProGeneral reasoning
GPQAAdvanced scientific reasoning
AIMEMathematical reasoning
LiveCodeBenchCoding performance
IFEvalInstruction following

Because this is a cross-checkpoint delta transfer, direct evaluation against the unmodified Lumen checkpoint is particularly important.


πŸ› οΈ Usage

Use the same inference stack and chat template as the corresponding Lumen 3.1 Pulsar checkpoint.

Example with Transformers:

python
from transformers import AutoProcessor, AutoModelForCausalLM

MODEL_ID = "YOUR_MODEL_ID"

processor = AutoProcessor.from_pretrained(MODEL_ID)

model = AutoModelForCausalLM.from_pretrained(
    MODEL_ID,
    dtype="auto",
    device_map="auto",
)

messages = [
    {
        "role": "user",
        "content": "Explain why low-rank adaptation can approximate a fine-tuning delta."
    }
]

text = processor.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)

inputs = processor(
    text=text,
    return_tensors="pt",
).to(model.device)

outputs = model.generate(
    **inputs,
    max_new_tokens=1024,
)

response = processor.decode(
    outputs[0][inputs["input_ids"].shape[-1]:],
    skip_special_tokens=False,
)

print(response)

πŸ™ Acknowledgements

  • β€”Google β€” for the Gemma architecture.
  • β€”TeichAI β€” for releasing the Claude Opus reasoning-distilled Gemma checkpoint used as the source of the fine-tuning delta.
  • β€”Crownelius β€” for the Opus reasoning dataset used in the source model.
  • β€”Unsloth β€” for efficient Gemma fine-tuning tooling and ecosystem support.

πŸ“– Model Lineage

text
                         SOURCE BRANCH

unsloth/gemma-4-26B-A4B-it
             β”‚
             β–Ό
Claude Opus Reasoning Fine-Tuning
             β”‚
             β–Ό
TeichAI/gemma-4-26B-A4B-it-Claude-Opus-Distill-v2
             β”‚
             β”‚ subtract original checkpoint
             β–Ό
      Fine-Tuning Delta
             β”‚
             β”‚ low-rank decomposition
             β–Ό
       Extracted LoRA
             β”‚
             β”‚
             └────────────────────┐
                                  β”‚
                         DESTINATION BRANCH
                                  β”‚
                       Lumen 3.1 Pulsar
                                  β”‚
                                  β–Ό
                                 QAT
                                  β”‚
                                  β–Ό
                       Lumen 3.1 Pulsar QAT
                                  β”‚
                                  β”‚ + LoRA Merge
                                  β–Ό
                              Final Model

Citation

If you use this model, please also credit the upstream checkpoints and datasets from which the transferred training delta was derived.