CatQualia/gnarp-m2
gnarp-m2
A 360M-parameter language model fine-tuned via QLoRA on 69,945 cross-domain isomorphism and verification-labeled instruction pairs (74,395 rows read; every row labelled negative is discarded by the trainer). gnarp-m2 specializes in cross-domain structural transfer — mapping mechanisms from one domain (biology, physics, anime, economics, etc.) to software engineering constructs.
Model Details
Training Data
Corpus: clean_corpus_v5.jsonl — 74,395 rows read, of which 69,945 were train-eligible
Label composition, measured: 11,669 rows labelled positive, 4,450 labelled negative, 58,276 unlabelled. Positive and unlabelled rows only; negative rows are excluded from SFT targets — and in this pipeline they are dropped outright, before the normalised corpus is written (run_qlora_gnarpm2.py: if r.get("label") == "negative": continue). So all 4,450 refuted rows were excluded from the loss, and the model did not train on them as context, as masked targets, or as contrastive pairs.
Training Configuration
Training Results
Evaluation: Cross-Domain Transfer Benchmark v2
36 cross-domain transfer tasks spanning anime, biology, physics, economics, fiction, geography, music, cooking, ecology, martial arts, psychology, logistics, chemistry, sports, agriculture, linguistics, city planning, finance, navigation, and architecture.
Scoring: Heuristic rubric (keyword + structural analysis). Trust DELTAS between models on the same tasks, not absolutes.
Prior Model Lineage (heldout benchmark, qwen3:8b judge)
m2 scored on transfer_benchmark_v2 (heuristic-only), not the qwen3:8b-judged heldout benchmark. Cross-benchmark comparisons should be treated with caution. The lineage is also not monotonic* on this table: v1 (0.218) scores far below base (0.709). The 10%→26%→31%→37% progression quoted elsewhere refers to a different eval (anime→architecture, two LLM judges) and should not be read as the heldout benchmark.
m2's refusal rate was never measured. The row is blank because no computation was performed, not because the result was unfavourable. The 0.130→0.400 figure sometimes quoted alongside this model belongs to the earlier lineage (base→v3) and does not describe m2. Calibration drift for m2 is therefore unmeasured.
Corrections
Two claims previously made about this model did not survive re-reading its own artefacts, and are recorded here rather than quietly edited.
- "Trained on data that includes the claims the system refuted against itself." False, and the opposite of what happens. The trainer discards every row labelled negative before the normalised corpus is written; 4,450 of 74,395 rows carry that label. The model never saw a refuted example, so it cannot have learnt "the shape of a wrong answer" from one. Any downstream claim resting on contrastive training over refutations does not hold for this checkpoint.
- "failure_class from the 17,801-row moat." The moat corpus does carry a non-empty
failure_classon all 17,801 of its rows. The corpus that trained this model carries the field on zero of its rows, and the normaliser emits onlyinstruction/input/output, so such a label could not have reached the loss even if present. The only per-row taxonomy in the corpus isverdict_raw, with four values (VERIFIED/GENERATIVE 7,758; REFUTED/DECORATIVE 4,231; GENERATIVE 3,160; REFUTED 12). The fine-grained error grammar is absent from the training signal, not collapsed within it.
- Row count. The headline figure of 74,395 is the pre-filter count; 69,945 rows were train-eligible. Every row in the corpus also carries
train_eligible: true, including the 4,450 the trainer discards — two mechanisms in one pipeline disagreeing about the same fact.
Not yet run: the ablation that would separate the value of the corpus structure from the value of the labels — same rows, failure labels stripped or shuffled, same benchmark. It cannot be run as originally stated, because the labels it would strip are not in the corpus.
Limitations
- 360M parameters. Small model. Cannot match larger models on complex reasoning, long-form generation, or nuanced instruction following.
- Single GPU, single epoch. Trained on consumer hardware (RTX 3080 8GB) for one epoch. More training could improve results but risks overfitting.
- Heuristic eval. The transfer benchmark v2 uses keyword/structural heuristic scoring, not a strong LLM judge. The +14.1% delta is directionally meaningful but not precisely calibrated.
- Cross-benchmark caveat. v1/v2/v3 were scored with a qwen3:8b judge; m2 was scored with heuristic-only. Direct numerical comparison across the two benchmarks is not valid.
- Domain-specific training data. Over 71% of training data is isomorphism pairs. The model is optimized for cross-domain structural transfer and may underperform on general chat or coding tasks.
- No safety fine-tuning beyond refusal data. The model includes 19 refusal pairs but is not extensively safety-tuned.
How to Use
With Ollama (recommended for local inference)
# Create the Modelfile
cat > Modelfile << 'EOF'
FROM ./model/merged_gnarpm2
TEMPLATE """### Instruction: {{ .Prompt }} ### Response: """
PARAMETER num_ctx 4096
PARAMETER temperature 0.3
PARAMETER num_predict 512
SYSTEM You are gnarp-m2, a cross-domain transfer specialist.
EOF
ollama create gnarp-m2 -f Modelfile
ollama run gnarp-m2With Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "gnarp/gnarp-m2" # or local path
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
prompt = "### Instruction:\nApply the concept of biological apoptosis to software deployment strategy.\n### Response:\n"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512, temperature=0.3)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))With PEFT (adapter only)
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained("HuggingFaceTB/SmolLM2-360M-Instruct")
model = PeftModel.from_pretrained(base, "path/to/adapter_gnarpm2")
model = model.merge_and_unload()License
- Model weights and code: CatQualia Open-Or-Pay License (COPL) v1.0 — free for research, individuals and academia with attribution; commercial use requires either opening the derivative stack or a commercial waiver.
- Training data (corpus): CatQualia Structural Isomorphism License (CSIL) v3.0 — derived from the WaveMotionExpansion isomorphism engine.
- Base model: Apache 2.0 (SmolLM2-360M-Instruct by HuggingFace)
Full terms in this repository: LICENSE (CSIL v3.0) and COPL-LICENSE.md (COPL v1.0). Both also at https://catqualia.com/licensing
Citation
If you use gnarp-m2, please cite the model together with the publication that describes the corpus it was trained on.
@model{gnarp-m2,
title={gnarp-m2: Cross-Domain Transfer Fine-Tuned Language Model},
author={Betances, Christopher},
year={2026},
base_model={HuggingFaceTB/SmolLM2-360M-Instruct},
method={QLoRA},
training_rows={69945},
corpus_rows_read={74395},
transfer_benchmark={0.7839},
url={https://huggingface.co/CatQualia/gnarp-m2}
}Related records:
- Training corpus — Self-Falsifying Research Infrastructure Void Atlas: https://doi.org/10.5281/zenodo.22751797
- OmniLingua Interpreter Core (prompt architecture the corpus derives from): https://doi.org/10.5281/zenodo.22751511
- The complete CatQualia record — 135 dated defensive publications, each with its own DOI: https://doi.org/10.5281/zenodo.22751107
