mergeability
zoo-curriculum-beetle-bilingual-l2-50-simultaneous-b2-fineweb-100m-ita-eng__beetle-bili-a14287f4zoo-language-beetle-monolingual-fineweb-100m-ell__beetle-monolingual-fineweb-100m-islzoo-curriculum-beetle-bilingual-l2-50-simultaneous-b2-fineweb-100m-rus-eng__beetle-bili-27e39cb9zoo-curriculum-beetle-bilingual-balanced-b1-fineweb-100m-nld-eng__beetle-bilingual-l2-5-98710115zoo-curriculum-beetle-bilingual-balanced-b1-fineweb-100m-tur-eng__beetle-bilingual-l2-5-c2c7275fzoo-curriculum-beetle-bilingual-l2-50-simultaneous-b2-fineweb-100m-fil-eng__beetle-bili-b638e8aczoo-curriculum-beetle-bilingual-balanced-b1-humanscale-deu-eng__beetle-bilingual-l2-50-59c91440beetle-humanscale-nld-eng
aim-activation-informed-merging
AIM: does activation-informed merging change what makes a merge work?
Headline
AIM does exactly what it claims, the targeting is what makes it work — and it changes
nothing about what predicts a good merge.
AIM is exactly what it says on the tin, and that is verifiable from public artefacts alone.
The published with-AIM checkpoints are recovered, to R² = 0.9992, as a closed-form
per-input-channel shrinkage of their baseline twins toward the base model, with ω̂ =… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/aim-activation-informed-merging.beetle-merge-eval
Beetle merged models — benchmark evaluation against their parents
Minimal-pair benchmark accuracy for the Beetle merged models published in the
Mergeability org, scored against their own parent models and, where one
exists, the jointly-trained ceiling on the same harness.
The existing sweep datasets (Mergeability/merge-sweep-results,
Mergeability-2/mergeability-results) record merge quality in nats (NLL,
delta_floor, rel_damage, barrier, geometry). They contain no downstream… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/beetle-merge-eval.compose-audit
Compose-audit: putting the alignment map and the merging payoff on the SAME real models
Generated 2026-08-26 23:25 UTC · training-free · code: /root/compose-audit · operators/aligners/metrics imported unmodified from mergeschool.core (/root/mergeability, treated as read-only).
Read this first: what substrate, and what metric
SET 1
SET 4
Substrate
EleutherAI/pythia-{14m,31m,70m,160m,410m}-seed{1..9} (PolyPythia) — real reseeded LMs… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/compose-audit.mergebench-property-sweep
MergeBench property sweep
Pre-merge pairwise properties for every mergeable pair in the
MergeBench suite (40 checkpoints, 8 base families, 5 domains),
computed with the metric panel behind Figure F8 of the Heterogeneous Mergeability project.
Read this before you read a number
MergeBench publishes no pairwise merge. Every merge score in their release
(arXiv:2505.10833, Tables 8-17, and the two eval dumps in their
GitHub repo) is for a merge of all five domain… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/mergebench-property-sweep.crossarch-1b-diagnostics
Cross-architecture mergeability diagnostics for five ~1B monolingual LMs
A third model family for the mergeability project, alongside Goldfish and Beetle/MergeBench.
Read this first: what "merging" means for these five models
Five independently trained ~1B monolingual models were requested: Pythia-1.4B (EN),
Zh-Pythia-1.4B (ZH), Tucano-1b1 (PT), Bielik-1.5B-v3 (PL), Minerva-1B (IT). They differ in
architecture family, hidden dimension (1536 vs 2048), depth (16 /… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/crossarch-1b-diagnostics.crossarch-accuracy
crossarch-accuracy — downstream accuracy for the ~1B merge results
Companion to Mergeability-2/crossarch-1b-diagnostics,
which measured 106 Pythia-1.4B / Zh-Pythia-1.4B checkpoint-merge pairs, 1 native cross-model
merge and 20 transport merges entirely in nats/token. This repo adds the accuracy axis:
does any of it produce a usable model?
Headline question (Figure 1). The diagnostics found that checkpoint pairs within about half
a decade of training are linearly mode-connected… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/crossarch-accuracy.
