suchirsalhan/goldfish-crosslingual-cka
Cross-lingual representational alignment (CKA) of the Goldfish bilingual merge/train models Reproduction of the §3 Representational Alignment protocol of the Goldfish cross-lingual paper, applied to the 111 bilingual merge/train checkpoints under suchirsalhan/* created 2025-12. Generated 2026-08-26 19:17 UTC. Code: scripts/ in this repo. TL;DR — read this before the numbers The uploaded checkpoints do not contain the merges. All 111 repos ship a LoRA adapter over… See the full description on the dataset page: https://huggingface.co/datasets/suchirsalhan/goldfish-crosslingual-cka.
Cross-lingual representational alignment (CKA) of the Goldfish bilingual merge/train models
Reproduction of the §3 Representational Alignment protocol of the Goldfish cross-lingual paper, applied to the 111 bilingual merge/train checkpoints under suchirsalhan/* created 2025-12. Generated 2026-08-26 19:17 UTC. Code: `scripts/` in this repo.
TL;DR — read this before the numbers
The uploaded checkpoints do not contain the merges. All 111 repos ship a LoRA adapter over goldfish-models/eng_latn_1000mb (adapter_config.json → peft_type: LORA, r: 16, modules_to_save: null), plus a per-pair bilingual SentencePiece tokenizer. Two things are wrong with that as a bilingual model:
- 75/111 adapters are exact identities. Every
-mergedand everyTOKENWISE_MERGEadapter haslora_B == 0— its initialisation value — soB @ A == 0and the adapter changes nothing. The-trainedadapters are non-zero but tiny (max |ΔW| element ≤ 1.5e-2). The merge-weight α and the vocabulary-overlap top-K therefore have literally zero effect on the weights of the uploaded artefacts. - The bilingual embedding matrix was never uploaded. A tokenwise merge is an embedding-table operation (English backbone + tokenwise-averaged eng/L2 embeddings over a bilingual vocabulary), but
modules_to_saveis null, so no embedding ships. The shipped bilingual tokenizer's ids do not correspond to the English base's embedding rows: only 7–13 of 50 000 vocabulary positions agree. Loading a repo the obvious way (AutoModelForCausalLM.from_pretrained(base)+PeftModel.from_pretrained(repo)+ the repo's tokenizer) yields a model at ≈9.2–10.4 nats/token — chance is ln(51200)=10.84 — i.e. not a language model at all. So the analysis is reported in two configurations: - `as-uploaded` — exactly what
from_pretrainedgives you. Included as the audit trail; its CKA is the CKA of a broken model and should not be interpreted as an alignment measurement. - `reconstructed` — the described method re-implemented here: English Goldfish backbone (+ the uploaded LoRA), and the embedding table rebuilt as a tokenwise merge of
goldfish-models/eng_latn_1000mband the matching L2 Goldfish (ell_grek/spa_latn/nld_latn/pol_latn_1000mb) over the shipped bilingual vocabulary. This restores a working model (7.4–7.9 nats/token) and makes α and top-K real variables again. It is a reconstruction, not the authors' checkpoints — the exact recipe (mixing convention, unseen-token handling) is inferred, not recovered.
0. Headline correlations (last layer, SGPT pooling)
Read the `reconstructed` column. The as-uploaded column is what the published artefacts actually give you, and for the α and top-K sweeps it is a measurement of a constant.
1. Method (as reproduced)
- Linear CKA = HSIC(X,Y)/√(HSIC(X,X)·HSIC(Y,Y)), computed in the feature-space form (algebraically identical to the Gram form, O(n·d²)). Invariant to orthogonal transformation and isotropic rescaling.
- X, Y = the same bilingual model's sentence representations of the English side and the L2 side of parallel FLORES-200 devtest text.
- Matched = parallel order. Shuffled control = side A held in order, side B reordered by a single random permutation per language pair, drawn once (seed 42+1000) and applied identically at every layer and for every model, so every B sentence appears exactly once and the control preserves architecture and corpus statistics. The permutations used are stored in
shuffle_perms.json. (A fresh permutation per layer would be a strictly weaker control; it was not used.) - Every layer including the embedding layer: 13 layers (0 = embeddings, 1–12 = transformer blocks).
- Pooling: (i) attention-masked mean; (ii) SGPT position-weighted (Muennighoff et al. 2022) — the paper's strongest for causal LMs and the one quoted below; (iii) token-level word-aligned via SimAlign (
bert/bpe, itermax), averaging subword states within each aligned word span. - Data: FLORES-200 devtest (
openlanguagedata/flores_plus), 1012 sentences per language; n = 500 sampled per pair (seed 42) for (i)/(ii), n = 250 for (iii). Tatoeba/OPUS/BouQUET not added — FLORES alone, as in the paper's headline numbers. - Model quality = mean per-token NLL over the same sentences, both sides, teacher-forced.
- For pooling (iii) the CKA rows are aligned words, not sentences (≈5 400–6 100 word pairs per language pair), so the shuffled control there is a single permutation over those word rows — again drawn once per language pair and applied identically at every layer.
2. Model inventory
111 models, created 2025-12 under suchirsalhan. Parsed grid:
No cells are missing. The 4 × (2 + 8 + 16) grid is full; the 7 extra repos are redundant re-uploads, not new cells.
Adapter audit (adapter_audit.csv)
delta = B@A, the actual weight change the adapter applies. Every merged arm is an exact no-op.
Tokenizer provenance
The `merged` and `trained` arms never share a tokenizer file. In all four pairs and both families the two arms ship different tokenizer.jsons. Any merged-vs-trained difference measured on the uploaded artefacts is therefore confounded with tokenization: since the merged weights are the untouched base model and the trained weights differ from it by ≤1.5e-2, the tokenizer is the only substantial difference between the arms. The reconstructed configuration removes this confound by giving both arms the same bilingual tokenizer, which is why its merged-vs-trained Δ collapses to ≈0 while the as-uploaded Δ does not.
3. Results
3.1 reconstructed — last-layer summary by language pair (sgpt pooling)
3.2 as-uploaded — last-layer summary by language pair (sgpt pooling)
3.3 Stage 1 — merged vs trained (the paired baseline contrast)
Paired over (pair × top-K × α) cells:
With tokenization controlled, merging preserves cross-lingual alignment exactly as well as joint training. In reconstructed — where both arms use the same bilingual tokenizer and differ only by the uploaded LoRA — the mean trained−merged difference in last-layer matched CKA is +0.00019 (Wilcoxon p = 0.846, paired t p = 0.337, n = 32 paired cells): indistinguishable from zero.
In as-uploaded the same contrast is +0.02275 (Wilcoxon p = 7.1e-07, n = 36) — apparently significant, but this is a tokenizer effect, not a training effect: the two arms never ship the same tokenizer.json (see Tokenizer provenance above), and the merged arm's weights are the untouched base model. Note also that the word-aligned pooling, which reduces both arms to the same aligned word spans, gives −0.00000 for the same comparison — consistent with the difference being tokenization.
Either way this is a null by construction, not a scientific finding about merging: the audit above shows the merged arm is the base model and the trained arm differs from it by ≤1.5e-2 in any weight. A real merged-vs-trained comparison needs the merged embedding tables, which were never uploaded.
3.4 Stage 2 — merge weight α
`as-uploaded` is degenerate. Distinct last-layer matched-CKA values actually observed along the sweep:
Bit-identical CKA means bit-identical models. Seven of the eight en_el-alpha* repos ship tokenizer 377d2055dc and return matched CKA 0.10644523 to eight decimal places; only alpha90, which ships a different tokenizer file (0b6e8ebd2f), differs. The variation in this sweep is tokenizer variation, not merge-weight variation.
Sweep range in last-layer matched CKA:
Reconstructed sweep (last-layer matched CKA, sgpt):
Best setting per pair:
α is not monotone — there is an interior optimum. Alignment climbs steeply from α=10 to α≈50–75, then flattens or falls slightly toward α=100 (per-pair argmax: enel=75, enes=50, ennl=66, enpl=66). Spearman is a monotone statistic and understates this; read the curve in figures/alignment_vs_alpha_reconstructed.png. Quote the mean of the four within-pair ρ from the table above, not the pooled p.
3.5 Stage 3 — vocabulary overlap top-K
`as-uploaded` is degenerate. Distinct last-layer matched-CKA values actually observed along the sweep:
Bit-identical CKA means bit-identical models. The merged arm returns a single value at every K — all eight repos are the same identity-adapter model with the same tokenizer. The trained arm varies only because its LoRA deltas are non-zero (by ≤1.5e-2 in any weight) and because some cells ship a different tokenizer file.
Sweep range in last-layer matched CKA:
Reconstructed sweep (last-layer matched CKA, sgpt):
Best setting per pair:
top-K is monotone increasing in every pair and in both arms (within-pair ρ ≈ +0.99; argmax = 100 in 8/8 cells): the more of the L2 vocabulary is merged in, the better the English and L2 representations correspond. No interaction with merged-vs-trained — the two arms' curves sit within ~0.003 CKA of one another at every K.
3.6 Language pair vs URIEL typological distance
URIEL syntax_knn cosine distance from English: nld 0.0757, spa 0.1784, pol 0.2136, ell 0.2169.
n = 4. The smallest p a Spearman ρ can attain at n = 4 is 0.0833, so nothing here can be significant; the rank ordering is the whole content.
3.7 Model quality (mean sentence NLL) vs alignment — the Platonic-Representation test
4. Non-independence — how to read the p-values
Every p above is computed as if the rows were independent draws. They are not.
- All 111 uploaded models share one base model (
goldfish-models/eng_latn_1000mb); 75 of them are that base model. - Within a language pair, the α sweep and the top-K sweep are repeated measures on one base model + one tokenizer; they are 8 perturbations of a single object, not 8 samples.
- Within a design cell,
mergedandtrainedshare the tokenizer and the backbone; the merged/trained comparison is therefore reported paired (Wilcoxon signed-rank and paired t over cells), which is the right test, but the cells themselves are still nested in 4 language pairs. - The honest effective n for anything that varies across language pairs is 4, not 8/32/104. The
mean of within-pair rhorows incorrelations.csvare the repeated-measures-safe summary: compute ρ inside each pair, then average the four ρ's. Quote those, not the pooled p.
5. What ran and what did not
6. Files
7. What would make this a real result
The design here is good and the reconstruction shows the questions are answerable. What is missing is one file per repo:
- Upload the merged embedding table. Re-save with
modules_to_save=["wte"](or push the full merged model rather than a PEFT adapter). The tokenwise merge lives entirely in the embeddings; without them the repo carries none of the method. - Check `lora_B` before pushing.
lora_Bis zero-initialised; 75 of these repos were pushed before any gradient reached it. An assert thatmax|B| > 0at save time would have caught all 75. - Keep the tokenizer fixed within a design cell. The
mergedandtrainedarms currently differ in tokenizer, which confounds the headline comparison the design was built to make. With those, re-runningscripts/run_cka.pyagainst the real checkpoints reproduces every table here in about 20 minutes on one GPU.
