CoolFace
Datasetpublic

suchirsalhan/goldfish-crosslingual-cka

Cross-lingual representational alignment (CKA) of the Goldfish bilingual merge/train models Reproduction of the §3 Representational Alignment protocol of the Goldfish cross-lingual paper, applied to the 111 bilingual merge/train checkpoints under suchirsalhan/* created 2025-12. Generated 2026-08-26 19:17 UTC. Code: scripts/ in this repo. TL;DR — read this before the numbers The uploaded checkpoints do not contain the merges. All 111 repos ship a LoRA adapter over… See the full description on the dataset page: https://huggingface.co/datasets/suchirsalhan/goldfish-crosslingual-cka.

sourceHugging Facecc-by-4.0updated 1mo agoView on Hugging Face
0likes52downloads
Dataset Card

Cross-lingual representational alignment (CKA) of the Goldfish bilingual merge/train models

Reproduction of the §3 Representational Alignment protocol of the Goldfish cross-lingual paper, applied to the 111 bilingual merge/train checkpoints under suchirsalhan/* created 2025-12. Generated 2026-08-26 19:17 UTC. Code: `scripts/` in this repo.


TL;DR — read this before the numbers

The uploaded checkpoints do not contain the merges. All 111 repos ship a LoRA adapter over goldfish-models/eng_latn_1000mb (adapter_config.json → peft_type: LORA, r: 16, modules_to_save: null), plus a per-pair bilingual SentencePiece tokenizer. Two things are wrong with that as a bilingual model:

  1. 1.75/111 adapters are exact identities. Every -merged and every TOKENWISE_MERGE adapter has lora_B == 0 — its initialisation value — so B @ A == 0 and the adapter changes nothing. The -trained adapters are non-zero but tiny (max |ΔW| element ≤ 1.5e-2). The merge-weight α and the vocabulary-overlap top-K therefore have literally zero effect on the weights of the uploaded artefacts.
  2. 2.The bilingual embedding matrix was never uploaded. A tokenwise merge is an embedding-table operation (English backbone + tokenwise-averaged eng/L2 embeddings over a bilingual vocabulary), but modules_to_save is null, so no embedding ships. The shipped bilingual tokenizer's ids do not correspond to the English base's embedding rows: only 7–13 of 50 000 vocabulary positions agree. Loading a repo the obvious way (AutoModelForCausalLM.from_pretrained(base) + PeftModel.from_pretrained(repo) + the repo's tokenizer) yields a model at ≈9.2–10.4 nats/token — chance is ln(51200)=10.84 — i.e. not a language model at all. So the analysis is reported in two configurations:
  3. 3.`as-uploaded` — exactly what from_pretrained gives you. Included as the audit trail; its CKA is the CKA of a broken model and should not be interpreted as an alignment measurement.
  4. 4.`reconstructed` — the described method re-implemented here: English Goldfish backbone (+ the uploaded LoRA), and the embedding table rebuilt as a tokenwise merge of goldfish-models/eng_latn_1000mb and the matching L2 Goldfish (ell_grek/spa_latn/nld_latn/pol_latn_1000mb) over the shipped bilingual vocabulary. This restores a working model (7.4–7.9 nats/token) and makes α and top-K real variables again. It is a reconstruction, not the authors' checkpoints — the exact recipe (mixing convention, unseen-token handling) is inferred, not recovered.

0. Headline correlations (last layer, SGPT pooling)

design variablereconstructedas-uploadedreading
merge weight αρ = +0.744 (mean of 4 within-pair ρ)ρ = +0.206 (mean of 4 within-pair ρ)alignment rises with the English share of the merged embedding, then plateaus — interior optimum α≈50–75
vocabulary overlap top-Kρ = +0.994 (mean of 4 within-pair ρ)ρ = -0.054 (mean of 4 within-pair ρ)alignment increases monotonically with shared vocabulary; no interaction with merged-vs-trained
model quality (mean NLL)ρ = -0.743 (p = 2.3e-17, n = 92)ρ = -0.054 (p = 0.58, n = 111)better language models align better — the Platonic-Representation prediction, and stronger than the paper's ρ=−0.43
URIEL syntactic distanceρ = -0.800 (p = 0.2, n = 4)ρ = -0.800 (p = 0.2, n = 4)typologically closer to English ⇒ better aligned; same sign as the paper's ρ = −0.64. n = 4, so p can never fall below 0.0833
merged vs trained (paired)Δ = +0.00019 (Wilcoxon p = 0.846, n = 32 paired cells)Δ = +0.02275 (Wilcoxon p = 0.000, n = 36 paired cells)null once tokenization is controlled. Merging preserves alignment exactly as well as joint training — but only because the two arms are nearly the same weights; the as-uploaded Δ is a tokenizer effect (the arms never share a tokenizer.json)

Read the `reconstructed` column. The as-uploaded column is what the published artefacts actually give you, and for the α and top-K sweeps it is a measurement of a constant.


1. Method (as reproduced)

  • —Linear CKA = HSIC(X,Y)/√(HSIC(X,X)·HSIC(Y,Y)), computed in the feature-space form (algebraically identical to the Gram form, O(n·d²)). Invariant to orthogonal transformation and isotropic rescaling.
  • —X, Y = the same bilingual model's sentence representations of the English side and the L2 side of parallel FLORES-200 devtest text.
  • —Matched = parallel order. Shuffled control = side A held in order, side B reordered by a single random permutation per language pair, drawn once (seed 42+1000) and applied identically at every layer and for every model, so every B sentence appears exactly once and the control preserves architecture and corpus statistics. The permutations used are stored in shuffle_perms.json. (A fresh permutation per layer would be a strictly weaker control; it was not used.)
  • —Every layer including the embedding layer: 13 layers (0 = embeddings, 1–12 = transformer blocks).
  • —Pooling: (i) attention-masked mean; (ii) SGPT position-weighted (Muennighoff et al. 2022) — the paper's strongest for causal LMs and the one quoted below; (iii) token-level word-aligned via SimAlign (bert/bpe, itermax), averaging subword states within each aligned word span.
  • —Data: FLORES-200 devtest (openlanguagedata/flores_plus), 1012 sentences per language; n = 500 sampled per pair (seed 42) for (i)/(ii), n = 250 for (iii). Tatoeba/OPUS/BouQUET not added — FLORES alone, as in the paper's headline numbers.
  • —Model quality = mean per-token NLL over the same sentences, both sides, teacher-forced.
  • —For pooling (iii) the CKA rows are aligned words, not sentences (≈5 400–6 100 word pairs per language pair), so the shuffled control there is a single permutation over those word rows — again drawn once per language pair and applied identically at every layer.

2. Model inventory

111 models, created 2025-12 under suchirsalhan. Parsed grid:

familyngrid
baseline84 pairs × {merged, trained} — complete
alpha (-alpha{K}-TOKENWISE_MERGE)324 pairs × α ∈ {10,25,33,50,66,75,90,100} — complete
top-K644 pairs × K ∈ {10,20,25,33,50,66,75,100} × {merged, trained} — complete
alternate naming7en_es-top{10,20,66,75}_top…, en_nl-top{25,33,50}_top… — duplicate namings of top-K cells that already exist under <pair>_merge_top{K}-merged-top{K}; excluded from the design tables, still audited

No cells are missing. The 4 × (2 + 8 + 16) grid is full; the 7 extra repos are redundant re-uploads, not new cells.

Adapter audit (adapter_audit.csv)

familyarmnexact_identitiesmax_abs_delta_Wmedian_frobenius
alphamerged32320.0000000.000000
baselinemerged440.0000000.000000
baselinetrained400.0029940.923666
topkmerged32320.0000000.000000
topktrained3200.0154403.618423

delta = B@A, the actual weight change the adapter applies. Every merged arm is an exact no-op.

Tokenizer provenance

pairfamilyarmtokenizer md5 (first 10)
en_elalphamerged0b6e8ebd2f, 377d2055dc
en_elbaselinemerged0b6e8ebd2f
en_elbaselinetrained5d785f18bc
en_eltopkmerged377d2055dc
en_eltopktrainedbb4c05fbf8
en_esalphamerged248c7914cb, cad7bb89e0
en_esbaselinemergedcad7bb89e0
en_esbaselinetrainedbaba93735d
en_estopkmerged248c7914cb
en_estopktrained804273be16
en_nlalphamerged8986b4dd2a, d2a8e7b606
en_nlbaselinemergedd2a8e7b606
en_nlbaselinetrained3eb4a951cb
en_nltopkmerged8986b4dd2a
en_nltopktrained824f7f899a
en_plalphamergedac737037a8, c4295ff30e
en_plbaselinemergedac737037a8
en_plbaselinetrained7945a3a515
en_pltopkmergedc4295ff30e
en_pltopktrained6e9eb8ff4d

The `merged` and `trained` arms never share a tokenizer file. In all four pairs and both families the two arms ship different tokenizer.jsons. Any merged-vs-trained difference measured on the uploaded artefacts is therefore confounded with tokenization: since the merged weights are the untouched base model and the trained weights differ from it by ≤1.5e-2, the tokenizer is the only substantial difference between the arms. The reconstructed configuration removes this confound by giving both arms the same bilingual tokenizer, which is why its merged-vs-trained Δ collapses to ≈0 while the as-uploaded Δ does not.


3. Results

3.1 reconstructed — last-layer summary by language pair (sgpt pooling)

pairn_modelsmatchedshuffleddeltanllURIEL_syntax_knn
en–el (Greek, Greek script)230.16600.09360.07248.12860.2169
en–es (Spanish)230.25630.10230.15407.48940.1784
en–nl (Dutch)230.24700.09650.15057.47960.0757
en–pl (Polish)230.21850.10750.11117.85290.2136

3.2 as-uploaded — last-layer summary by language pair (sgpt pooling)

pairn_modelsmatchedshuffleddeltanllURIEL_syntax_knn
en–el (Greek, Greek script)260.10830.08780.02059.81340.2169
en–es (Spanish)300.13680.08850.04839.68870.1784
en–nl (Dutch)290.15230.09300.05939.51970.0757
en–pl (Polish)260.10630.07770.02869.71880.2136

3.3 Stage 1 — merged vs trained (the paired baseline contrast)

configpairarmcka_matchedcka_shuffleddeltanll_mean
as-uploadeden_elmerged0.12370.08800.03589.8013
as-uploadeden_eltrained0.12390.08870.03529.7972
as-uploadeden_esmerged0.14440.08790.05659.2884
as-uploadeden_estrained0.14640.08910.05739.3032
as-uploadeden_nlmerged0.11610.08710.02909.1630
as-uploadeden_nltrained0.11690.08760.02939.1845
as-uploadeden_plmerged0.11490.08160.03339.6213
as-uploadeden_pltrained0.11470.08150.03329.6226
reconstructeden_elmerged0.18900.10400.08507.9466
reconstructeden_eltrained0.18940.10450.08497.9430
reconstructeden_esmerged0.28930.10590.18347.4213
reconstructeden_estrained0.28830.10690.18147.4276
reconstructeden_nlmerged0.28360.10240.18127.3902
reconstructeden_nltrained0.28210.10090.18127.3868
reconstructeden_plmerged0.24890.11520.13377.7426
reconstructeden_pltrained0.24900.11570.13337.7483

Paired over (pair × top-K × α) cells:

configpoolingn_pairsmean_mergedmean_trainedmean_diff_trained_minus_mergedsd_diffwilcoxon_Wwilcoxon_ppaired_tt_p
reconstructedmean320.222670.222930.000260.00087211.000000.330921.696680.09978
reconstructedsgpt320.214890.215080.000190.00111253.000000.846460.974570.33732
reconstructedword_aligned40.177030.177200.000170.001014.000000.875000.337010.75831
as-uploadedmean360.122020.141910.019890.0223050.000000.000005.350930.00001
as-uploadedsgpt360.119950.142700.022750.0279248.000000.000004.889600.00002
as-uploadedword_aligned40.138250.13825-0.000000.000665.000001.00000-0.001710.99875

With tokenization controlled, merging preserves cross-lingual alignment exactly as well as joint training. In reconstructed — where both arms use the same bilingual tokenizer and differ only by the uploaded LoRA — the mean trained−merged difference in last-layer matched CKA is +0.00019 (Wilcoxon p = 0.846, paired t p = 0.337, n = 32 paired cells): indistinguishable from zero.

In as-uploaded the same contrast is +0.02275 (Wilcoxon p = 7.1e-07, n = 36) — apparently significant, but this is a tokenizer effect, not a training effect: the two arms never ship the same tokenizer.json (see Tokenizer provenance above), and the merged arm's weights are the untouched base model. Note also that the word-aligned pooling, which reduces both arms to the same aligned word spans, gives −0.00000 for the same comparison — consistent with the difference being tokenization.

Either way this is a null by construction, not a scientific finding about merging: the audit above shows the merged arm is the base model and the trained arm differs from it by ≤1.5e-2 in any weight. A real merged-vs-trained comparison needs the merged embedding tables, which were never uploaded.

3.4 Stage 2 — merge weight α

configsubsetspearman_rhop_valuennote
reconstructedall pairs pooled0.41950.016832sweep points share a base model -> repeated measures; p assumes indepe
reconstructedwithin en_el0.92860.00098within-pair sweep; n = sweep points
reconstructedwithin en_es0.52380.18278within-pair sweep; n = sweep points
reconstructedwithin en_nl0.76190.02808within-pair sweep; n = sweep points
reconstructedwithin en_pl0.76190.02808within-pair sweep; n = sweep points
reconstructedmean of within-pair rho0.74404average of the per-pair rhos (the repeated-measures-safe summary)
as-uploadedall pairs pooled0.06030.742932sweep points share a base model -> repeated measures; p assumes indepe
as-uploadedwithin en_el0.41240.31008within-pair sweep; n = sweep points
as-uploadedwithin en_es0.41240.31008within-pair sweep; n = sweep points
as-uploadedwithin en_nl-0.41240.31008within-pair sweep; n = sweep points
as-uploadedwithin en_pl0.41240.31008within-pair sweep; n = sweep points
as-uploadedmean of within-pair rho0.20624average of the per-pair rhos (the repeated-measures-safe summary)

`as-uploaded` is degenerate. Distinct last-layer matched-CKA values actually observed along the sweep:

pairarmn_pointsn_distinct
en_elmerged82
en_esmerged82
en_nlmerged82
en_plmerged82

Bit-identical CKA means bit-identical models. Seven of the eight en_el-alpha* repos ship tokenizer 377d2055dc and return matched CKA 0.10644523 to eight decimal places; only alpha90, which ships a different tokenizer file (0b6e8ebd2f), differs. The variation in this sweep is tokenizer variation, not merge-weight variation.

Sweep range in last-layer matched CKA:

configpairarmn_pointscka_mincka_maxcka_range
reconstructeden_elmerged80.1473270.1907150.043388
reconstructeden_esmerged80.2436470.2893140.045667
reconstructeden_nlmerged80.2256770.2910760.065399
reconstructeden_plmerged80.1850760.2622110.077135
as-uploadeden_elmerged80.1064450.1237270.017282
as-uploadeden_esmerged80.1216220.1443930.022771
as-uploadeden_nlmerged80.1161020.1462940.030191
as-uploadeden_plmerged80.1030310.1148750.011844

Reconstructed sweep (last-layer matched CKA, sgpt):

alphaen_elen_esen_nlen_pl
10.00000.14730.24360.22570.1851
25.00000.16980.25900.25170.1997
33.00000.18560.26610.25830.2128
50.00000.18920.28880.28290.2489
66.00000.19070.28520.29110.2622
75.00000.19070.27910.28890.2600
90.00000.19070.27640.28620.2553
100.00000.19070.27320.28410.2526

Best setting per pair:

pairarmargmaxcka_at_argmaxcka_at_min_settingcka_at_max_setting
en_elmerged75.00000.19070.14730.1907
en_esmerged50.00000.28930.24360.2732
en_nlmerged66.00000.29110.22570.2841
en_plmerged66.00000.26220.18510.2526

α is not monotone — there is an interior optimum. Alignment climbs steeply from α=10 to α≈50–75, then flattens or falls slightly toward α=100 (per-pair argmax: enel=75, enes=50, ennl=66, enpl=66). Spearman is a monotone statistic and understates this; read the curve in figures/alignment_vs_alpha_reconstructed.png. Quote the mean of the four within-pair ρ from the table above, not the pooled p.

3.5 Stage 3 — vocabulary overlap top-K

configsubsetspearman_rhop_valuennote
reconstructedall pairs pooled0.54860.000064sweep points share a base model -> repeated measures; p assumes indepe
reconstructedwithin en_el0.99410.000016within-pair sweep; n = sweep points
reconstructedwithin en_es0.99410.000016within-pair sweep; n = sweep points
reconstructedwithin en_nl0.99410.000016within-pair sweep; n = sweep points
reconstructedwithin en_pl0.99410.000016within-pair sweep; n = sweep points
reconstructedmean of within-pair rho0.99414average of the per-pair rhos (the repeated-measures-safe summary)
as-uploadedall pairs pooled-0.06040.635464sweep points share a base model -> repeated measures; p assumes indepe
as-uploadedwithin en_el0.62570.009516within-pair sweep; n = sweep points
as-uploadedwithin en_es-0.25280.344816within-pair sweep; n = sweep points
as-uploadedwithin en_nl-0.19590.467116within-pair sweep; n = sweep points
as-uploadedwithin en_pl-0.39190.133316within-pair sweep; n = sweep points
as-uploadedmean of within-pair rho-0.05374average of the per-pair rhos (the repeated-measures-safe summary)

`as-uploaded` is degenerate. Distinct last-layer matched-CKA values actually observed along the sweep:

pairarmn_pointsn_distinct
en_elmerged81
en_eltrained88
en_esmerged81
en_estrained88
en_nlmerged81
en_nltrained88
en_plmerged81
en_pltrained88

Bit-identical CKA means bit-identical models. The merged arm returns a single value at every K — all eight repos are the same identity-adapter model with the same tokenizer. The trained arm varies only because its LoRA deltas are non-zero (by ≤1.5e-2 in any weight) and because some cells ship a different tokenizer file.

Sweep range in last-layer matched CKA:

configpairarmn_pointscka_mincka_maxcka_range
reconstructeden_elmerged80.1357860.1890020.053216
reconstructeden_eltrained80.1365020.1893510.052849
reconstructeden_esmerged80.2054700.2893140.083844
reconstructeden_estrained80.2081400.2882740.080134
reconstructeden_nlmerged80.1984680.2835980.085131
reconstructeden_nltrained80.2006210.2821180.081497
reconstructeden_plmerged80.1876460.2488830.061236
reconstructeden_pltrained80.1870900.2489830.061893
as-uploadeden_elmerged80.1064450.1064450.000000
as-uploadeden_eltrained80.1026100.1080300.005419
as-uploadeden_esmerged80.1216220.1216220.000000
as-uploadeden_estrained80.1459460.2161200.070174
as-uploadeden_nlmerged80.1462940.1462940.000000
as-uploadeden_nltrained80.1703120.2109250.040612
as-uploadeden_plmerged80.1030310.1030310.000000
as-uploadeden_pltrained80.1016510.1167910.015140

Reconstructed sweep (last-layer matched CKA, sgpt):

topken_elen_esen_nlen_pl
10.00000.13610.20680.19950.1874
20.00000.13720.22300.21060.1908
25.00000.14240.23100.21260.1956
33.00000.14990.24110.21950.2068
50.00000.16140.25760.24300.2119
66.00000.17800.27740.26070.2258
75.00000.18190.28060.26890.2323
100.00000.18920.28880.28290.2489

Best setting per pair:

pairarmargmaxcka_at_argmaxcka_at_min_settingcka_at_max_setting
en_elmerged100.00000.18900.13580.1890
en_eltrained100.00000.18940.13650.1894
en_esmerged100.00000.28930.20550.2893
en_estrained100.00000.28830.20810.2883
en_nlmerged100.00000.28360.19850.2836
en_nltrained100.00000.28210.20060.2821
en_plmerged100.00000.24890.18760.2489
en_pltrained100.00000.24900.18710.2490

top-K is monotone increasing in every pair and in both arms (within-pair ρ ≈ +0.99; argmax = 100 in 8/8 cells): the more of the L2 vocabulary is merged in, the better the English and L2 representations correspond. No interaction with merged-vs-trained — the two arms' curves sit within ~0.003 CKA of one another at every K.

3.6 Language pair vs URIEL typological distance

design_variableconfigpoolingspearman_rhop_valuennote
URIEL syntax_knn vs ckareconstructedmean-0.80000.20004n = 4 language pairs; minimum attainable Spearman p at n=4 i
URIEL syntax_knn vs deltareconstructedmean-1.00000.00004n = 4 language pairs; minimum attainable Spearman p at n=4 i
URIEL syntax_knn vs ckareconstructedsgpt-0.80000.20004n = 4 language pairs; minimum attainable Spearman p at n=4 i
URIEL syntax_knn vs deltareconstructedsgpt-0.80000.20004n = 4 language pairs; minimum attainable Spearman p at n=4 i
URIEL syntax_knn vs ckareconstructedword_aligned0.20000.80004n = 4 language pairs; minimum attainable Spearman p at n=4 i
URIEL syntax_knn vs deltareconstructedword_aligned0.20000.80004n = 4 language pairs; minimum attainable Spearman p at n=4 i
URIEL syntax_knn vs ckaas-uploadedmean-1.00000.00004n = 4 language pairs; minimum attainable Spearman p at n=4 i
URIEL syntax_knn vs deltaas-uploadedmean-1.00000.00004n = 4 language pairs; minimum attainable Spearman p at n=4 i
URIEL syntax_knn vs ckaas-uploadedsgpt-0.80000.20004n = 4 language pairs; minimum attainable Spearman p at n=4 i
URIEL syntax_knn vs deltaas-uploadedsgpt-1.00000.00004n = 4 language pairs; minimum attainable Spearman p at n=4 i
URIEL syntax_knn vs ckaas-uploadedword_aligned0.40000.60004n = 4 language pairs; minimum attainable Spearman p at n=4 i
URIEL syntax_knn vs deltaas-uploadedword_aligned0.40000.60004n = 4 language pairs; minimum attainable Spearman p at n=4 i

URIEL syntax_knn cosine distance from English: nld 0.0757, spa 0.1784, pol 0.2136, ell 0.2169.

n = 4. The smallest p a Spearman ρ can attain at n = 4 is 0.0833, so nothing here can be significant; the rank ordering is the whole content.

3.7 Model quality (mean sentence NLL) vs alignment — the Platonic-Representation test

configpoolingsubsetspearman_rhop_valuen
reconstructedmeanall models pooled-0.77150.000092
reconstructedmeanwithin en_el-0.81820.000023
reconstructedmeanwithin en_es-0.68970.000323
reconstructedmeanwithin en_nl-0.87650.000023
reconstructedmeanwithin en_pl-0.95850.000023
reconstructedsgptall models pooled-0.74290.000092
reconstructedsgptwithin en_el-0.81520.000023
reconstructedsgptwithin en_es-0.71840.000123
reconstructedsgptwithin en_nl-0.80830.000023
reconstructedsgptwithin en_pl-0.90610.000023
reconstructedword_alignedall models pooled0.02380.95548
reconstructedword_alignedwithin en_el2
reconstructedword_alignedwithin en_es2
reconstructedword_alignedwithin en_nl2
reconstructedword_alignedwithin en_pl2
as-uploadedmeanall models pooled-0.13010.1736111
as-uploadedmeanwithin en_el-0.00590.977126
as-uploadedmeanwithin en_es0.56220.001230
as-uploadedmeanwithin en_nl0.66240.000129
as-uploadedmeanwithin en_pl0.50510.008526
as-uploadedsgptall models pooled-0.05380.5752111
as-uploadedsgptwithin en_el-0.25380.210926
as-uploadedsgptwithin en_es0.59120.000630
as-uploadedsgptwithin en_nl0.66580.000129
as-uploadedsgptwithin en_pl0.18950.353826
as-uploadedword_alignedall models pooled0.40480.31998
as-uploadedword_alignedwithin en_el2
as-uploadedword_alignedwithin en_es2
as-uploadedword_alignedwithin en_nl2
as-uploadedword_alignedwithin en_pl2

4. Non-independence — how to read the p-values

Every p above is computed as if the rows were independent draws. They are not.

  • —All 111 uploaded models share one base model (goldfish-models/eng_latn_1000mb); 75 of them are that base model.
  • —Within a language pair, the α sweep and the top-K sweep are repeated measures on one base model + one tokenizer; they are 8 perturbations of a single object, not 8 samples.
  • —Within a design cell, merged and trained share the tokenizer and the backbone; the merged/trained comparison is therefore reported paired (Wilcoxon signed-rank and paired t over cells), which is the right test, but the cells themselves are still nested in 4 language pairs.
  • —The honest effective n for anything that varies across language pairs is 4, not 8/32/104. The mean of within-pair rho rows in correlations.csv are the repeated-measures-safe summary: compute ρ inside each pair, then average the four ρ's. Quote those, not the pooled p.

5. What ran and what did not

itemstatus
Model enumeration (111 repos, HF API)ran
Adapter integrity audit (all 111)ran → adapter_audit.csv
FLORES-200 devtest, 4 pairs × 1012 sentsran
Linear CKA, matched + single-permutation shuffled control, all 13 layersran
Pooling (i) meanran
Pooling (ii) SGPT position-weightedran
Pooling (iii) word-aligned (SimAlign)ran on the 16 stage-1 models (8 as-uploaded + 8 reconstructed), n=250 — cka_layerwise_wordaligned.csv
Stage 1 merged vs trained, 4 pairsran
Stage 2 α sweepran (as-uploaded: null by construction; reconstructed: real)
Stage 3 top-K sweepran (same caveat)
URIEL typological distance (lang2vec syntax_knn, cosine)ran
Mean sentence NLL per modelran
Tatoeba / OPUS / BouQUET parallel datanot run — FLORES alone, which is what the paper's headline numbers use
B-GPT jointly-trained upper boundnot run — out of scope for this brief

6. Files

filecontents
cka_layerwise.csvmodel × layer × pooling × {matched, shuffled, Δ}, both configs
cka_lastlayer.csvlast-layer slice of the above
lastlayer_summary.csvlast-layer summary matrix per config × pair × pooling
correlations.csvone row per design variable × config × pooling × subset
corr_merged_vs_trained.csvthe paired merged/trained contrast
adapter_audit.csvper-repo LoRA weight-delta audit (the integrity finding)
model_inventory.jsonthe parsed design grid
uriel_distances.jsonURIEL/lang2vec distances from English
shuffle_perms.jsonthe exact shuffled-control permutations
tables.mdall tables in markdown
figures/Figure-2 equivalent, per-pair heatmaps, alignment-vs-α, alignment-vs-top-K, CKA-vs-NLL
sweep_flatness.csvmin/max/range of matched CKA along each sweep (the flatness evidence)
sweep_optimum.csvargmax setting per sweep × pair × arm
lastlayer_matrix_<pair>.csvthe per-language-pair last-layer summary matrix
cka_layerwise_wordaligned.csvthe SimAlign word-aligned pooling run
scripts/run_cka.py (pipeline), analyse.py, adapter_audit.py, prep_data.py, recon_probe.py, write_report.py, publish.py

7. What would make this a real result

The design here is good and the reconstruction shows the questions are answerable. What is missing is one file per repo:

  1. 1.Upload the merged embedding table. Re-save with modules_to_save=["wte"] (or push the full merged model rather than a PEFT adapter). The tokenwise merge lives entirely in the embeddings; without them the repo carries none of the method.
  2. 2.Check `lora_B` before pushing. lora_B is zero-initialised; 75 of these repos were pushed before any gradient reached it. An assert that max|B| > 0 at save time would have caught all 75.
  3. 3.Keep the tokenizer fixed within a design cell. The merged and trained arms currently differ in tokenizer, which confounds the headline comparison the design was built to make. With those, re-running scripts/run_cka.py against the real checkpoints reproduces every table here in about 20 minutes on one GPU.