CoolFace
Datasetpublic

BrainAlign/cdl-devai-results-ds002236

ds002236 (Lytle et al. 2020) — brain × interpretability × localisation, per model per checkpoint Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children. Cohort: children 8.7–15.5 years; presentation: auditory and visual word presentation. Tasks Phon, Sem × sessions ses-9, ses-11, ses-11+ = 6 task × session cells, all of them scored here. Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds002236.

sourceHugging Facecc-by-4.0updated 20d agoView on Hugging Face
0likes375downloads
Dataset Card

ds002236 (Lytle et al. 2020) — brain × interpretability × localisation, per model per checkpoint

Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children. Cohort: children 8.7–15.5 years; presentation: auditory and visual word presentation. Tasks Phon, Sem × sessions ses-9, ses-11, ses-11+ = 6 task × session cells, all of them scored here.

Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4 checkpoints, pythia-6.9b-full has 1 checkpoint. pythia-6.9b-full's single checkpoint is step 0, i.e. the untrained initialisation; it is not the top of the scale ladder, it is the ladder's zero rung.

Same schema, same metric and same upstream analysis code as `BrainAlign/cdl-devai-results`, so rows are directly comparable across the three developmental datasets.

axiswhat it askssource tables
braindoes the model's representational geometry match the brain's?brain_alignment
interphow is the representation organised internally?interp_mechanistic, interp_layerwise
localisationare linguistic phenomena isolated into dedicated units?localisation_isolation, localisation_onset
behaviourdoes it get the minimal pairs right?behaviour

Start with `summary_by_checkpoint` (the default): one row per model × checkpoint with all axes side by side. 294 rows, 29 model families, 1764 alignment rows.


⚠️ READ THIS BEFORE USING THE ALIGNMENT NUMBERS

Nothing here shows that language models align with these brain data, and nothing here shows that they fail to. The reasons, measured rather than assumed — and note that the first one means a masked rebuild may change these numbers substantially:

1. These RDMs are WHOLE-BRAIN and UNMASKED, and that is a defect, not a design choice. Anatomical masking was not applied when these RDMs were built on the GPU cluster, so every number here is computed over the whole acquired volume rather than over language- or phonology-responsive cortex. On ds003604 the consequence is measured: the RDMs occupy ~4 effective dimensions of a 72-stimulus space and track whole-brain signal level. Masked versions are being rebuilt (ROI_SET=language, ROI_SET=phonology, ROI_SET=all) and will be published as separate -roi* datasets. Until then, treat every alignment number here as a whole-brain measurement and do not read it as a claim about the language network.

1b. The positive control on THIS dataset is one control, and its gate plumbing was faulty. The card previously said "the positive control failed", carried over from the ds003604 battery. That is an over-read here twice over. First, control/control_summary.csv in the predecessor repo contains exactly one control — text_edit_distance — against the eight-control battery (duration, intensity, word length, syllables, phonemes, log frequency, run identity, presentation order) used on ds003604; the stimulus-characteristics tables needed for the rest were not on the machine that ran it. A single non-significant edit-distance control does not establish that no stimulus property is recoverable. Second, the control labels feeding those gates were empty, so the gate outcomes are not interpretable at all and are being re-plumbed. No claim on this card depends on the positive control, and none should be read as supported by it.

2. The reference that matters is the untrained one, and it is now measured. An earlier version of this card said no random-initialisation baseline existed in this collection. That was wrong: the Pythia/PolyPythia step0 branches ARE the initialisation before any optimiser step, 15 independent random inits were swept across every cell, and they were sitting in the results unlabelled. untrained_reference and untrained_vs_trained_paired report them. Paired within each family, on this dataset:

61/84 family x cell pairs (73%) improve with training
mean rsa   untrained +0.00122   trained +0.02203
Wilcoxon signed-rank p = 4.4e-05
VERDICT: training IMPROVES alignment on this dataset

Read frac_of_ceiling alongside it: the absolute magnitudes are tiny either way, so this is a statement about sign and consistency, not about a model matching the brain.

2b. `parc_reference` is a between-model check, not a null, and must be read two-sided. The PARC families (parc-pythia, parc-mamba, parc-rwkv, seeds 0--2) are trained models --- 160M, OpenWebText, 4000 steps --- differing only in architecture and seed, so they are a matched-scale seed reference. Both this collection's upstream notes and the `cdl-devai-results` README call them "pure-noise runs"; that label is wrong and is corrected here. On this dataset 13 of 120 family x cell combinations exceed the PARC band by 2 SD and 13 fall below it. Where those two counts are comparable, the excursions are a variance artifact rather than evidence of alignment, and the cell contributes nothing either way.

parc_seed_summary reports that reference as a measurement in its own right: per architecture, the mean rsa over this dataset's cells with the SD and a t(2) 95% interval taken across the three seeds, plus the same in units of the noise ceiling. parc_by_seed_cell is the per (architecture x seed x cell) table it is built from. Three seeds is a very small sample for an interval; read the SD, not the CI width, as the error bar.

2c. The instrument was tested, and it works --- which is what makes the null mean something. A null is only interpretable with a detection floor attached, so we measured one (diagnostics/).

external group RDM, different cohort, same stimuli : rho = 0.592 (median over cells)
best language model anywhere in this grid            : rsa = 0.1032
model as a fraction of that external benchmark       : 15.1%
smallest detectable mixed probe (w*)                 : 0.02
additive per-stimulus control                        : rsa median 0.2248, max 0.3391, significant in 6/6 cells

Two consequences. First, the estimator is not deaf: an external group RDM --- a different cohort of children scanned on the same stimuli --- is recovered by this exact pipeline (z-normalise both, Spearman on the upper triangle) at the rho above, and a probe mixed at w is detected in every cell. So "no LM alignment is detectable here" is a bounded statement about models, not a suspicion about the measurement. Second, and less comfortably, a trivial control outscores every language model*: an additive per-stimulus main-effect RDM (meanD_i + meanD_j, carrying no relational structure at all) is significant by stimulus-label permutation in 6 of 6 cells. Whatever these RDMs are dominated by, it is closer to a per-stimulus offset than to the relational geometry RSA is meant to compare. additive_control reports it per cell; no previously published artifact in this collection does.

2c-ii. The scanner-run confound is not what suppresses alignment. ds006239/SemLocal is the only genuinely run x stimulus crossed cell in the whole collection --- the run confound cannot arise there by design --- so it is the sharpest available test of the "it is the confound" explanation. Expressed as a ratio to each cell's own external-RDM benchmark (which controls for the very different ceilings), SemLocal scores R = 0.106 against a median of 0.046 across the 24 confounded cells (IQR 0.035--0.086, range 0.029--0.199). It is the high end of that distribution but inside it, at the 83rd percentile. The clean cell behaves like the dirty ones. Whatever is holding model alignment near zero here, the scanner-run confound is not it, and that explanation should stop being offered.

2d. The pipeline is exactly reproducible, except for numerical precision. Five contrasts were run on the same family and cells, changing only things that should be no-ops. Ratios are the mean \|delta rsa\| at the final checkpoint as a fraction of the between-family sd at that cell --- i.e. how much of the signal a nuisance parameter can move:

contrastratiomaxcells over 0.5
different GPU (1 vs 2)0.0000.0000/26
different GPU + different day (vs the main sweep)0.0000.0000/26
batch size 16 vs 40.00070.0030/26
batch size 16 vs 320.00050.0020/26
fp32 vs bf160.3160.9546/26

Device and run-to-run results are bit-identical to 1e-12, and batch size is negligible. Precision is the sole exception, and it grows monotonically with training: at step 0 the two precisions agree to 2e-4, but by the final checkpoint the disagreement reaches 32% of the entire between-family spread and exceeds half of it in 6 of 26 cells. This whole grid is fp32, so its numbers are internally consistent and comparable to each other; do not compare a row here against a bf16 number computed elsewhere, and treat small between-family differences at trained checkpoints as noise.

3. Ceilings are low on this dataset, so use `frac_of_ceiling`, not `rsa`. The inter-subject noise ceilings are in noise_ceilings and joined onto every alignment row. Best raw rsa anywhere in this grid is 0.1032; as a fraction of the cell's own ceiling that is 0.404.

All RDMs are the `within-run-normalised` ones. The uncorrected RDMs carry a scanner-run confound in which run identity predicts brain dissimilarity at ρ = +0.49…+0.87 while no stimulus property predicts it at all; z-scoring each voxel within run drops that to ≈−0.04. Do not mix the two.


Why this repo exists

`BrainAlign/brain-lm-alignment-ds002236` covered 2 of 6 task × session cells. The cause is a launcher bug, not a scientific choice: slurm/run_devai_grid.sh never passes --sessions, so the runner silently fell back to the ds003604 default ["ses-5","ses-7","ses-9"] — sessions this dataset does not have. Sessions and tasks are derived from the RDM tree here, which is what takes coverage to all 6 cells.

Method

Per (checkpoint × task × session): feed the RDM file's own stimulus_texts through the model, mean-pool the final block's hidden states over tokens, build the model RDM as 1 − corrcoef across stimuli, then correlate the upper triangles of the z-normalised model and brain RDMs. rsa is Spearman (headline); rsa_pearson and rsa_kendall are also reported, and frac_of_ceiling is rsa / ceiling_lower.

Checkpoints are subsampled log-uniformly across each family's training trajectory, so a family's rows trace its development rather than only its final state.

Sessions and tasks

ses-9, ses-11, ses-11+ × Phon, Sem.

claim_tests carries the upstream per-family tests (R1 alignment-rises, R2 alignment-vs-mechanistic with a step-partialled control, R5 isolation-vs-mechanistic, R2b behaviour tests) computed by the same mechanistic_brain_analysis.py used for cdl-devai-results.