CoolFace
Datasetpublic

BrainAlign/brain-lm-alignment-ds002236

Brain–language-model alignment: ds002236 (whole-brain) Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children (8.7–15.5), auditory and visual. Paper: https://pubmed.ncbi.nlm.nih.gov/31956678/ Data: https://openneuro.org/datasets/ds002236/versions/1.0.1 Generated: 2026-09-21 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number in… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds002236.

sourceHugging Faceupdated 33m agoView on Hugging Face
0likes4kdownloads
Dataset Card

Brain–language-model alignment: ds002236 (whole-brain)

Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children (8.7–15.5), auditory and visual.

  • Paper: https://pubmed.ncbi.nlm.nih.gov/31956678/
  • Data: https://openneuro.org/datasets/ds002236/versions/1.0.1
  • Generated: 2026-09-21
  • Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms

Read this first: does the measurement work?

Every alignment number in this dataset is only as meaningful as the brain RDMs it was computed against. So before any model result, the same pipeline is asked whether anything stimulus-driven correlates with those RDMs — stimulus duration, intensity, word length, frequency, phoneme and syllable counts, an acoustic model of the audio where the stimuli are audio, and the study's own condition contrast — each tested by a permutation test that shuffles stimulus identity.

GATE: FAILED. 0/30 stimulus tests are significant after Holm correction — not the acoustic model of the audio the children actually heard, not the study's own experimental contrast.

The alignment numbers below are therefore uninterpretable as evidence about language models. They measure a representational geometry that does not demonstrably encode the stimuli. They are published for completeness and for whoever fixes the estimator, not as a result. Do not cite them as evidence that models fail to align with the developing brain.

Measured cause, from control/:

  • RDM effective rank: 62 of 84 stimuli

Note that this is NOT ds003604's failure mode. There, the RDM effective rank was ~3 of 40-48 stimuli -- near-degenerate betas that could not express stimulus-level structure at all. The rank recorded above is a large fraction of the stimulus count, so these RDMs do carry stimulus structure and the control failing here means the specific controls tested did not reach significance, not that the measurement is uninterpretable. Check control/ for which controls ran: an acoustic or visual control needs the dataset's stimulus files present, and reports zero features if they are not.

What was built

126 task × session cells, each an RDM over the stimuli shared by that cell's subjects, with voxel patterns z-scored within run before aggregation (without that, the RDM measures scanner drift rather than language) and an inter-subject noise ceiling.

tasksessionn_stimceiling_lowerceiling_upperceiling_n
Phonses-11+960.3141080.39645826
Phonses-11960.25480.37017818
Phonses-9960.2747330.40354615
Semses-11+720.4648550.51918827
Semses-11720.399310.4798220
Semses-9720.3876120.46931321
Phonses-11+960.2705560.36424326
Phonses-11960.2279850.3601418
Phonses-9960.2310320.38222115
Semses-11+720.3607750.43444727
Semses-11720.3121740.41919820
Semses-9720.3244440.42041221
Phonses-11+960.2308440.5027736
Phonses-1196nannannan
Phonses-9960.172550.5448114
Semses-11+720.2701510.5031897
Semses-1172nannannan
Semses-9720.1941780.6267673
Phonses-11+960.2429620.40405113
Phonses-11960.1565860.4602976
Phonses-9960.1864860.4742526
Semses-11+720.3273530.49233710
Semses-11720.2846810.537446
Semses-9720.3021820.4840299
Phonses-11+960.2615790.36909622
Phonses-11960.2278120.38504814
Phonses-9960.2275630.40242712
Semses-11+720.3526810.44057622
Semses-11720.2974310.44478213
Semses-9720.33060.45101115
Phonses-11+960.2439640.6526723
Phonses-1196nannannan
Phonses-996nannannan
Semses-11+720.2397790.6551663
Semses-1172nannannan
Semses-972nannannan
Phonses-11+960.2531740.35165826
Phonses-11960.2200810.35257618
Phonses-9960.2238490.37646415
Semses-11+720.3711680.44477427
Semses-11720.3303330.42952520
Semses-9720.330790.42720621
Phonses-11+960.2234970.4915346
Phonses-1196nannannan
Phonses-9960.16550.5363684
Semses-11+720.3215590.5295867
Semses-1172nannannan
Semses-9720.245930.6493453
Phonses-11+960.2281010.39358413
Phonses-11960.1764170.4582396
Phonses-9960.1988750.4763766
Semses-11+720.3594910.51055710
Semses-11720.3284610.5587456
Semses-9720.3095080.48979
Phonses-11+960.2423230.35563522
Phonses-11960.2142440.37408414
Phonses-9960.2160550.39607412
Semses-11+720.3590660.44787122
Semses-11720.3118710.45381513
Semses-9720.3299280.45184615
Phonses-11+960.2119850.6275343
Phonses-1196nannannan
Phonses-996nannannan
Semses-11+720.2731450.6573693
Semses-1172nannannan
Semses-972nannannan
Phonses-11+960.2046120.32270226
Phonses-11960.1645030.32011118
Phonses-9960.174350.34601415
Semses-11+720.3093340.39777927
Semses-11720.2547840.37534320
Semses-9720.2760020.38609721
Phonses-11+960.2076480.492226
Phonses-1196nannannan
Phonses-9960.1326450.5178834
Semses-11+720.2672850.5019887
Semses-1172nannannan
Semses-9720.1490760.606933
Phonses-11+960.2063340.3861913
Phonses-11960.1360020.4432616
Phonses-9960.1486890.4516956
Semses-11+720.3129910.48452110
Semses-11720.2825540.5342156
Semses-9720.2672490.464589
Phonses-11+960.2057690.33643322
Phonses-11960.1562610.34042714
Phonses-9960.1607540.36170312
Semses-11+720.3079130.41185622
Semses-11720.2440360.40771313
Semses-9720.2788330.41714915
Phonses-11+960.2190810.6468953
Phonses-1196nannannan
Phonses-996nannannan
Semses-11+720.2014610.6380933
Semses-1172nannannan
Semses-972nannannan
Phonses-11+960.283370.37403126
Phonses-11960.2308840.3617318
Phonses-9960.2521890.39201115
Semses-11+720.3978350.46406827
Semses-11720.3393230.43769220
Semses-9720.3556210.44312721
Phonses-11+960.2482840.512466
Phonses-1196nannannan
Phonses-9960.1690660.5328194
Semses-11+720.3249820.5296167
Semses-1172nannannan
Semses-9720.1887550.6176893
Phonses-11+960.2554420.41455413
Phonses-11960.1813340.4655646
Phonses-9960.19060.4706826
Semses-11+720.3860070.52920810
Semses-11720.3574670.5738246
Semses-9720.3097230.4911489
Phonses-11+960.2689210.37671922
Phonses-11960.230840.38502814
Phonses-9960.239890.40612312
Semses-11+720.3905830.4695422
Semses-11720.3316680.46684813
Semses-9720.3599160.4704715
Phonses-11+960.2821380.6689093
Phonses-1196nannannan
Phonses-996nannannan
Semses-11+720.2731810.6611823
Semses-1172nannannan
Semses-972nannannan

Model grid: 15 families, 33012 alignment rows across 6 cells.

mean noise ceiling0.264
best alignment anywhere0.1219
as a fraction of ceiling89.6%
families equivalent to zero (TOST ±0.05)13/15
Pythia scale trendρ = +0.147, p = 0.44

Per family

familyn_checkpointsrsa_meanrsa_sdrsa_abs_maxfrac_of_ceiling_abs_maxp_equivalence_tost
babylm-gpt290.02290.03830.07660.56290.0716
pythia-1b-full210.02280.01520.09230.67840.0036
pico-decoder-medium210.02220.02610.12190.89650.0237
babylm-gpt2-790.02220.03510.07250.53340.055
babylm-gpt2-390.02110.03350.06650.48930.0439
pico-decoder-large210.01990.02710.09290.70060.0209
babylm-gpt2-590.01940.03460.06990.51380.0413
pythia-1.4b-full210.01620.02080.11150.81960.0053
pythia-70m-full210.0150.01450.0880.65230.001
pythia-410m-full210.01490.01470.09310.68450.001
pythia-160m-full210.01220.01370.08720.62630.0005
pico-decoder-small210.010.02070.09570.72140.0026
pico-decoder-tiny210.00780.01680.08240.5570.0008
beetle-humanscale-eng180.00630.00460.0790.58060
beetle-fineweb3-eng190.00040.00480.07430.56050

Dataset-specific notes

The accession is not stated in the data article; it was resolved to ds002236 by matching OpenNeuro's own dataset name ("Cross-Sectional Multidomain Lexical Processing") AND the per-subject age range in participants.tsv (8.67–15.5) against the range the article reports. Best developmental axis of the four datasets: explicit per-subject age at scan, continuous rather than binned. Six tasks crossing modality (auditory/visual) with judgement (rhyme/spelling/semantic) — a modality control no other dataset here provides. A third of trials are coded null (Tones/nullsilence.WAV) and are excluded from the stimulus set.

Files

pathwhatpresent here
alignment_by_checkpoint.csvevery model × checkpoint × cell, with ceiling
alignment_by_family.csvper family, with equivalence tests
alignment_by_cell.csvper task × session
ceilings_ds002236.csvnoise ceiling per cell
control/the positive control and RDM dimensionality — the gate
scale_ladder.csvthe Pythia 70M→1.4B scale test
fig_*.pdf, fig_*.pngfigures

Method

Representational similarity analysis. For each cell, a brain RDM over stimuli (correlation distance between per-stimulus GLM beta patterns, within-run z-scored, aggregated across subjects) is compared by Spearman correlation with a model RDM over the same stimuli, taken from each checkpoint's hidden states. Alignment is reported raw and as a fraction of the inter-subject noise ceiling, and judged against a null built from the PARC suite — 18 models differing only by random seed, which is what 'no effect' looks like on this measurement.

Null and fixation trials are excluded from the stimulus set. For paired designs the stimulus identity is the pair, not either word alone.