CoolFace
Datasetpublic

BrainAlign/brain-lm-alignment-ds006239

Brain–language-model alignment: ds006239 (whole-brain) Wang et al. 2025 — word-level phonological and semantic reading tasks in children and adolescents aged 10–17. Paper: https://www.sciencedirect.com/science/article/pii/S2352340925009692 Data: https://openneuro.org/datasets/ds006239/versions/1.0.5 Generated: 2026-09-21 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds006239.

sourceHugging Faceupdated 2h agoView on Hugging Face
2likes3.8kdownloads
Dataset Card

Brain–language-model alignment: ds006239 (whole-brain)

Wang et al. 2025 — word-level phonological and semantic reading tasks in children and adolescents aged 10–17.

  • Paper: https://www.sciencedirect.com/science/article/pii/S2352340925009692
  • Data: https://openneuro.org/datasets/ds006239/versions/1.0.5
  • Generated: 2026-09-21
  • Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms

Read this first: does the measurement work?

Every alignment number in this dataset is only as meaningful as the brain RDMs it was computed against. So before any model result, the same pipeline is asked whether anything stimulus-driven correlates with those RDMs — stimulus duration, intensity, word length, frequency, phoneme and syllable counts, an acoustic model of the audio where the stimuli are audio, and the study's own condition contrast — each tested by a permutation test that shuffles stimulus identity.

GATE: FAILED. 0/38 stimulus tests are significant after Holm correction — not the acoustic model of the audio the children actually heard, not the study's own experimental contrast.

The alignment numbers below are therefore uninterpretable as evidence about language models. They measure a representational geometry that does not demonstrably encode the stimuli. They are published for completeness and for whoever fixes the estimator, not as a result. Do not cite them as evidence that models fail to align with the developing brain.

Measured cause, from control/:

  • RDM effective rank: 61 of 84 stimuli

Note that this is NOT ds003604's failure mode. There, the RDM effective rank was ~3 of 40-48 stimuli -- near-degenerate betas that could not express stimulus-level structure at all. The rank recorded above is a large fraction of the stimulus count, so these RDMs do carry stimulus structure and the control failing here means the specific controls tested did not reach significance, not that the measurement is uninterpretable. Check control/ for which controls ran: an acoustic or visual control needs the dataset's stimulus files present, and reports zero features if they are not.

What was built

168 task × session cells, each an RDM over the stimuli shared by that cell's subjects, with voxel patterns z-scored within run before aggregation (without that, the RDM measures scanner drift rather than language) and an inter-subject noise ceiling.

tasksessionn_stimceiling_lowerceiling_upperceiling_n
Orthses-11+960.5653630.59655648
Orthses-11960.5256370.56432739
Phonses-11+960.5653630.59655648
Phonses-11960.5256370.56432739
Semses-11+720.3901390.43324250
Semses-11720.3510290.40524239
SemLocalses-11+480.3100010.35995849
SemLocalses-11480.264770.33166934
Orthses-11+960.5653630.59655648
Orthses-11960.5256370.56432739
Phonses-11+960.5653630.59655648
Phonses-11960.5256370.56432739
Semses-11+720.3901390.43324250
Semses-11720.3510290.40524239
SemLocalses-11+480.3100010.35995849
SemLocalses-11480.264770.33166934
Orthses-11+960.5524020.7073025
Orthses-11960.5084390.6451197
Phonses-11+960.5524020.7073025
Phonses-11960.5084390.6451197
Semses-11+720.3707440.5917365
Semses-11720.29280.4890777
SemLocalses-11+480.2654940.5265545
SemLocalses-11480.2354270.4770016
Orthses-11+960.554340.62872914
Orthses-11960.5287290.62308511
Phonses-11+960.554340.62872914
Phonses-11960.5287290.62308511
Semses-11+720.4063670.50507214
Semses-11720.3234610.45930211
SemLocalses-11+480.2910450.42037414
SemLocalses-11480.240010.41541110
Orthses-11+960.5641850.60966826
Orthses-11960.5329460.58366124
Phonses-11+960.5641850.60966826
Phonses-11960.5329460.58366124
Semses-11+720.4048850.47248226
Semses-11720.3460740.42430124
SemLocalses-11+480.305260.38228827
SemLocalses-11480.2484810.34895821
Orthses-11+960.5275090.7608263
Orthses-11960.3957010.6947773
Phonses-11+960.5275090.7608263
Phonses-11960.3957010.6947773
Semses-11+720.3135650.6580193
Semses-11720.1854780.5869453
SemLocalses-11+480.2085210.6017443
SemLocalses-11480.2164610.6070913
Orthses-11+960.5653630.59655648
Orthses-11960.5256370.56432739
Phonses-11+960.5653630.59655648
Phonses-11960.5256370.56432739
Semses-11+720.3901390.43324250
Semses-11720.3510290.40524239
SemLocalses-11+480.3100010.35995849
SemLocalses-11480.264770.33166934
Orthses-11+960.5524020.7073025
Orthses-11960.5084390.6451197
Phonses-11+960.5524020.7073025
Phonses-11960.5084390.6451197
Semses-11+720.3707440.5917365
Semses-11720.29280.4890777
SemLocalses-11+480.2654940.5265545
SemLocalses-11480.2354270.4770016
Orthses-11+960.554340.62872914
Orthses-11960.5287290.62308511
Phonses-11+960.554340.62872914
Phonses-11960.5287290.62308511
Semses-11+720.4063670.50507214
Semses-11720.3234610.45930211
SemLocalses-11+480.2910450.42037414
SemLocalses-11480.240010.41541110
Orthses-11+960.5641850.60966826
Orthses-11960.5329460.58366124
Phonses-11+960.5641850.60966826
Phonses-11960.5329460.58366124
Semses-11+720.4048850.47248226
Semses-11720.3460740.42430124
SemLocalses-11+480.305260.38228827
SemLocalses-11480.2484810.34895821
Orthses-11+960.5275090.7608263
Orthses-11960.3957010.6947773
Phonses-11+960.5275090.7608263
Phonses-11960.3957010.6947773
Semses-11+720.3135650.6580193
Semses-11720.1854780.5869453
SemLocalses-11+480.2085210.6017443
SemLocalses-11480.2164610.6070913
Orthses-11+960.5653630.59655648
Orthses-11960.5256370.56432739
Phonses-11+960.5653630.59655648
Phonses-11960.5256370.56432739
Semses-11+720.3901390.43324250
Semses-11720.3510290.40524239
SemLocalses-11+480.3100010.35995849
SemLocalses-11480.264770.33166934
Orthses-11+960.5524020.7073025
Orthses-11960.5084390.6451197
Phonses-11+960.5524020.7073025
Phonses-11960.5084390.6451197
Semses-11+720.3707440.5917365
Semses-11720.29280.4890777
SemLocalses-11+480.2654940.5265545
SemLocalses-11480.2354270.4770016
Orthses-11+960.554340.62872914
Orthses-11960.5287290.62308511
Phonses-11+960.554340.62872914
Phonses-11960.5287290.62308511
Semses-11+720.4063670.50507214
Semses-11720.3234610.45930211
SemLocalses-11+480.2910450.42037414
SemLocalses-11480.240010.41541110
Orthses-11+960.5641850.60966826
Orthses-11960.5329460.58366124
Phonses-11+960.5641850.60966826
Phonses-11960.5329460.58366124
Semses-11+720.4048850.47248226
Semses-11720.3460740.42430124
SemLocalses-11+480.305260.38228827
SemLocalses-11480.2484810.34895821
Orthses-11+960.5275090.7608263
Orthses-11960.3957010.6947773
Phonses-11+960.5275090.7608263
Phonses-11960.3957010.6947773
Semses-11+720.3135650.6580193
Semses-11720.1854780.5869453
SemLocalses-11+480.2085210.6017443
SemLocalses-11480.2164610.6070913
Orthses-11+960.5653630.59655648
Orthses-11960.5256370.56432739
Phonses-11+960.5653630.59655648
Phonses-11960.5256370.56432739
Semses-11+720.3901390.43324250
Semses-11720.3510290.40524239
SemLocalses-11+480.3100010.35995849
SemLocalses-11480.264770.33166934
Orthses-11+960.5524020.7073025
Orthses-11960.5084390.6451197
Phonses-11+960.5524020.7073025
Phonses-11960.5084390.6451197
Semses-11+720.3707440.5917365
Semses-11720.29280.4890777
SemLocalses-11+480.2654940.5265545
SemLocalses-11480.2354270.4770016
Orthses-11+960.554340.62872914
Orthses-11960.5287290.62308511
Phonses-11+960.554340.62872914
Phonses-11960.5287290.62308511
Semses-11+720.4063670.50507214
Semses-11720.3234610.45930211
SemLocalses-11+480.2910450.42037414
SemLocalses-11480.240010.41541110
Orthses-11+960.5641850.60966826
Orthses-11960.5329460.58366124
Phonses-11+960.5641850.60966826
Phonses-11960.5329460.58366124
Semses-11+720.4048850.47248226
Semses-11720.3460740.42430124
SemLocalses-11+480.305260.38228827
SemLocalses-11480.2484810.34895821
Orthses-11+960.5275090.7608263
Orthses-11960.3957010.6947773
Phonses-11+960.5275090.7608263
Phonses-11960.3957010.6947773
Semses-11+720.3135650.6580193
Semses-11720.1854780.5869453
SemLocalses-11+480.2085210.6017443
SemLocalses-11480.2164610.6070913

Model grid: 15 families, 44016 alignment rows across 8 cells.

mean noise ceiling0.413
best alignment anywhere0.0691
as a fraction of ceiling25.6%
families equivalent to zero (TOST ±0.05)15/15
Pythia scale trendρ = -0.348, p = 0.03

Per family

familyn_checkpointsrsa_meanrsa_sdrsa_abs_maxfrac_of_ceiling_abs_maxp_equivalence_tost
pico-decoder-tiny210.00530.00580.05380.1770
pico-decoder-small210.00350.01250.05220.24110
babylm-gpt290.0020.01040.03890.13370
pico-decoder-large210.00150.0090.04740.25570
pico-decoder-medium21-0.00070.00860.04760.21990
beetle-humanscale-eng18-0.00140.01220.04230.2030
pythia-70m-full21-0.00160.01030.03820.16990
pythia-410m-full21-0.00460.01020.04580.21980
beetle-fineweb3-eng19-0.00570.01190.05040.23490
pythia-1b-full21-0.00980.00450.04610.17790
pythia-160m-full21-0.01160.00580.04230.19540
pythia-1.4b-full21-0.01350.00440.05950.21160
babylm-gpt2-59-0.02620.03360.06630.16770.0427
babylm-gpt2-79-0.02650.0340.06910.17460.0457
babylm-gpt2-39-0.0270.03360.0670.16940.0471

Dataset-specific notes

Contains LocalSem, the only genuinely run/stimulus-CROSSED language cell across all four datasets in this project: its stimuli recur across runs, so run identity and stimulus identity are separable and the scanner-run confound that invalidated the first ds003604 analysis cannot arise. Per-subject age is NOT recoverable from the release — participants.tsv has birthdate but no scan date and there are no *_scans.tsv files — so this dataset is cohort-level only and cannot carry the developmental axis as published.

Files

pathwhatpresent here
alignment_by_checkpoint.csvevery model × checkpoint × cell, with ceiling
alignment_by_family.csvper family, with equivalence tests
alignment_by_cell.csvper task × session
ceilings_ds006239.csvnoise ceiling per cell
control/the positive control and RDM dimensionality — the gate
scale_ladder.csvthe Pythia 70M→1.4B scale test
fig_*.pdf, fig_*.pngfigures

Method

Representational similarity analysis. For each cell, a brain RDM over stimuli (correlation distance between per-stimulus GLM beta patterns, within-run z-scored, aggregated across subjects) is compared by Spearman correlation with a model RDM over the same stimuli, taken from each checkpoint's hidden states. Alignment is reported raw and as a fraction of the inter-subject noise ceiling, and judged against a null built from the PARC suite — 18 models differing only by random seed, which is what 'no effect' looks like on this measurement.

Null and fixation trials are excluded from the stimulus set. For paired designs the stimulus identity is the pair, not either word alone.