datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
brain-lm-alignment-ds002236
Brain–language-model alignment: ds002236 (whole-brain)
Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children (8.7–15.5), auditory and visual.
Paper: https://pubmed.ncbi.nlm.nih.gov/31956678/
Data: https://openneuro.org/datasets/ds002236/versions/1.0.1
Generated: 2026-09-26
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number in… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds002236.brain-lm-alignment-ds006239
Brain–language-model alignment: ds006239 (whole-brain)
Wang et al. 2025 — word-level phonological and semantic reading tasks in children and adolescents aged 10–17.
Paper: https://www.sciencedirect.com/science/article/pii/S2352340925009692
Data: https://openneuro.org/datasets/ds006239/versions/1.0.5
Generated: 2026-09-26
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds006239.M4Raw_brain
M4Raw Brain v1.6
M4Raw Brain is a multi-contrast, multi-repetition, four-channel k-space
dataset acquired on a 0.3-T whole-body MRI system. The release contains T1w,
T2w, FLAIR, and GRE brain acquisitions from healthy volunteers, including
explicit motion subsets and an expanded repeated-acquisition test cohort.
Companion dataset: M4Raw-Abdomen
is a separate low-field abdominal MRI k-space and segmentation dataset and
will be made public soon. Until then, the linked private… See the full description on the dataset page: https://huggingface.co/datasets/mylyu/M4Raw_brain.brain-lm-alignment-ds001894
Brain–language-model alignment: ds001894 (whole-brain)
Lytle et al. 2019 — longitudinal word-level phonological processing in children scanned twice, at roughly 10 and 12 years old.
Paper: https://www.nature.com/articles/s41597-019-0338-5
Data: https://openneuro.org/datasets/ds001894/versions/1.4.2
Generated: 2026-09-26
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds001894.cdl-devai-results-ds006239
ds006239 (Wang et al. 2025) — brain × interpretability × localisation, per model per checkpoint
Wang et al. 2025 — word-level phonological and semantic reading in children and adolescents. Cohort: children and adolescents 10–17 years; presentation: visual (reading).
Tasks Orth, Phon, Sem, SemLocal × sessions ses-11, ses-11+ = 8 task × session cells,
all of them scored here.
Duplicate cells. Phon/ses-11+ is bit-identical to Orth/ses-11+, Phon/ses-11 is bit-identical to… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds006239.cdl-devai-results-ds002236
ds002236 (Lytle et al. 2020) — brain × interpretability × localisation, per model per checkpoint
Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children. Cohort: children 8.7–15.5 years; presentation: auditory and visual word presentation.
Tasks Phon, Sem × sessions ses-9, ses-11, ses-11+ = 6 task × session cells,
all of them scored here.
Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds002236.cdl-devai-results-ds003604-roiauditory
ds003604 (Wang et al. 2022) — brain × interpretability × localisation, per model per checkpoint
Wang et al. 2022 — auditory language comprehension in children. Cohort: children scanned at 5, 7 and 9; presentation: auditory.
Tasks Sem, Phon, Gram, Plaus × sessions ses-5, ses-7, ses-9 = 12 task × session cells,
all of them scored here.
Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4 checkpoints, pythia-6.9b-full has 1 checkpoint.… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds003604-roiauditory.cdl-devai-results-ds003604-roiphonology
ds003604 (Wang et al. 2022) — brain × interpretability × localisation, per model per checkpoint
Wang et al. 2022 — auditory language comprehension in children. Cohort: children scanned at 5, 7 and 9; presentation: auditory.
Tasks Sem, Phon, Gram, Plaus × sessions ses-5, ses-7, ses-9 = 12 task × session cells,
all of them scored here.
Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4 checkpoints. pythia-6.9b-full's single checkpoint is step… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds003604-roiphonology.cdl-devai-results-ds003604-roimotor
ds003604 (Wang et al. 2022) — brain × interpretability × localisation, per model per checkpoint
Wang et al. 2022 — auditory language comprehension in children. Cohort: children scanned at 5, 7 and 9; presentation: auditory.
Tasks Sem, Phon, Gram, Plaus × sessions ses-5, ses-7, ses-9 = 12 task × session cells,
all of them scored here.
Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4 checkpoints. pythia-6.9b-full's single checkpoint is step… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds003604-roimotor.cdl-devai-results-ds003604
ds003604 (Wang et al. 2022) — brain × interpretability × localisation, per model per checkpoint
Wang et al. 2022 — auditory language comprehension in children. Cohort: children scanned at 5, 7 and 9; presentation: auditory.
Tasks Sem, Phon, Gram, Plaus × sessions ses-5, ses-7, ses-9 = 12 task × session cells,
all of them scored here.
Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4 checkpoints, pythia-6.9b-full has 1 checkpoint.… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds003604.cdl-devai-results
CDL DevAI results — brain × interpretability × localisation, per model per checkpoint
Developmental analysis of 10 language-model families against the ds003604 auditory
language fMRI dataset. For every training checkpoint of every model we measured three
things and here report them side by side:
axis
what it asks
source tables
brain
does the model's representational geometry match the brain's?
brain_alignment
interp
how is the representation organised internally?… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results.pubmed_arxiv_abstracts_dataannotations_creators:
machine-generated
language:
en
license:
apache-2.0
multilinguality:
monolingual
task_categories:
classification
generation
pretty_name: PubMed_ArXiv_Abstracts
features:
name: abstr
dtype: string
name: title
dtype: string
name: journal
dtype: string
name: field
dtype: string
name: label_journal
dtype: int64
name: label_field
dtype: int64
brain-lm-alignment-ds003604
Brain-LM alignment: ds003604
Representational-similarity alignment between language-model hidden states and
child fMRI RDMs for ds003604 (children ages 5/7/9, auditory).
Tasks: Sem, Phon, Gram, Plaus Sessions: ses-5, ses-7, ses-9 Cells: 12
Models: 14 families (5 real + 9 PARC noise-seed baselines)
Rows: 1848 (family x checkpoint x task x session)
Generated: 2026-08-29
Headline: no model is distinguishable from a random seed
Alignment is computed as Spearman… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds003604.enhanced_brain_magnet_hg38For details and usage of these datasets please see my GitHub repository: https://github.com/caenrigen/enhanced_brain_magnet.
brainteaser_splitedbrainprint-benchmarks
DMIT BrainPrint Analytics Benchmarks
Benchmark dataset of 20 DMIT assessment cases with individual scores for BrainPrint, cognitive profile, learning style, personality insight, career pathway, and leadership & workplace signals.
Guided by Mrs. Priyanka Swain, Founder of Merit Teacher. Built by DMIT.fyi.
Dataset Description
This dataset contains benchmark data for an intelligent assessment platform supporting fingerprint analysis, cognitive profiling, learning… See the full description on the dataset page: https://huggingface.co/datasets/dmit-fyi/brainprint-benchmarks.Brain-Computer-Interaction-Eye-State-Classificationtheme-indices-5y3d_nyxusmed_mr_braintest-embeddingsldwcnet-brain-mri-modelinsult_external_dataldwcnet-brain-mri-2ldwcnet-brain-mri-3brain_drain_competition
