datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
brain-lm-alignment-ds002236
Brain–language-model alignment: ds002236 (whole-brain)
Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children (8.7–15.5), auditory and visual.
Paper: https://pubmed.ncbi.nlm.nih.gov/31956678/
Data: https://openneuro.org/datasets/ds002236/versions/1.0.1
Generated: 2026-09-25
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number in… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds002236.brain-lm-alignment-ds006239
Brain–language-model alignment: ds006239 (whole-brain)
Wang et al. 2025 — word-level phonological and semantic reading tasks in children and adolescents aged 10–17.
Paper: https://www.sciencedirect.com/science/article/pii/S2352340925009692
Data: https://openneuro.org/datasets/ds006239/versions/1.0.5
Generated: 2026-09-25
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds006239.M4Raw_brain
M4Raw Brain v1.6
M4Raw Brain is a multi-contrast, multi-repetition, four-channel k-space
dataset acquired on a 0.3-T whole-body MRI system. The release contains T1w,
T2w, FLAIR, and GRE brain acquisitions from healthy volunteers, including
explicit motion subsets and an expanded repeated-acquisition test cohort.
Companion dataset: M4Raw-Abdomen
is a separate low-field abdominal MRI k-space and segmentation dataset and
will be made public soon. Until then, the linked private… See the full description on the dataset page: https://huggingface.co/datasets/mylyu/M4Raw_brain.brain-lm-alignment-ds001894
Brain–language-model alignment: ds001894 (whole-brain)
Lytle et al. 2019 — longitudinal word-level phonological processing in children scanned twice, at roughly 10 and 12 years old.
Paper: https://www.nature.com/articles/s41597-019-0338-5
Data: https://openneuro.org/datasets/ds001894/versions/1.4.2
Generated: 2026-09-26
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds001894.cdl-devai-results-ds006239
ds006239 (Wang et al. 2025) — brain × interpretability × localisation, per model per checkpoint
Wang et al. 2025 — word-level phonological and semantic reading in children and adolescents. Cohort: children and adolescents 10–17 years; presentation: visual (reading).
Tasks Orth, Phon, Sem, SemLocal × sessions ses-11, ses-11+ = 8 task × session cells,
all of them scored here.
Duplicate cells. Phon/ses-11+ is bit-identical to Orth/ses-11+, Phon/ses-11 is bit-identical to… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds006239.cdl-devai-results-ds002236
ds002236 (Lytle et al. 2020) — brain × interpretability × localisation, per model per checkpoint
Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children. Cohort: children 8.7–15.5 years; presentation: auditory and visual word presentation.
Tasks Phon, Sem × sessions ses-9, ses-11, ses-11+ = 6 task × session cells,
all of them scored here.
Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds002236.cdl-devai-results-ds003604-roiauditory
ds003604 (Wang et al. 2022) — brain × interpretability × localisation, per model per checkpoint
Wang et al. 2022 — auditory language comprehension in children. Cohort: children scanned at 5, 7 and 9; presentation: auditory.
Tasks Sem, Phon, Gram, Plaus × sessions ses-5, ses-7, ses-9 = 12 task × session cells,
all of them scored here.
Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4 checkpoints, pythia-6.9b-full has 1 checkpoint.… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds003604-roiauditory.cdl-devai-results-ds003604-roiphonology
ds003604 (Wang et al. 2022) — brain × interpretability × localisation, per model per checkpoint
Wang et al. 2022 — auditory language comprehension in children. Cohort: children scanned at 5, 7 and 9; presentation: auditory.
Tasks Sem, Phon, Gram, Plaus × sessions ses-5, ses-7, ses-9 = 12 task × session cells,
all of them scored here.
Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4 checkpoints. pythia-6.9b-full's single checkpoint is step… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds003604-roiphonology.cdl-devai-results-ds003604-roimotor
ds003604 (Wang et al. 2022) — brain × interpretability × localisation, per model per checkpoint
Wang et al. 2022 — auditory language comprehension in children. Cohort: children scanned at 5, 7 and 9; presentation: auditory.
Tasks Sem, Phon, Gram, Plaus × sessions ses-5, ses-7, ses-9 = 12 task × session cells,
all of them scored here.
Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4 checkpoints. pythia-6.9b-full's single checkpoint is step… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds003604-roimotor.cdl-devai-results-ds003604
ds003604 (Wang et al. 2022) — brain × interpretability × localisation, per model per checkpoint
Wang et al. 2022 — auditory language comprehension in children. Cohort: children scanned at 5, 7 and 9; presentation: auditory.
Tasks Sem, Phon, Gram, Plaus × sessions ses-5, ses-7, ses-9 = 12 task × session cells,
all of them scored here.
Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4 checkpoints, pythia-6.9b-full has 1 checkpoint.… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds003604.cdl-devai-results
CDL DevAI results — brain × interpretability × localisation, per model per checkpoint
Developmental analysis of 10 language-model families against the ds003604 auditory
language fMRI dataset. For every training checkpoint of every model we measured three
things and here report them side by side:
axis
what it asks
source tables
brain
does the model's representational geometry match the brain's?
brain_alignment
interp
how is the representation organised internally?… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results.pubmed_arxiv_abstracts_dataannotations_creators:
machine-generated
language:
en
license:
apache-2.0
multilinguality:
monolingual
task_categories:
classification
generation
pretty_name: PubMed_ArXiv_Abstracts
features:
name: abstr
dtype: string
name: title
dtype: string
name: journal
dtype: string
name: field
dtype: string
name: label_journal
dtype: int64
name: label_field
dtype: int64
brain-lm-alignment-ds003604
Brain-LM alignment: ds003604
Representational-similarity alignment between language-model hidden states and
child fMRI RDMs for ds003604 (children ages 5/7/9, auditory).
Tasks: Sem, Phon, Gram, Plaus Sessions: ses-5, ses-7, ses-9 Cells: 12
Models: 14 families (5 real + 9 PARC noise-seed baselines)
Rows: 1848 (family x checkpoint x task x session)
Generated: 2026-08-29
Headline: no model is distinguishable from a random seed
Alignment is computed as Spearman… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds003604.enhanced_brain_magnet_hg38For details and usage of these datasets please see my GitHub repository: https://github.com/caenrigen/enhanced_brain_magnet.
brainteaser_splitedbrainstorming-thinking
brainstorming-thinking
Using Qwen3-14b to synthetically generate the reasoning traces and answers for Explore_Instruct_Brainstorming_10k
Suitable for LLM post-training, especially RL.
Brain_Tumor_pubmed_abstracts
This Dataset has been downloaded from PubMed
It has abstracts and titles that are related to Brain Tumors
the data has been cleaned before uploading
it could be used for any NLP task, such as Domain Adaptation
Brainrot-xK
About
A 3.81K rows synthetic dataset made using llama-3.1-8b-instant.
Sample
conversation_id,role,content
1,user,"I just watched the new season of ""Euphoria"" and I'm still thinking about the plot twists."
1,assistant,"omg sameeee!! i'm literallyyyy still shook fr the finale episode was hella sus, no cap. i was lowkey predicting the whole thang tho, bet u didnt see it coming rn"
2,user,I'm trying to decide between studying for my math exam or playing Overwatch with my… See the full description on the dataset page: https://huggingface.co/datasets/GoofyLM/Brainrot-xK.Brainrot-xK-large
About
A 10K rows synthetic dataset made using llama-3.1-8b-instant.
Sample
conversation_id,role,content
1,user,I'm having a lowkey dramatic day because I forgot my math homework at home and now I'm gonna be moooed out of class.
1,assistant,"omg u ate that math test?? nooooo!! u better get ur dad to drvie u home and grt that hw lol iykyk (btw, r u still stanin ur calc teacher??)"
2,user,I'm feeling super anxious today.
2,assistant,"ik what u r goin thru rn... dont 4get 2… See the full description on the dataset page: https://huggingface.co/datasets/GoofyLM/Brainrot-xK-large.Jennifer-brainwashbrainprint-benchmarks
DMIT BrainPrint Analytics Benchmarks
Benchmark dataset of 20 DMIT assessment cases with individual scores for BrainPrint, cognitive profile, learning style, personality insight, career pathway, and leadership & workplace signals.
Guided by Mrs. Priyanka Swain, Founder of Merit Teacher. Built by DMIT.fyi.
Dataset Description
This dataset contains benchmark data for an intelligent assessment platform supporting fingerprint analysis, cognitive profiling, learning… See the full description on the dataset page: https://huggingface.co/datasets/dmit-fyi/brainprint-benchmarks.Brain-tumor-effect-of-bloodpressure-sugarlevel-and-BMI-ON-ITS-PRESSENCEBrain-Computer-Interaction-Eye-State-ClassificationBrainOdatatheme-indices-5ybrain_mri3d_nyxusmed_mr_braintest-embeddingsbrainomixgenz_brainrot_dataset
