datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
brain-lm-alignment-ds002236
Brain–language-model alignment: ds002236 (whole-brain)
Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children (8.7–15.5), auditory and visual.
Paper: https://pubmed.ncbi.nlm.nih.gov/31956678/
Data: https://openneuro.org/datasets/ds002236/versions/1.0.1
Generated: 2026-09-25
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number in… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds002236.Mega-Brain-Distill
Mega-Brain-Distill
Curated merge of the top 10% highest-scoring examples from
584 community-uploaded LLM distillation/reasoning-trace datasets
on the Hub (Fable-5, Opus, GLM, Kimi, DeepSeek, GPT, MiniMax, Qwen traces,
etc.), deduplicated within and across all of them — many of these source
repos are the same underlying dump re-uploaded by different users.
Auto-generated by run.py — do not hand-edit, it will be overwritten on
the next run. Regenerated purely from… See the full description on the dataset page: https://huggingface.co/datasets/ShinMK3/Mega-Brain-Distill.brain-lm-alignment-ds006239
Brain–language-model alignment: ds006239 (whole-brain)
Wang et al. 2025 — word-level phonological and semantic reading tasks in children and adolescents aged 10–17.
Paper: https://www.sciencedirect.com/science/article/pii/S2352340925009692
Data: https://openneuro.org/datasets/ds006239/versions/1.0.5
Generated: 2026-09-25
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds006239.M4Raw_brain
M4Raw Brain v1.6
M4Raw Brain is a multi-contrast, multi-repetition, four-channel k-space
dataset acquired on a 0.3-T whole-body MRI system. The release contains T1w,
T2w, FLAIR, and GRE brain acquisitions from healthy volunteers, including
explicit motion subsets and an expanded repeated-acquisition test cohort.
Companion dataset: M4Raw-Abdomen
is a separate low-field abdominal MRI k-space and segmentation dataset and
will be made public soon. Until then, the linked private… See the full description on the dataset page: https://huggingface.co/datasets/mylyu/M4Raw_brain.brain-lm-alignment-ds001894
Brain–language-model alignment: ds001894 (whole-brain)
Lytle et al. 2019 — longitudinal word-level phonological processing in children scanned twice, at roughly 10 and 12 years old.
Paper: https://www.nature.com/articles/s41597-019-0338-5
Data: https://openneuro.org/datasets/ds001894/versions/1.4.2
Generated: 2026-09-25
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds001894.G1_WBT_Brainco_Collect_Plates_Into_Dishwasher
Data Structure
Observations
observation.state.ee_state (12)
End-effector states of the robot.
Computed via forward kinematics (FK) from the root link to the left and right end-effectors.
Includes the contribution of the waist.
Represented as concatenated poses of both end-effectors.
observation.state.hand_state (12 or 2)
Finger states for both hands. The dimensionality depends on the hand type.
Inspire Hand (range: 0.0… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_WBT_Brainco_Collect_Plates_Into_Dishwasher.general-master-en-202608
General · Master · English · 2026-08
English pretraining text, assembled from three public sources, cleaned with one
character-level cleaner, and filtered for repetition.
109,337,531 documents and 468,064,046,462 characters.
Composition
Config
Documents
Characters
What it is
fineweb-edu-dedup
65,010,430
297,544,916,118
Web text an educational classifier kept
cosmopedia-v2
38,591,146
144,011,993,012
Synthetic prose from a seeded generator… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-master-en-202608.G1_WBT_Brainco_Pickup_Pillow
Data Structure
Observations
observation.state.ee_state (12)
End-effector states of the robot.
Computed via forward kinematics (FK) from the root link to the left and right end-effectors.
Includes the contribution of the waist.
Represented as concatenated poses of both end-effectors.
observation.state.hand_state (12 or 2)
Finger states for both hands. The dimensionality depends on the hand type.
Inspire Hand (range: 0.0… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_WBT_Brainco_Pickup_Pillow.brainformer-e-largeG1_WBT_Brainco_Make_The_Bed
Data Structure
Observations
observation.state.ee_state (12)
End-effector states of the robot.
Computed via forward kinematics (FK) from the root link to the left and right end-effectors.
Includes the contribution of the waist.
Represented as concatenated poses of both end-effectors.
observation.state.hand_state (12 or 2)
Finger states for both hands. The dimensionality depends on the hand type.
Inspire Hand (range: 0.0… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_WBT_Brainco_Make_The_Bed.G1_WBT_Brainco_Pick_Up_Medicine_v0
Data Structure
Observations
observation.state.ee_state (12)
End-effector states of the robot.
Computed via forward kinematics (FK) from the root link to the left and right end-effectors.
Includes the contribution of the waist.
Represented as concatenated poses of both end-effectors.
observation.state.hand_state (12 or 2)
Finger states for both hands. The dimensionality depends on the hand type.
Inspire Hand (range: 0.0… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_WBT_Brainco_Pick_Up_Medicine_v0.general-web-it-202608
General · Web · Italian · 2026-08
Italian pretraining text, built from the Italian portion of EPFL's FineWeb2-HQ, which is the
high quality slice of FineWeb-2. Every document passes one character-level cleaner and a
repetition filter.
21,065,052 documents and 66,158,573,443 characters of Italian prose.
Contents
Config
Documents
Characters
Upstream
fineweb2-hq-ita_Latn
21,065,052
66,158,573,443
epfml/FineWeb2-HQ, ita_Latn
The character count is… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-it-202608.general-web-fr-202608
General · Web · French · 2026-08
French pretraining text, built from the French portion of EPFL's FineWeb2-HQ, which is the
high quality slice of FineWeb-2. Every document passes one character-level cleaner and a
repetition filter.
31,999,309 documents and 118,346,333,763 characters of French prose.
Contents
Config
Documents
Characters
Upstream
fineweb2-hq-fra_Latn
31,999,309
118,346,333,763
epfml/FineWeb2-HQ, fra_Latn
The character count is exact.… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-fr-202608.G1_Brainco_PickApple_DatasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Brainco",
"total_episodes": 200,
"total_frames": 153668,
"total_tasks": 1,
"total_videos": 800,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Brainco_PickApple_Dataset.G1_Brainco_GraspRubiksCube_DatasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Brainco",
"total_episodes": 197,
"total_frames": 220788,
"total_tasks": 1,
"total_videos": 788,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:197"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Brainco_GraspRubiksCube_Dataset.G1_Brainco_GraspOreo_DatasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Brainco",
"total_episodes": 201,
"total_frames": 234959,
"total_tasks": 1,
"total_videos": 804,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:201"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Brainco_GraspOreo_Dataset.brainformer-e-smallcdl-devai-results-ds006239
ds006239 (Wang et al. 2025) — brain × interpretability × localisation, per model per checkpoint
Wang et al. 2025 — word-level phonological and semantic reading in children and adolescents. Cohort: children and adolescents 10–17 years; presentation: visual (reading).
Tasks Orth, Phon, Sem, SemLocal × sessions ses-11, ses-11+ = 8 task × session cells,
all of them scored here.
Duplicate cells. Phon/ses-11+ is bit-identical to Orth/ses-11+, Phon/ses-11 is bit-identical to… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds006239.cdl-devai-results-ds002236
ds002236 (Lytle et al. 2020) — brain × interpretability × localisation, per model per checkpoint
Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children. Cohort: children 8.7–15.5 years; presentation: auditory and visual word presentation.
Tasks Phon, Sem × sessions ses-9, ses-11, ses-11+ = 6 task × session cells,
all of them scored here.
Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds002236.brainformer-e-mediumG1_Brainco_PickDoll_DatasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Brainco",
"total_episodes": 200,
"total_frames": 313401,
"total_tasks": 1,
"total_videos": 800,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Brainco_PickDoll_Dataset.cdl-devai-results-ds003604-roiauditory
ds003604 (Wang et al. 2022) — brain × interpretability × localisation, per model per checkpoint
Wang et al. 2022 — auditory language comprehension in children. Cohort: children scanned at 5, 7 and 9; presentation: auditory.
Tasks Sem, Phon, Gram, Plaus × sessions ses-5, ses-7, ses-9 = 12 task × session cells,
all of them scored here.
Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4 checkpoints, pythia-6.9b-full has 1 checkpoint.… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds003604-roiauditory.G1_Brainco_PickCharger_DatasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Brainco",
"total_episodes": 200,
"total_frames": 216532,
"total_tasks": 1,
"total_videos": 800,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Brainco_PickCharger_Dataset.brain-hackathon-2023-embed-datamulti-modal-derived-brain-network
PPMI Connectivity Graphs — HF Staging (Derivatives)
This dataset ships ready-to-use functional brain connectivity graphs derived from the PPMI cohort in a BIDS-ish derivatives layout. For each subject and parcellation, we include:
ROI time-series (*_desc-timeseries_parc-<name>.mat)
Pearson correlation connectivity matrix (*_desc-correlation_matrix_parc-<name>.mat)
JSON sidecars with summary fields (nodes, measure, symmetric/weighted flags)
Contents
data/… See the full description on the dataset page: https://huggingface.co/datasets/pakkinlau/multi-modal-derived-brain-network.G1_Brainco_PickDrink_DatasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Brainco",
"total_episodes": 201,
"total_frames": 174483,
"total_tasks": 1,
"total_videos": 804,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:201"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Brainco_PickDrink_Dataset.cdl-devai-results-ds003604-roiphonology
ds003604 (Wang et al. 2022) — brain × interpretability × localisation, per model per checkpoint
Wang et al. 2022 — auditory language comprehension in children. Cohort: children scanned at 5, 7 and 9; presentation: auditory.
Tasks Sem, Phon, Gram, Plaus × sessions ses-5, ses-7, ses-9 = 12 task × session cells,
all of them scored here.
Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4 checkpoints. pythia-6.9b-full's single checkpoint is step… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds003604-roiphonology.G1_Brainco_PickToothpaste_DatasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Brainco",
"total_episodes": 193,
"total_frames": 143863,
"total_tasks": 1,
"total_videos": 772,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:193"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Brainco_PickToothpaste_Dataset.G1_Brainco_PickTissues_DatasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Brainco",
"total_episodes": 206,
"total_frames": 198622,
"total_tasks": 1,
"total_videos": 824,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:206"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Brainco_PickTissues_Dataset.cdl-devai-results-ds003604-roimotor
ds003604 (Wang et al. 2022) — brain × interpretability × localisation, per model per checkpoint
Wang et al. 2022 — auditory language comprehension in children. Cohort: children scanned at 5, 7 and 9; presentation: auditory.
Tasks Sem, Phon, Gram, Plaus × sessions ses-5, ses-7, ses-9 = 12 task × session cells,
all of them scored here.
Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4 checkpoints. pythia-6.9b-full's single checkpoint is step… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds003604-roimotor.
