datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mega-Brain-Distill
Mega-Brain-Distill
Curated merge of the top 10% highest-scoring examples from
584 community-uploaded LLM distillation/reasoning-trace datasets
on the Hub (Fable-5, Opus, GLM, Kimi, DeepSeek, GPT, MiniMax, Qwen traces,
etc.), deduplicated within and across all of them — many of these source
repos are the same underlying dump re-uploaded by different users.
Auto-generated by run.py — do not hand-edit, it will be overwritten on
the next run. Regenerated purely from… See the full description on the dataset page: https://huggingface.co/datasets/ShinMK3/Mega-Brain-Distill.brain-lm-alignment-ds006239
Brain–language-model alignment: ds006239 (whole-brain)
Wang et al. 2025 — word-level phonological and semantic reading tasks in children and adolescents aged 10–17.
Paper: https://www.sciencedirect.com/science/article/pii/S2352340925009692
Data: https://openneuro.org/datasets/ds006239/versions/1.0.5
Generated: 2026-09-22
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds006239.brainteasers
Dataset Card for "brainteasers"
More Information needed
M4Raw_brain
M4Raw Brain v1.6
M4Raw Brain is a multi-contrast, multi-repetition, four-channel k-space
dataset acquired on a 0.3-T whole-body MRI system. The release contains T1w,
T2w, FLAIR, and GRE brain acquisitions from healthy volunteers, including
explicit motion subsets and an expanded repeated-acquisition test cohort.
Companion dataset: M4Raw-Abdomen
is a separate low-field abdominal MRI k-space and segmentation dataset and
will be made public soon. Until then, the linked private… See the full description on the dataset page: https://huggingface.co/datasets/mylyu/M4Raw_brain.single-cell-brain-zarr
Single-Cell Brain Zarr Collection
Production-ready brain single-cell RNA-seq data exported from the CellxGene Census into native Zarr stores for chunked, on-demand access on the Hugging Face Hub.
Why Zarr
Single-cell expression matrices get impractical fast if you treat them like ordinary dense files. Zarr is the point of this repo: it makes large atlas-scale data usable without forcing users to download or materialize the whole matrix before they can do anything useful.… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/single-cell-brain-zarr.brain-mri-dataset-140
Brain Tumor MRI Dataset (Visual Viewer Enabled)
This dataset contains structural MRI cross-sections processed from clinical scans, formatted into interactive image columns for direct streaming.
Dataset Features Map
image: Viewable cross-section slice image.
volume / slice: Core scan extraction reference coordinates.
Age / Survival_Days: Patient clinical records metrics.
Grade: Tumor classifications status (e.g., HGG, LGG).
anatomical_location: Specific scan… See the full description on the dataset page: https://huggingface.co/datasets/Satavisha2026/brain-mri-dataset-140.ocr-synthetic-cheque-datatsetbrain-lm-alignment-ds001894
Brain–language-model alignment: ds001894 (whole-brain)
Lytle et al. 2019 — longitudinal word-level phonological processing in children scanned twice, at roughly 10 and 12 years old.
Paper: https://www.nature.com/articles/s41597-019-0338-5
Data: https://openneuro.org/datasets/ds001894/versions/1.4.2
Generated: 2026-09-22
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds001894.brainformer-largebrainformer-mediumBrain-Stroke-Diagnosisgeneral-master-en-202608
General · Master · English · 2026-08
English pretraining text, assembled from three public sources, cleaned with one
character-level cleaner, and filtered for repetition.
109,337,531 documents and 468,064,046,462 characters.
Composition
Config
Documents
Characters
What it is
fineweb-edu-dedup
65,010,430
297,544,916,118
Web text an educational classifier kept
cosmopedia-v2
38,591,146
144,011,993,012
Synthetic prose from a seeded generator… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-master-en-202608.brainformer-e-largebrainformer-e-mediumBrainly_datasetbrainformer-smallbrain-teasersRadGenome-Brain_MRI_parquetbrainformer-small-v2general-web-it-202608
General · Web · Italian · 2026-08
Italian pretraining text, built from the Italian portion of EPFL's FineWeb2-HQ, which is the
high quality slice of FineWeb-2. Every document passes one character-level cleaner and a
repetition filter.
21,065,052 documents and 66,158,573,443 characters of Italian prose.
Contents
Config
Documents
Characters
Upstream
fineweb2-hq-ita_Latn
21,065,052
66,158,573,443
epfml/FineWeb2-HQ, ita_Latn
The character count is… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-it-202608.brain-instruction-tuningbrainformer-e-smallgeneral-web-fr-202608
General · Web · French · 2026-08
French pretraining text, built from the French portion of EPFL's FineWeb2-HQ, which is the
high quality slice of FineWeb-2. Every document passes one character-level cleaner and a
repetition filter.
31,999,309 documents and 118,346,333,763 characters of French prose.
Contents
Config
Documents
Characters
Upstream
fineweb2-hq-fra_Latn
31,999,309
118,346,333,763
epfml/FineWeb2-HQ, fra_Latn
The character count is exact.… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-fr-202608.Brain-tumor
Ultralytics Brain-tumor Dataset
Introduction
Ultralytics brain tumor detection dataset consists of medical images from MRI or CT scans, containing information about brain tumor presence, location, and characteristics. This dataset is essential for training computer vision algorithms to automate brain tumor identification, aiding in early diagnosis and treatment planning.
Sample Images and Annotations
Here are some examples of images from the dataset, along… See the full description on the dataset page: https://huggingface.co/datasets/Ultralytics/Brain-tumor.NeuronSpark-V1
NeuronSpark-V1 Pretraining Dataset
Bilingual (English + Chinese) pretraining corpus for NeuronSpark, a bio-inspired Spiking Neural Network language model.
Dataset Summary
Metric
Value
Total documents
17,174,734
Estimated tokens
~14.5B
Languages
English (55%), Chinese (42%), Bilingual Math (3%)
Format
Parquet (35 shards, ~39 GB)
Columns
text (string), source (string)
Sources & Composition
Source
Documents
Ratio
Est. Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Brain2nd/NeuronSpark-V1.NeuronSpark-Pretrain-v3
NeuronSpark-Pretrain-v3
Bilingual pretraining corpus for NeuronSpark v3, a bio-inspired Spiking Neural
Network language model with selective PLIF neurons and dynamic per-token compute
budget (PonderNet-v3).
Composition
Metric
Value
Total documents
18.2 M
Estimated tokens
~20 B
Format
37 Parquet shards (~1 GB each, zstd)
Schema
text: string, source: string
Languages
EN 55.6%, ZH 28.1%, code 16.3%
Deduplication
All source sampling is weighted so each… See the full description on the dataset page: https://huggingface.co/datasets/Brain2nd/NeuronSpark-Pretrain-v3.nejm-brain-to-text-sonified-istft
NEJM Brain-to-Text Sonified (iSTFT)
Pre-shuffled dataset (seed: 42) at 16kHz, 0-8000Hz range.
Sharded into 1000 files per shard for efficient loading.
Usage
from datasets import load_dataset
ds = load_dataset("ljcamargo/nejm-brain-to-text-sonified-istft")
cdl-devai-results-ds006239
ds006239 (Wang et al. 2025) — brain × interpretability × localisation, per model per checkpoint
Wang et al. 2025 — word-level phonological and semantic reading in children and adolescents. Cohort: children and adolescents 10–17 years; presentation: visual (reading).
Tasks Orth, Phon, Sem, SemLocal × sessions ses-11, ses-11+ = 8 task × session cells,
all of them scored here.
Duplicate cells. Phon/ses-11+ is bit-identical to Orth/ses-11+, Phon/ses-11 is bit-identical to… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds006239.cdl-devai-results-ds002236
ds002236 (Lytle et al. 2020) — brain × interpretability × localisation, per model per checkpoint
Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children. Cohort: children 8.7–15.5 years; presentation: auditory and visual word presentation.
Tasks Phon, Sem × sessions ses-9, ses-11, ses-11+ = 6 task × session cells,
all of them scored here.
Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds002236.cdl-devai-results-ds003604-roiauditory
ds003604 (Wang et al. 2022) — brain × interpretability × localisation, per model per checkpoint
Wang et al. 2022 — auditory language comprehension in children. Cohort: children scanned at 5, 7 and 9; presentation: auditory.
Tasks Sem, Phon, Gram, Plaus × sessions ses-5, ses-7, ses-9 = 12 task × session cells,
all of them scored here.
Incomplete families --- do not read these as scale-ladder points. pythia-2.8b-full has 4 checkpoints, pythia-6.9b-full has 1 checkpoint.… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/cdl-devai-results-ds003604-roiauditory.
