datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mega-Brain-Distill
Mega-Brain-Distill
Curated merge of the top 10% highest-scoring examples from
584 community-uploaded LLM distillation/reasoning-trace datasets
on the Hub (Fable-5, Opus, GLM, Kimi, DeepSeek, GPT, MiniMax, Qwen traces,
etc.), deduplicated within and across all of them — many of these source
repos are the same underlying dump re-uploaded by different users.
Auto-generated by run.py — do not hand-edit, it will be overwritten on
the next run. Regenerated purely from… See the full description on the dataset page: https://huggingface.co/datasets/ShinMK3/Mega-Brain-Distill.brainteasers
Dataset Card for "brainteasers"
More Information needed
brain-mri-dataset-140
Brain Tumor MRI Dataset (Visual Viewer Enabled)
This dataset contains structural MRI cross-sections processed from clinical scans, formatted into interactive image columns for direct streaming.
Dataset Features Map
image: Viewable cross-section slice image.
volume / slice: Core scan extraction reference coordinates.
Age / Survival_Days: Patient clinical records metrics.
Grade: Tumor classifications status (e.g., HGG, LGG).
anatomical_location: Specific scan… See the full description on the dataset page: https://huggingface.co/datasets/Satavisha2026/brain-mri-dataset-140.G1_WBT_Brainco_Collect_Plates_Into_Dishwasher
Data Structure
Observations
observation.state.ee_state (12)
End-effector states of the robot.
Computed via forward kinematics (FK) from the root link to the left and right end-effectors.
Includes the contribution of the waist.
Represented as concatenated poses of both end-effectors.
observation.state.hand_state (12 or 2)
Finger states for both hands. The dimensionality depends on the hand type.
Inspire Hand (range: 0.0… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_WBT_Brainco_Collect_Plates_Into_Dishwasher.general-master-en-202608
General · Master · English · 2026-08
English pretraining text, assembled from three public sources, cleaned with one
character-level cleaner, and filtered for repetition.
109,337,531 documents and 468,064,046,462 characters.
Composition
Config
Documents
Characters
What it is
fineweb-edu-dedup
65,010,430
297,544,916,118
Web text an educational classifier kept
cosmopedia-v2
38,591,146
144,011,993,012
Synthetic prose from a seeded generator… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-master-en-202608.G1_WBT_Brainco_Pickup_Pillow
Data Structure
Observations
observation.state.ee_state (12)
End-effector states of the robot.
Computed via forward kinematics (FK) from the root link to the left and right end-effectors.
Includes the contribution of the waist.
Represented as concatenated poses of both end-effectors.
observation.state.hand_state (12 or 2)
Finger states for both hands. The dimensionality depends on the hand type.
Inspire Hand (range: 0.0… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_WBT_Brainco_Pickup_Pillow.Brain-Stroke-DiagnosisG1_WBT_Brainco_Make_The_Bed
Data Structure
Observations
observation.state.ee_state (12)
End-effector states of the robot.
Computed via forward kinematics (FK) from the root link to the left and right end-effectors.
Includes the contribution of the waist.
Represented as concatenated poses of both end-effectors.
observation.state.hand_state (12 or 2)
Finger states for both hands. The dimensionality depends on the hand type.
Inspire Hand (range: 0.0… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_WBT_Brainco_Make_The_Bed.G1_WBT_Brainco_Pick_Up_Medicine_v0
Data Structure
Observations
observation.state.ee_state (12)
End-effector states of the robot.
Computed via forward kinematics (FK) from the root link to the left and right end-effectors.
Includes the contribution of the waist.
Represented as concatenated poses of both end-effectors.
observation.state.hand_state (12 or 2)
Finger states for both hands. The dimensionality depends on the hand type.
Inspire Hand (range: 0.0… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_WBT_Brainco_Pick_Up_Medicine_v0.Brainly_datasetbrain-teasersgeneral-web-it-202608
General · Web · Italian · 2026-08
Italian pretraining text, built from the Italian portion of EPFL's FineWeb2-HQ, which is the
high quality slice of FineWeb-2. Every document passes one character-level cleaner and a
repetition filter.
21,065,052 documents and 66,158,573,443 characters of Italian prose.
Contents
Config
Documents
Characters
Upstream
fineweb2-hq-ita_Latn
21,065,052
66,158,573,443
epfml/FineWeb2-HQ, ita_Latn
The character count is… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-it-202608.general-web-fr-202608
General · Web · French · 2026-08
French pretraining text, built from the French portion of EPFL's FineWeb2-HQ, which is the
high quality slice of FineWeb-2. Every document passes one character-level cleaner and a
repetition filter.
31,999,309 documents and 118,346,333,763 characters of French prose.
Contents
Config
Documents
Characters
Upstream
fineweb2-hq-fra_Latn
31,999,309
118,346,333,763
epfml/FineWeb2-HQ, fra_Latn
The character count is exact.… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-fr-202608.G1_Brainco_PickApple_DatasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Brainco",
"total_episodes": 200,
"total_frames": 153668,
"total_tasks": 1,
"total_videos": 800,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Brainco_PickApple_Dataset.G1_Brainco_GraspRubiksCube_DatasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Brainco",
"total_episodes": 197,
"total_frames": 220788,
"total_tasks": 1,
"total_videos": 788,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:197"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Brainco_GraspRubiksCube_Dataset.G1_Brainco_GraspOreo_DatasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Brainco",
"total_episodes": 201,
"total_frames": 234959,
"total_tasks": 1,
"total_videos": 804,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:201"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Brainco_GraspOreo_Dataset.nejm-brain-to-text-sonified-istft
NEJM Brain-to-Text Sonified (iSTFT)
Pre-shuffled dataset (seed: 42) at 16kHz, 0-8000Hz range.
Sharded into 1000 files per shard for efficient loading.
Usage
from datasets import load_dataset
ds = load_dataset("ljcamargo/nejm-brain-to-text-sonified-istft")
G1_Brainco_PickDoll_DatasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Brainco",
"total_episodes": 200,
"total_frames": 313401,
"total_tasks": 1,
"total_videos": 800,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Brainco_PickDoll_Dataset.G1_Brainco_PickCharger_DatasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Brainco",
"total_episodes": 200,
"total_frames": 216532,
"total_tasks": 1,
"total_videos": 800,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Brainco_PickCharger_Dataset.brain-hackathon-2023-embed-dataNeuronSpark-Pretrain-v3
NeuronSpark-Pretrain-v3
Bilingual pretraining corpus for NeuronSpark v3, a bio-inspired Spiking Neural
Network language model with selective PLIF neurons and dynamic per-token compute
budget (PonderNet-v3).
Composition
Metric
Value
Total documents
18.2 M
Estimated tokens
~20 B
Format
37 Parquet shards (~1 GB each, zstd)
Schema
text: string, source: string
Languages
EN 55.6%, ZH 28.1%, code 16.3%
Deduplication
All source sampling is weighted so each… See the full description on the dataset page: https://huggingface.co/datasets/Brain2nd/NeuronSpark-Pretrain-v3.multi-modal-derived-brain-network
PPMI Connectivity Graphs — HF Staging (Derivatives)
This dataset ships ready-to-use functional brain connectivity graphs derived from the PPMI cohort in a BIDS-ish derivatives layout. For each subject and parcellation, we include:
ROI time-series (*_desc-timeseries_parc-<name>.mat)
Pearson correlation connectivity matrix (*_desc-correlation_matrix_parc-<name>.mat)
JSON sidecars with summary fields (nodes, measure, symmetric/weighted flags)
Contents
data/… See the full description on the dataset page: https://huggingface.co/datasets/pakkinlau/multi-modal-derived-brain-network.G1_Brainco_PickDrink_DatasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Brainco",
"total_episodes": 201,
"total_frames": 174483,
"total_tasks": 1,
"total_videos": 804,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:201"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Brainco_PickDrink_Dataset.Brain-IT_Results
🧠 Brain-IT Results Dataset
This dataset contains the official reconstructed images and corresponding tensor files (.pt) produced by the Brain-IT model, as presented in the paper:
Brain-IT: Image Reconstruction from fMRI via Brain-Interaction TransformerRoman Beliy, Amit Zalcher, Jonathan Kogman, Navve Wasserman, Michal Irani
🔗 Project Page: https://AmitZalcher.github.io/Brain-IT/
🧩 Splits
Split name pattern
Description
ses40_subi
Full-training results… See the full description on the dataset page: https://huggingface.co/datasets/Amitz244/Brain-IT_Results.G1_Brainco_PickToothpaste_DatasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Brainco",
"total_episodes": 193,
"total_frames": 143863,
"total_tasks": 1,
"total_videos": 772,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:193"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Brainco_PickToothpaste_Dataset.G1_Brainco_PickTissues_DatasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Brainco",
"total_episodes": 206,
"total_frames": 198622,
"total_tasks": 1,
"total_videos": 824,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:206"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Brainco_PickTissues_Dataset.NeuronSpark-V1
NeuronSpark-V1 Pretraining Dataset
Bilingual (English + Chinese) pretraining corpus for NeuronSpark, a bio-inspired Spiking Neural Network language model.
Dataset Summary
Metric
Value
Total documents
17,174,734
Estimated tokens
~14.5B
Languages
English (55%), Chinese (42%), Bilingual Math (3%)
Format
Parquet (35 shards, ~39 GB)
Columns
text (string), source (string)
Sources & Composition
Source
Documents
Ratio
Est. Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Brain2nd/NeuronSpark-V1.brain-to-text-25-text-cleaned
An Accurate and Rapidly Calibrating Speech Neuroprosthesis
The New England Journal of Medicine (2024)
Nicholas S. Card, Maitreyee Wairagkar, Carrina Iacobacci,
Xianda Hou, Tyler Singer-Clark, Francis R. Willett,
Erin M. Kunz, Chaofei Fan, Maryam Vahdati Nia,
Darrel R. Deo, Aparna Srinivasan, Eun Young Choi,
Matthew F. Glasser, Leigh R. Hochberg,
Jaimie M. Henderson, Kiarash Shahlaie,
Sergey D. Stavisky*, and David M. Brandman*.
Text labels are represented as ASCII.
Phoneme labels… See the full description on the dataset page: https://huggingface.co/datasets/riverjiang/brain-to-text-25-text-cleaned.medtrinity_brain_30k_hftrain_valid_split_pmc_neuroscience_2002-2022_filtered_subsetData from PubMed for abstracts and PubMed Central Open Access Subset (PMC OAS) for full-text articles using the Entrez Programming Utilities (E-utilities) API
and the pubget Python package, respectively. The data span publication dates from 2002 to 2022. For science general journals, a keyword filter of ``Neuroscience" was applied (all sourced journals are below).
Data extraction efforts yielded 332,807 abstracts and 123,085 full-text articles, totaling 1.3 billion tokens.
Figures and tables… See the full description on the dataset page: https://huggingface.co/datasets/BrainGPT/train_valid_split_pmc_neuroscience_2002-2022_filtered_subset.
