datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mega-Brain-Distill
Mega-Brain-Distill
Curated merge of the top 10% highest-scoring examples from
584 community-uploaded LLM distillation/reasoning-trace datasets
on the Hub (Fable-5, Opus, GLM, Kimi, DeepSeek, GPT, MiniMax, Qwen traces,
etc.), deduplicated within and across all of them — many of these source
repos are the same underlying dump re-uploaded by different users.
Auto-generated by run.py — do not hand-edit, it will be overwritten on
the next run. Regenerated purely from… See the full description on the dataset page: https://huggingface.co/datasets/ShinMK3/Mega-Brain-Distill.general-master-en-202608
General · Master · English · 2026-08
English pretraining text, assembled from three public sources, cleaned with one
character-level cleaner, and filtered for repetition.
109,337,531 documents and 468,064,046,462 characters.
Composition
Config
Documents
Characters
What it is
fineweb-edu-dedup
65,010,430
297,544,916,118
Web text an educational classifier kept
cosmopedia-v2
38,591,146
144,011,993,012
Synthetic prose from a seeded generator… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-master-en-202608.general-web-it-202608
General · Web · Italian · 2026-08
Italian pretraining text, built from the Italian portion of EPFL's FineWeb2-HQ, which is the
high quality slice of FineWeb-2. Every document passes one character-level cleaner and a
repetition filter.
21,065,052 documents and 66,158,573,443 characters of Italian prose.
Contents
Config
Documents
Characters
Upstream
fineweb2-hq-ita_Latn
21,065,052
66,158,573,443
epfml/FineWeb2-HQ, ita_Latn
The character count is… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-it-202608.general-web-fr-202608
General · Web · French · 2026-08
French pretraining text, built from the French portion of EPFL's FineWeb2-HQ, which is the
high quality slice of FineWeb-2. Every document passes one character-level cleaner and a
repetition filter.
31,999,309 documents and 118,346,333,763 characters of French prose.
Contents
Config
Documents
Characters
Upstream
fineweb2-hq-fra_Latn
31,999,309
118,346,333,763
epfml/FineWeb2-HQ, fra_Latn
The character count is exact.… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-fr-202608.NeuronSpark-Pretrain-v3
NeuronSpark-Pretrain-v3
Bilingual pretraining corpus for NeuronSpark v3, a bio-inspired Spiking Neural
Network language model with selective PLIF neurons and dynamic per-token compute
budget (PonderNet-v3).
Composition
Metric
Value
Total documents
18.2 M
Estimated tokens
~20 B
Format
37 Parquet shards (~1 GB each, zstd)
Schema
text: string, source: string
Languages
EN 55.6%, ZH 28.1%, code 16.3%
Deduplication
All source sampling is weighted so each… See the full description on the dataset page: https://huggingface.co/datasets/Brain2nd/NeuronSpark-Pretrain-v3.NeuronSpark-V1
NeuronSpark-V1 Pretraining Dataset
Bilingual (English + Chinese) pretraining corpus for NeuronSpark, a bio-inspired Spiking Neural Network language model.
Dataset Summary
Metric
Value
Total documents
17,174,734
Estimated tokens
~14.5B
Languages
English (55%), Chinese (42%), Bilingual Math (3%)
Format
Parquet (35 shards, ~39 GB)
Columns
text (string), source (string)
Sources & Composition
Source
Documents
Ratio
Est. Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Brain2nd/NeuronSpark-V1.TopoSense-Bench
TopoSense-Bench: A Campus-Scale Benchmark for Semantic-Spatial Sensor Scheduling
TopoSense-Bench is a large-scale, rigorous benchmark designed to evaluate Large Language Models (LLMs) and agents on the Semantic-Spatial Sensor Scheduling (S³) problem. It features a realistic digital twin of a university campus equipped with 2,510 cameras and contains 5,250 natural language queries grounded in physical topology.
This dataset is the official benchmark for the ACM MobiCom 2026 paper:… See the full description on the dataset page: https://huggingface.co/datasets/IoT-Brain/TopoSense-Bench.brain-memory
🧠 NIFTY AI Agent: Memory OS Cloud Snapshot
Cloud backup repository for the NIFTY 50 Autonomous AI Agent Memory OS.
• Repository: nagarhimanshu37/brain-memory• Total Stored Records: 231• Last Synchronized: 2026-09-24 12:55:53 UTC
📊 Partition Statistics
Partition
Records
Description
conversation_memory
78
Multi-turn trader dialogues & intent logs
episodic_memory
50
Trading day episodes (facts vs interpretations)
experience_memory
50
Crystallized… See the full description on the dataset page: https://huggingface.co/datasets/nagarhimanshu37/brain-memory.generate-narrate-tinystories-pretrain
Narrative · TinyStories · Pretraining (Cleaned)
Microsoft's TinyStories V2, cleaned and stored as parquet. 2,745,100 stories, 441 million
words, one story per row with provenance on every record.
Composition
Config
Records
%
Source
all
2,745,100
100.00
the single config (default)
gpt-4
2,745,100
100.00
TinyStoriesV2-GPT4-train
TinyStories V2 holds samples generated by GPT-3.5 and samples generated by GPT-4. Only
the GPT-4 samples are here… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/generate-narrate-tinystories-pretrain.code-training-il
Code-Training-IL
A 40,330-example instruction-tuning dataset for code: 20K Python (NVIDIA OpenCodeInstruct, test-filtered) + 20K TypeScript + 330 hand-written bilingual identity examples.
Overview
code-training-il is a curated, filtered instruction-tuning corpus for training small coding assistants. It is the dataset used to fine-tune code-il-E4B, a 4B on-device model.
The dataset was designed around a thesis: less data, better filtered, beats more data. The… See the full description on the dataset page: https://huggingface.co/datasets/BrainboxAI/code-training-il.medical-training-il
Medical-Training-IL
A bilingual (Hebrew / English) medical instruction-tuning corpus — curated for training small, on-device medical models for Israeli residents preparing for Stage A exams.
Overview
medical-training-il is a curated, bilingual medical instruction-tuning dataset designed to fine-tune language models for Israeli clinical reasoning. It combines high-quality English medical QA (USMLE-style, basic sciences, research-grounded) with ~5,000 Hebrew-native… See the full description on the dataset page: https://huggingface.co/datasets/BrainboxAI/medical-training-il.MedCortex-v1
MedCortex — Bilingual Medical Reasoning + Consultation Corpus
86,006 provenance-tracked, decontaminated medical examples in one uniform schema, fusing two
complementary strengths without flattening either:
task_type
Rows
What it is
Why it is here
reasoning
44,736
Verified chain-of-thought (KG-grounded + verified CoT), English
Drives exam-style benchmark reasoning — the source of the MedReason paper's measured gains
consultation
41,270
Bilingual (EN/FR) clinical Q&A… See the full description on the dataset page: https://huggingface.co/datasets/BrainHealthAI/MedCortex-v1.BrainScore-Code-EnglishDataset containing clean code and the TinyStories dataset
brainly
brainly.co.id dataset
Data Structure
The keys in each JSONL object include:
"id": An integer value representing the page of task from url (e.g. brainly.co.id/tugas/117).
"subject": A string indicating the subject of the question (e.g., "Fisika", "Matematika", "Sejarah").
"author": A string representing the author of the question.
"instruction": A string providing the instruction or prompt for the question.
"answerer_1", "answer_2": Strings representing the answerers for… See the full description on the dataset page: https://huggingface.co/datasets/lesserfield/brainly.swebench-django
SWE-bench Django (30-task subset)
Thirty real bug-fix tasks from the Django project, derived from SWE-bench (Jimenez, Yang, et al., ICLR 2024). Each task is a merged pull request rewound to its buggy commit: an agent gets only the issue text, must locate and fix the bug in the codebase, and the fix is checked against the PR's held-out test.
This subset powers a Braintrust eval on behavior-vs-output scoring — whether a coding agent obeys a "locate code via vector search only"… See the full description on the dataset page: https://huggingface.co/datasets/BraintrustDataDev/swebench-django.reason-qa-biology-finetune-preview
Reasoning · Biology · Finetuning · Preview (Synthetic)
A public, single-generator preview of a larger private biology reasoning corpus.
This dataset has been created with gpt-oss-20b output and uses a simplified three-field format.
The full set spans many generator models, two reasoning styles (linear and
branching), and a richer schema (metadata, instruction, thinking, reasoning, answer).
Synthetic question-reasoning-answer data for domain finetuning on biology and
biochemistry… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/reason-qa-biology-finetune-preview.uv-brain-s03_custom_with_rehearsal_v2BrainMedCoT
BrainMedCoT — Trilingual Medical Chain-of-Thought Dataset
BrainMedCoT is a trilingual (French / English / Arabic Darija) medical Q&A dataset enriched
with structured chain-of-thought reasoning (<think> block), grounded in real biomedical
sources (PubMed / RxNorm / DailyMed / MedlinePlus). It is the CoT fine-tuning stage of the
HELIX-FT medical-LLM curriculum (SFT → SASR/GRPO → CoT).
Stat
Value
Total examples
3458
Splits (train / val / test)
2768 / 345 / 345
Source… See the full description on the dataset page: https://huggingface.co/datasets/BrainHealthAI/BrainMedCoT.brainstorming-ideation-sft-100k
Brainstorming and Ideation SFT (100K)
100,000 ShareGPT conversations demonstrating structured, high-quality brainstorming and ideation across 22 professional domains. Each example takes a realistic context and constraint, then generates specific, actionable, well-reasoned ideas — not generic advice dressed as creativity.
Motivation
Brainstorming and ideation is one of the highest-value use cases for AI assistants, and one where models routinely underperform:… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/brainstorming-ideation-sft-100k.uv-brain-s03_custom_with_rehearsalPashto-Brain-Extraction-Dataset
🧠 Pashto Brain Extraction Dataset
A small experimental Pashto reasoning dataset designed to extract and preserve useful model reasoning/planning while discarding the final answer.
Keep the brain 🧠 — throw away the mouth 🗣️
🎯 Purpose
A language model may understand a question and produce useful reasoning while still generating poor, unnatural, or grammatically incorrect Pashto in its final answer.
Instead of throwing away the entire generation, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Brain-Extraction-Dataset.brainblast-verified-footgun-corpus
Brainblast — Verified SDK Footgun Corpus (free sample)
The only code-training data that ships with a machine-checkable proof. Each
record is a real insecure→fixed code footgun with a replayable RED→GREEN
receipt: a deterministic checker fails the insecure version and passes the fixed
one. You don't trust the labels — you replay the proof.
This repo is a free 40-record sample (receipt-only tier). The full corpus is
4,183 proven records across 154 SDKs and 9 vulnerability classes… See the full description on the dataset page: https://huggingface.co/datasets/dsb117/brainblast-verified-footgun-corpus.ptv3-bericht-lora-de-300
ptv3-bericht-lora-de-300
Synthetic German dataset for fine-tuning LLMs to generate structured psychotherapy reports (PTV-3 / Bericht an den Gutachter) from therapy session transcripts.
Overview
Property
Value
Samples
311 (280 train / 31 val)
Language
German
Format
ChatML JSONL (system / user / assistant)
Teacher model
Qwen2.5-27B (local)
Generation
Two-stage: seed → session transcript → PTV-3 JSON report
Schema
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/speed-brain-ai/ptv3-bericht-lora-de-300.MedQADataEnglishSaad
Health QA English — Medical Question Answering Dataset
Dataset Description
A curated dataset of 13,812 medical question-answer pairs sourced from real patient-doctor consultations. Each entry contains a patient's clinical scenario, a focused medical question, and a doctor's professional response, enriched with named medical entities (symptoms, diseases, medications, tests).
Key Features
13,812 high-quality entries across 15 medical specialties
Structured… See the full description on the dataset page: https://huggingface.co/datasets/BrainHealthAI/MedQADataEnglishSaad.legal-training-il
Legal-Training-IL
A 17,613-example bilingual instruction-tuning corpus for Israeli legal reasoning — covering rulings, statutes, citizen-rights pages, and contract clauses.
Overview
legal-training-il is a curated, bilingual (Hebrew / English) instruction-tuning dataset designed to adapt general-purpose language models to Israeli legal work. It was built to train law-il-E2B, a 2B-parameter on-device legal assistant.
The dataset is not a scraped dump. Every example… See the full description on the dataset page: https://huggingface.co/datasets/BrainboxAI/legal-training-il.LLM_BRAIn_datasetLLM_BRAIn: AI-driven Fast Generation of Robot Behaviour Tree based on Large Language Model
Original paper preprint: https://arxiv.org/abs/2305.19352
This paper introduces a pioneering methodology in autonomous robot control, denoted as LLM-BRAIn, enabling the generation of adaptive behaviors in robots in response to operator commands, while simultaneously considering a multitude of potential future events. LLM-BRAIn is a transformer-based Large Language Model (LLM) fine-tuned from the… See the full description on the dataset page: https://huggingface.co/datasets/ArtemLykov/LLM_BRAIn_dataset.xio-compliance-brain-triad-prompts
XIO Compliance Brain — Triad Reviewer Prompts
Reusable system prompts for running a multi-voice compliance debate against the same matter — the heart of XIO Compliance Brain's "Triad Review Engine" pattern.
This dataset extracts the production prompts from the open-source compliance-AI hackathon branch so others can replicate the Triad pattern (three reviewer voices + synthesis + optional Round 2) on any LLM that follows OpenAI-compatible chat APIs.
What's in this… See the full description on the dataset page: https://huggingface.co/datasets/slavazeph/xio-compliance-brain-triad-prompts.r-braincels-instructThis is a chat formatted dataset of r/braincels posts. Each row in the JSONL contains one user message, which is a submission to r/braincels, and one assistant message, which is a reply to that submission.
Built from this dataset
brainstorming-thinking
brainstorming-thinking
Using Qwen3-14b to synthetically generate the reasoning traces and answers for Explore_Instruct_Brainstorming_10k
Suitable for LLM post-training, especially RL.
MedUnified
MedUnified — Trilingual Medical SFT Mix (HELIX-FT v2)
MedUnified is the continued-SFT data mix for the second iteration of the HELIX-FT
medical LLM (BrainHealthAI/MedQA-Llama3.1-8B-HELIX-v2). It unifies five complementary
medical sources — real and synthetic, English / French / Moroccan Darija — into one
single-language-per-row training corpus, decontaminated against the standard medical eval
benchmarks.
Stat
Value
Total rows
23 000 (22 505 train + 495 validation)… See the full description on the dataset page: https://huggingface.co/datasets/BrainHealthAI/MedUnified.
