CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ShinMK3 /Mega-Brain-Distill Mega-Brain-Distill Curated merge of the top 10% highest-scoring examples from 584 community-uploaded LLM distillation/reasoning-trace datasets on the Hub (Fable-5, Opus, GLM, Kimi, DeepSeek, GPT, MiniMax, Qwen traces, etc.), deduplicated within and across all of them — many of these source repos are the same underlying dump re-uploaded by different users. Auto-generated by run.py — do not hand-edit, it will be overwritten on the next run. Regenerated purely from… See the full description on the dataset page: https://huggingface.co/datasets/ShinMK3/Mega-Brain-Distill.tabulartext-generation10K<n<100K2 likes5.4k downloads2mo agoHugging Face02Brainquiver /general-master-en-202608 General · Master · English · 2026-08 English pretraining text, assembled from three public sources, cleaned with one character-level cleaner, and filtered for repetition. 109,337,531 documents and 468,064,046,462 characters. Composition Config Documents Characters What it is fineweb-edu-dedup 65,010,430 297,544,916,118 Web text an educational classifier kept cosmopedia-v2 38,591,146 144,011,993,012 Synthetic prose from a seeded generator… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-master-en-202608.tabulartext-generation100M<n<1B1 likes966 downloads25d agoHugging Face03Brainquiver /general-web-it-202608 General · Web · Italian · 2026-08 Italian pretraining text, built from the Italian portion of EPFL's FineWeb2-HQ, which is the high quality slice of FineWeb-2. Every document passes one character-level cleaner and a repetition filter. 21,065,052 documents and 66,158,573,443 characters of Italian prose. Contents Config Documents Characters Upstream fineweb2-hq-ita_Latn 21,065,052 66,158,573,443 epfml/FineWeb2-HQ, ita_Latn The character count is… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-it-202608.tabulartext-generation10M<n<100M0 likes554 downloads25d agoHugging Face04Brainquiver /general-web-fr-202608 General · Web · French · 2026-08 French pretraining text, built from the French portion of EPFL's FineWeb2-HQ, which is the high quality slice of FineWeb-2. Every document passes one character-level cleaner and a repetition filter. 31,999,309 documents and 118,346,333,763 characters of French prose. Contents Config Documents Characters Upstream fineweb2-hq-fra_Latn 31,999,309 118,346,333,763 epfml/FineWeb2-HQ, fra_Latn The character count is exact.… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-fr-202608.tabulartext-generation10M<n<100M0 likes495 downloads25d agoHugging Face05Brain2nd /NeuronSpark-Pretrain-v3 NeuronSpark-Pretrain-v3 Bilingual pretraining corpus for NeuronSpark v3, a bio-inspired Spiking Neural Network language model with selective PLIF neurons and dynamic per-token compute budget (PonderNet-v3). Composition Metric Value Total documents 18.2 M Estimated tokens ~20 B Format 37 Parquet shards (~1 GB each, zstd) Schema text: string, source: string Languages EN 55.6%, ZH 28.1%, code 16.3% Deduplication All source sampling is weighted so each… See the full description on the dataset page: https://huggingface.co/datasets/Brain2nd/NeuronSpark-Pretrain-v3.texttext-generation10M<n<100M0 likes412 downloads5mo agoHugging Face06Brain2nd /NeuronSpark-V1 NeuronSpark-V1 Pretraining Dataset Bilingual (English + Chinese) pretraining corpus for NeuronSpark, a bio-inspired Spiking Neural Network language model. Dataset Summary Metric Value Total documents 17,174,734 Estimated tokens ~14.5B Languages English (55%), Chinese (42%), Bilingual Math (3%) Format Parquet (35 shards, ~39 GB) Columns text (string), source (string) Sources & Composition Source Documents Ratio Est. Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Brain2nd/NeuronSpark-V1.texttext-generation10M<n<100M0 likes300 downloads6mo agoHugging Face07IoT-Brain /TopoSense-Bench TopoSense-Bench: A Campus-Scale Benchmark for Semantic-Spatial Sensor Scheduling TopoSense-Bench is a large-scale, rigorous benchmark designed to evaluate Large Language Models (LLMs) and agents on the Semantic-Spatial Sensor Scheduling (S³) problem. It features a realistic digital twin of a university campus equipped with 2,510 cameras and contains 5,250 natural language queries grounded in physical topology. This dataset is the official benchmark for the ACM MobiCom 2026 paper:… See the full description on the dataset page: https://huggingface.co/datasets/IoT-Brain/TopoSense-Bench.texttext-generation1K<n<10K2 likes280 downloads10mo agoHugging Face08nagarhimanshu37 /brain-memory 🧠 NIFTY AI Agent: Memory OS Cloud Snapshot Cloud backup repository for the NIFTY 50 Autonomous AI Agent Memory OS. • Repository: nagarhimanshu37/brain-memory• Total Stored Records: 231• Last Synchronized: 2026-09-24 12:55:53 UTC 📊 Partition Statistics Partition Records Description conversation_memory 78 Multi-turn trader dialogues & intent logs episodic_memory 50 Trading day episodes (facts vs interpretations) experience_memory 50 Crystallized… See the full description on the dataset page: https://huggingface.co/datasets/nagarhimanshu37/brain-memory.texttext-generationn<1K0 likes172 downloads4h agoHugging Face09Brainquiver /generate-narrate-tinystories-pretrain Narrative · TinyStories · Pretraining (Cleaned) Microsoft's TinyStories V2, cleaned and stored as parquet. 2,745,100 stories, 441 million words, one story per row with provenance on every record. Composition Config Records % Source all 2,745,100 100.00 the single config (default) gpt-4 2,745,100 100.00 TinyStoriesV2-GPT4-train TinyStories V2 holds samples generated by GPT-3.5 and samples generated by GPT-4. Only the GPT-4 samples are here… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/generate-narrate-tinystories-pretrain.texttext-generation1M<n<10M1 likes122 downloads23d agoHugging Face10BrainboxAI /code-training-il Code-Training-IL A 40,330-example instruction-tuning dataset for code: 20K Python (NVIDIA OpenCodeInstruct, test-filtered) + 20K TypeScript + 330 hand-written bilingual identity examples. Overview code-training-il is a curated, filtered instruction-tuning corpus for training small coding assistants. It is the dataset used to fine-tune code-il-E4B, a 4B on-device model. The dataset was designed around a thesis: less data, better filtered, beats more data. The… See the full description on the dataset page: https://huggingface.co/datasets/BrainboxAI/code-training-il.texttext-generation10K<n<100K1 likes95 downloads5mo agoHugging Face11BrainboxAI /medical-training-il Medical-Training-IL A bilingual (Hebrew / English) medical instruction-tuning corpus — curated for training small, on-device medical models for Israeli residents preparing for Stage A exams. Overview medical-training-il is a curated, bilingual medical instruction-tuning dataset designed to fine-tune language models for Israeli clinical reasoning. It combines high-quality English medical QA (USMLE-style, basic sciences, research-grounded) with ~5,000 Hebrew-native… See the full description on the dataset page: https://huggingface.co/datasets/BrainboxAI/medical-training-il.texttext-generation10K<n<100K0 likes84 downloads5mo agoHugging Face12BrainHealthAI /MedCortex-v1 MedCortex — Bilingual Medical Reasoning + Consultation Corpus 86,006 provenance-tracked, decontaminated medical examples in one uniform schema, fusing two complementary strengths without flattening either: task_type Rows What it is Why it is here reasoning 44,736 Verified chain-of-thought (KG-grounded + verified CoT), English Drives exam-style benchmark reasoning — the source of the MedReason paper's measured gains consultation 41,270 Bilingual (EN/FR) clinical Q&A… See the full description on the dataset page: https://huggingface.co/datasets/BrainHealthAI/MedCortex-v1.tabularquestion-answering10K<n<100K1 likes84 downloads2mo agoHugging Face13IParraMartin /BrainScore-Code-EnglishDataset containing clean code and the TinyStories dataset texttext-generation10M<n<100M0 likes66 downloads4mo agoHugging Face14lesserfield /brainly brainly.co.id dataset Data Structure The keys in each JSONL object include: "id": An integer value representing the page of task from url (e.g. brainly.co.id/tugas/117). "subject": A string indicating the subject of the question (e.g., "Fisika", "Matematika", "Sejarah"). "author": A string representing the author of the question. "instruction": A string providing the instruction or prompt for the question. "answerer_1", "answer_2": Strings representing the answerers for… See the full description on the dataset page: https://huggingface.co/datasets/lesserfield/brainly.textquestion-answering1M<n<10M6 likes63 downloads2y agoHugging Face15BraintrustDataDev /swebench-django SWE-bench Django (30-task subset) Thirty real bug-fix tasks from the Django project, derived from SWE-bench (Jimenez, Yang, et al., ICLR 2024). Each task is a merged pull request rewound to its buggy commit: an agent gets only the issue text, must locate and fix the bug in the codebase, and the fix is checked against the PR's held-out test. This subset powers a Braintrust eval on behavior-vs-output scoring — whether a coding agent obeys a "locate code via vector search only"… See the full description on the dataset page: https://huggingface.co/datasets/BraintrustDataDev/swebench-django.texttext-generationn<1K0 likes62 downloads1mo agoHugging Face16Brainquiver /reason-qa-biology-finetune-preview Reasoning · Biology · Finetuning · Preview (Synthetic) A public, single-generator preview of a larger private biology reasoning corpus. This dataset has been created with gpt-oss-20b output and uses a simplified three-field format. The full set spans many generator models, two reasoning styles (linear and branching), and a richer schema (metadata, instruction, thinking, reasoning, answer). Synthetic question-reasoning-answer data for domain finetuning on biology and biochemistry… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/reason-qa-biology-finetune-preview.texttext-generation10K<n<100K0 likes59 downloads3mo agoHugging Face17codex-master /uv-brain-s03_custom_with_rehearsal_v2texttext-generation1K<n<10K0 likes58 downloads27d agoHugging Face18BrainHealthAI /BrainMedCoT BrainMedCoT — Trilingual Medical Chain-of-Thought Dataset BrainMedCoT is a trilingual (French / English / Arabic Darija) medical Q&A dataset enriched with structured chain-of-thought reasoning (<think> block), grounded in real biomedical sources (PubMed / RxNorm / DailyMed / MedlinePlus). It is the CoT fine-tuning stage of the HELIX-FT medical-LLM curriculum (SFT → SASR/GRPO → CoT). Stat Value Total examples 3458 Splits (train / val / test) 2768 / 345 / 345 Source… See the full description on the dataset page: https://huggingface.co/datasets/BrainHealthAI/BrainMedCoT.textquestion-answering1K<n<10K1 likes56 downloads4mo agoHugging Face19stindardlogic /brainstorming-ideation-sft-100k Brainstorming and Ideation SFT (100K) 100,000 ShareGPT conversations demonstrating structured, high-quality brainstorming and ideation across 22 professional domains. Each example takes a realistic context and constraint, then generates specific, actionable, well-reasoned ideas — not generic advice dressed as creativity. Motivation Brainstorming and ideation is one of the highest-value use cases for AI assistants, and one where models routinely underperform:… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/brainstorming-ideation-sft-100k.texttext-generation100K<n<1M0 likes52 downloads2mo agoHugging Face20codex-master /uv-brain-s03_custom_with_rehearsaltexttext-generation1K<n<10K0 likes51 downloads27d agoHugging Face21nassimjp /Pashto-Brain-Extraction-Dataset 🧠 Pashto Brain Extraction Dataset A small experimental Pashto reasoning dataset designed to extract and preserve useful model reasoning/planning while discarding the final answer. Keep the brain 🧠 — throw away the mouth 🗣️ 🎯 Purpose A language model may understand a question and produce useful reasoning while still generating poor, unnatural, or grammatically incorrect Pashto in its final answer. Instead of throwing away the entire generation, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Brain-Extraction-Dataset.texttext-generationn<1K0 likes42 downloads9d agoHugging Face22dsb117 /brainblast-verified-footgun-corpus Brainblast — Verified SDK Footgun Corpus (free sample) The only code-training data that ships with a machine-checkable proof. Each record is a real insecure→fixed code footgun with a replayable RED→GREEN receipt: a deterministic checker fails the insecure version and passes the fixed one. You don't trust the labels — you replay the proof. This repo is a free 40-record sample (receipt-only tier). The full corpus is 4,183 proven records across 154 SDKs and 9 vulnerability classes… See the full description on the dataset page: https://huggingface.co/datasets/dsb117/brainblast-verified-footgun-corpus.tabulartext-generationn<1K0 likes36 downloads2mo agoHugging Face23speed-brain-ai /ptv3-bericht-lora-de-300 ptv3-bericht-lora-de-300 Synthetic German dataset for fine-tuning LLMs to generate structured psychotherapy reports (PTV-3 / Bericht an den Gutachter) from therapy session transcripts. Overview Property Value Samples 311 (280 train / 31 val) Language German Format ChatML JSONL (system / user / assistant) Teacher model Qwen2.5-27B (local) Generation Two-stage: seed → session transcript → PTV-3 JSON report Schema Each sample… See the full description on the dataset page: https://huggingface.co/datasets/speed-brain-ai/ptv3-bericht-lora-de-300.texttext-generationn<1K0 likes35 downloads6mo agoHugging Face24BrainHealthAI /MedQADataEnglishSaad Health QA English — Medical Question Answering Dataset Dataset Description A curated dataset of 13,812 medical question-answer pairs sourced from real patient-doctor consultations. Each entry contains a patient's clinical scenario, a focused medical question, and a doctor's professional response, enriched with named medical entities (symptoms, diseases, medications, tests). Key Features 13,812 high-quality entries across 15 medical specialties Structured… See the full description on the dataset page: https://huggingface.co/datasets/BrainHealthAI/MedQADataEnglishSaad.textquestion-answering10K<n<100K0 likes35 downloads5mo agoHugging Face25BrainboxAI /legal-training-il Legal-Training-IL A 17,613-example bilingual instruction-tuning corpus for Israeli legal reasoning — covering rulings, statutes, citizen-rights pages, and contract clauses. Overview legal-training-il is a curated, bilingual (Hebrew / English) instruction-tuning dataset designed to adapt general-purpose language models to Israeli legal work. It was built to train law-il-E2B, a 2B-parameter on-device legal assistant. The dataset is not a scraped dump. Every example… See the full description on the dataset page: https://huggingface.co/datasets/BrainboxAI/legal-training-il.texttext-generation10K<n<100K2 likes32 downloads5mo agoHugging Face26ArtemLykov /LLM_BRAIn_datasetLLM_BRAIn: AI-driven Fast Generation of Robot Behaviour Tree based on Large Language Model Original paper preprint: https://arxiv.org/abs/2305.19352 This paper introduces a pioneering methodology in autonomous robot control, denoted as LLM-BRAIn, enabling the generation of adaptive behaviors in robots in response to operator commands, while simultaneously considering a multitude of potential future events. LLM-BRAIn is a transformer-based Large Language Model (LLM) fine-tuned from the… See the full description on the dataset page: https://huggingface.co/datasets/ArtemLykov/LLM_BRAIn_dataset.textrobotics1K<n<10K4 likes31 downloads2y agoHugging Face27slavazeph /xio-compliance-brain-triad-prompts XIO Compliance Brain — Triad Reviewer Prompts Reusable system prompts for running a multi-voice compliance debate against the same matter — the heart of XIO Compliance Brain's "Triad Review Engine" pattern. This dataset extracts the production prompts from the open-source compliance-AI hackathon branch so others can replicate the Triad pattern (three reviewer voices + synthesis + optional Round 2) on any LLM that follows OpenAI-compatible chat APIs. What's in this… See the full description on the dataset page: https://huggingface.co/datasets/slavazeph/xio-compliance-brain-triad-prompts.texttext-generationn<1K0 likes31 downloads5mo agoHugging Face28trentmkelly /r-braincels-instructThis is a chat formatted dataset of r/braincels posts. Each row in the JSONL contains one user message, which is a submission to r/braincels, and one assistant message, which is a reply to that submission. Built from this dataset texttext-generation10K<n<100K0 likes27 downloads2y agoHugging Face29chimbiwide /brainstorming-thinking brainstorming-thinking Using Qwen3-14b to synthetically generate the reasoning traces and answers for Explore_Instruct_Brainstorming_10k Suitable for LLM post-training, especially RL. texttext-generation1K<n<10K0 likes25 downloads9mo agoHugging Face30BrainHealthAI /MedUnified MedUnified — Trilingual Medical SFT Mix (HELIX-FT v2) MedUnified is the continued-SFT data mix for the second iteration of the HELIX-FT medical LLM (BrainHealthAI/MedQA-Llama3.1-8B-HELIX-v2). It unifies five complementary medical sources — real and synthetic, English / French / Moroccan Darija — into one single-language-per-row training corpus, decontaminated against the standard medical eval benchmarks. Stat Value Total rows 23 000 (22 505 train + 495 validation)… See the full description on the dataset page: https://huggingface.co/datasets/BrainHealthAI/MedUnified.textquestion-answering10K<n<100K0 likes25 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.