CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01semran1 /dclm-stem-filteredtext1M<n<10M0 likes3.3k downloads1y agoHugging Face02EssentialAI /eai-taxonomy-stem-w-dclm 🔬 EAI-Taxonomy STEM w/ DCLM 🏆 Website | 🖥️ Code | 📖 Paper A high-quality STEM dataset curated from web data using taxonomy-based filtering, containing 1742 billion tokens of science, technology, engineering, and mathematics content. 🎯 Dataset Overview This dataset is part of the Essential-Web project, which introduces a new paradigm for dataset curation using expressive metadata and simple semantic filters. Unlike traditional STEM datasets that require… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/eai-taxonomy-stem-w-dclm.6 likes2.7k downloads1y agoHugging Face03TigreGotico /arabic-stem-lexicon Arabic Diacritized-Stem Lexicon An undiacritized Arabic surface form → its most frequent diacritized stem. Standard Arabic writes no short vowels, so anything that has to pronounce Arabic must first put them back. A neural diacritizer does that well on rare words, where inference is the only thing there is. On common words it is the wrong tool: which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.tabulartext-to-speech100K<n<1M0 likes2.2k downloads2mo agoHugging Face04TIGER-Lab /MMLU-STEMThis contains a subset of STEM subjects defined in MMLU by the original paper. The included subjects are 'abstract_algebra', 'anatomy', 'astronomy', 'college_biology', 'college_chemistry', 'college_computer_science', 'college_mathematics', 'college_physics', 'computer_security', 'conceptual_physics', 'electrical_engineering', 'elementary_mathematics', 'high_school_biology', 'high_school_chemistry', 'high_school_computer_science', 'high_school_mathematics', 'high_school_physics'… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/MMLU-STEM.text1K<n<10K19 likes2.1k downloads2y agoHugging Face05Isaac105 /melodix-stemsaudio1M<n<10M0 likes1.7k downloads1d agoHugging Face06stemdataset /STEM STEM Dataset 📃 [Paper] • 💻 [Github] • 🤗 [Dataset] • 🏆 [Leaderboard] • 📽 [Slides] • 📋 [Poster] This dataset is proposed in the ICLR 2024 paper: Measuring Vision-Language STEM Skills of Neural Models. We introduce a new challenge to test the STEM skills of neural models. The problems in the real world often require solutions, combining knowledge from STEM (science, technology, engineering, and math). Unlike existing datasets, our dataset requires the understanding of… See the full description on the dataset page: https://huggingface.co/datasets/stemdataset/STEM.text1M<n<10M6 likes1.3k downloads2y agoHugging Face07EssentialAI /eai-taxonomy-stem-w-dclm-100b-sample 🔬 EAI-Taxonomy STEM w/ DCLM (100B sample) 🏆 Website | 🖥️ Code | 📖 Paper A high-quality STEM dataset curated from web data using taxonomy-based filtering, containing 100 billion tokens of science, technology, engineering, and mathematics content. 🎯 Dataset Overview This dataset is part of the Essential-Web project, which introduces a new paradigm for dataset curation using expressive metadata and simple semantic filters. Unlike traditional STEM datasets that… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/eai-taxonomy-stem-w-dclm-100b-sample.text10M<n<100M5 likes947 downloads1y agoHugging Face08Logics-MLLM /Logics-STEM-SFT-Dataset-Open-1.6M Logics-STEM-SFT-Dataset-2.2M 📰 News [2026.01.05]🔥 Release of our Techinical Report. [2026.01.05]🔥 Release the first version of Logics-STEM-8B-SFT, Logics-STEM-8B-RL, /Logics-STEM-SFT-Dataset-Open-1.6M. Overview What is this dataset? Logics-STEM-SFT-Dataset-2.2M is a curated long Chain-of-Thought (CoT) SFT dataset for STEM reasoning, built on top of high-quality open-source data and enhanced through a rigorous curation and distillation… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-STEM-SFT-Dataset-Open-1.6M.text1M<n<10M33 likes809 downloads8mo agoHugging Face09si-m07 /stemdata STEM Dataset 📃 [Paper] • 💻 [Github] • 🤗 [Dataset] • 🏆 [Leaderboard] • 📽 [Slides] • 📋 [Poster] This dataset is proposed in the ICLR 2024 paper: Measuring Vision-Language STEM Skills of Neural Models. We introduce a new challenge to test the STEM skills of neural models. The problems in the real world often require solutions, combining knowledge from STEM (science, technology, engineering, and math). Unlike existing datasets, our dataset requires the understanding of… See the full description on the dataset page: https://huggingface.co/datasets/si-m07/stemdata.text1M<n<10M0 likes782 downloads2mo agoHugging Face10gary23ai /STEM2Crystal-Bench STEM2Crystal-Bench STEM2Crystal-Bench is the benchmark for the paper "From Noisy STEM to Crystal Structure: Evidence-Structure CoDiffusion under Composition Constraints" (Chen & You, KDD 2026, Oral), which introduces STEM2Crystal CoDiffusion (SCCD). It evaluates methods that reconstruct a crystal structure from a noisy STEM image when the composition is known. The release has a large synthetic set with controlled noise and a small set of real STEM images, with ground-truth CIFs… See the full description on the dataset page: https://huggingface.co/datasets/gary23ai/STEM2Crystal-Bench.imageimage-to-text1K<n<10K1 likes761 downloads3mo agoHugging Face11kaizen9 /stem-corpustext10M<n<100M0 likes700 downloads1y agoHugging Face12lfaviate /China-K12-STEM-10K-CoT-Reasoning K12-STEM-CoT-Chinese 1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams. The largest structured Chinese math/physics/chemistry reasoning dataset. This is a curated sample (10,000 problems) of the full 1.54M dataset available via API. Full Dataset Access Access the full 1,540,000+ problems via API → This Sample Full API Total problems 10,025 1,540,000+ With CoT solutions 10,025 1,490,000+ With diagrams 6,093 740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.tabularquestion-answering10K<n<100K3 likes647 downloads7mo agoHugging Face13sreearravind /AI-Research-Evaluation-Repository-STEM AI-STEM-Research-Eval-Dataset Overview This dataset contains AI-generated scientific reports across STEM domains, accompanied by structured metadata, prompt documentation, reference validation, and hallucination annotations. It is designed as an open research resource to study the capabilities, limitations, and reliability of large language models (LLMs) in generating scientific content. The dataset enables systematic analysis of how AI systems perform in… See the full description on the dataset page: https://huggingface.co/datasets/sreearravind/AI-Research-Evaluation-Repository-STEM.text-generationn<1K1 likes498 downloads3mo agoHugging Face14galaxyMindAiLabs /stem-reasoning-complex STEM-Reasoning-Complex: High-Fidelity Scientific CoT Dataset 1. Dataset Summary STEM-Reasoning-Complex is a curated collection of 118.255 high-quality samples designed for Supervised Fine-Tuning (SFT) and alignment of Large Language Models. The dataset focuses on four core disciplines: Biology, Mathematics, Physics, and Chemistry. Unlike standard QA datasets, each entry provides a structured Chain-of-Thought (CoT) reasoning process, enabling models to learn… See the full description on the dataset page: https://huggingface.co/datasets/galaxyMindAiLabs/stem-reasoning-complex.texttext-generation100K<n<1M79 likes457 downloads6mo agoHugging Face15qingyangzhang /Natural-Reasoning-STEM-25Ktabular10K<n<100K0 likes440 downloads1y agoHugging Face16mhurhangee /pat-stem-pretraintext1M<n<10M0 likes401 downloads1y agoHugging Face17scitomo /pfnc-gst-haadf-stem-eds-tomography-b2-d3 PFNC GST HAADF-STEM/EDS Tomography — B2 VIRGIN and D3 SET Summary This release contains two limited-angle HAADF-STEM/EDS tomography acquisitions of cross-sectional Ge-Sb-Te phase-change-memory devices acquired at the Platform for Nanocharacterisation (PFNC), CEA Grenoble, on a Thermo Fisher Scientific Titan Themis operated at 200 kV with a four-detector Super-X EDS system. B2 and D3 are internal CEA microscope sample identifiers. In this release, B2 is the VIRGIN… See the full description on the dataset page: https://huggingface.co/datasets/scitomo/pfnc-gst-haadf-stem-eds-tomography-b2-d3.1K<n<10K0 likes388 downloads2d agoHugging Face18yaotianvector /STEM2Mat AutoMat Benchmark: STEM Image to Crystal Structure The AutoMat Benchmark is a multimodal dataset designed to evaluate deep‑learning systems for iDPC-STEM‑based crystal‑structure reconstruction and property prediction. Code: https://github.com/yyt-2378/AutoMat 📁 Dataset Structure The dataset is organized into three tiers of increasing difficulty: benchmark/ ├── tier1/ │ ├── img/ # STEM images (e.g., PNG, TIFF) │ ├── label/ # Atomic position labels… See the full description on the dataset page: https://huggingface.co/datasets/yaotianvector/STEM2Mat.imageimage-to-3d1K<n<10K2 likes373 downloads1y agoHugging Face19shreyansh12183 /shreyansh-1B-SLM-pretrain-stem-english 📚 Vigyan Pretrain Corpus: 13GB Web-Scale Scientific & Technical Text The Vigyan Pretrain Corpus is a web-scale, curated raw text pre-training dataset comprising 13.14 GB of high-density STEM literature, textbooks, open-access research papers, and technical documentations. 🔬 Dataset Overview Designed specifically for pre-training and continuous pre-training (CPT) of Small Language Models (SLMs) in the 1B–3B parameter regime: High Information Density: Filtered to… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-1B-SLM-pretrain-stem-english.texttext-generation10M<n<100M0 likes346 downloads18h agoHugging Face20semran1 /ultrafine-stem-part-2text100K<n<1M0 likes336 downloads1y agoHugging Face21Tushe /tushe-grade-school-stem Tushe Community Grade School STEM Open dataset of grade-school STEM (Science, Technology, Engineering, Mathematics) textbooks, curated for Tushe Community and aligned with curriculum use (e.g. CAPS-aligned content). Data Fields (per book JSON) Field Type Description source_file string Original .txt filename title string Derived book title (e.g. "Grade 8A Mathematics") table_of_contents list [{ "section_id", "title" }, ...] front_matter string Intro… See the full description on the dataset page: https://huggingface.co/datasets/Tushe/tushe-grade-school-stem.text-generation10K<n<100K0 likes294 downloads8mo agoHugging Face22AxiomSetLabs /stem-scientific-code-sample AxiomSet Labs STEM Scientific-Code Sample A 30-task sample of STEM reasoning and scientific-code problems across five domains. Domains Biology: 6 tasks Chemistry: 6 tasks Materials Science: 6 tasks Mathematics: 6 tasks Physics: 6 tasks Each task contains two subproblems and one main problem, including prompts, scientific background, testing templates, and reference solutions. Files data/sample.jsonl — one task per line; used by the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/AxiomSetLabs/stem-scientific-code-sample.texttext-generationn<1K0 likes290 downloads1mo agoHugging Face23RedMod /STEM_sfttext10M<n<100M0 likes280 downloads5mo agoHugging Face24Logics-MLLM /Logics-STEM-SFT-Dataset-Open-5.3Mtext1M<n<10M4 likes279 downloads8mo agoHugging Face25Harmonic-Frontier-Audio /Celtic_Stems_Reference_Sessions_Preview Harmonic Frontier Audio – Celtic Constellation Reference Sessions (Preview, v0.9) A high-fidelity music-production dataset designed to connect isolated source performances, production processing, arrangement context, and finished musical outcomes. Celtic Constellation Reference Sessions (Preview), created by Harmonic Frontier Audio, introduces the Reference Sessions product vertical through a compact proof-of-concept built around purpose-recorded Celtic ensemble material.… See the full description on the dataset page: https://huggingface.co/datasets/Harmonic-Frontier-Audio/Celtic_Stems_Reference_Sessions_Preview.audioothern<1K2 likes260 downloads25d agoHugging Face26semran1 /ultrafine-stem-part-1text100K<n<1M0 likes250 downloads1y agoHugging Face27Jackrong /gpt-oss-120b-reasoning-STEM-5K GPT-OSS-120B-Distilled-Reasoning-STEM Dataset 1) Dataset Overview Data Source Model: gpt-oss-120b-high Task Type: STEM Reasoning and Problem Solving (Science, Technology, Engineering & Mathematics) Data Format: `JSON Lines Fields: generator, category, input, CoT_Native——reasoning, answer (Consistent with the math dataset, splitting the original 'output' into 'reasoning' and 'answer' for COT/SFT scenarios.) 2) Design Goals (Motivation) This dataset targets… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120b-reasoning-STEM-5K.textquestion-answering1K<n<10K12 likes244 downloads1y agoHugging Face28aeyxen /stem-diagrams STEM Diagrams 30,325 technical diagrams (block diagrams, schematics, flowcharts, architectures) extracted from arXiv papers across six engineering fields, each with a source attribution and a quality score. Built by an LLM-curated pipeline and used to show that a small frozen-feature classifier can replace the paid LLM labeling gate. Paper: Distilling an LLM Diagram-Curation Pipeline into Local Classifiers (Adnan Abbasi, Thothica, 2026) Code:… See the full description on the dataset page: https://huggingface.co/datasets/aeyxen/stem-diagrams.imageimage-classification10K<n<100K0 likes241 downloads2mo agoHugging Face29stem-content-ai-project /swahili-speech Swahili Speech-to-Text Dataset This dataset contains paired audio and text data for training and evaluating speech-to-text models in Swahili. The audio files have been processed to remove silence, converted to 44.1kHz mono FLAC format, and are paired with corresponding transcriptions. Structure audio_*.flac: Audio files in FLAC format, named by their corresponding text corpus ID. metadata.jsonl: JSON Lines file with metadata for each audio-text pair. Each line is a JSON… See the full description on the dataset page: https://huggingface.co/datasets/stem-content-ai-project/swahili-speech.audiotext-to-speech1K<n<10K0 likes228 downloads1y agoHugging Face30toksuite /toksuite_stem Dataset Card for Tokenization Robustness TokSuite Benchmark (STEM Collection) Dataset Description This dataset is the STEM subset of the TokSuite benchmark, designed to evaluate how tokenizer choice affects model behavior under realistic formatting, notation, and surface-form perturbations in technical text. TokSuite includes specialized benchmarks for mathematics and STEM, with the STEM subset containing 44 canonical technical questions paired with a… See the full description on the dataset page: https://huggingface.co/datasets/toksuite/toksuite_stem.textmultiple-choicen<1K0 likes204 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.