CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TMoC /SYNTH-Swallow-Math-Code-Mix Mixed dataset: SYNTH + SwallowMath-v2 + SwallowCode-v2 This high-signal, all synthetic dataset is a complete shuffled mix of the following four sources: SYNTH ~63.5% SwallowCode-v2 ~15.5% SwallowMath-v2-textbook ~10.5% SwallowMath-v2-qa ~10.0% The motivation to provide this on HF was the need for a convenient, pre-shuffled merge of the highest quality synthetic / augmented datasets for small language model pre-training experiments as of… See the full description on the dataset page: https://huggingface.co/datasets/TMoC/SYNTH-Swallow-Math-Code-Mix.texttext-generation100M<n<1B1 likes728 downloads8mo agoHugging Face02MathLLMs /MathCodeInstruct MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning Paper: https://arxiv.org/pdf/2310.03731.pdf Repo: https://github.com/mathllm/MathCoder Introduction We introduce MathCoder, a series of open-source large language models (LLMs) specifically tailored for general math problem-solving. Base Model: Llama-2 Base Model: Code Llama MathCoder-L-7B MathCoder-CL-7B MathCoder-L-13B MathCoder-CL-34B Training Data The models… See the full description on the dataset page: https://huggingface.co/datasets/MathLLMs/MathCodeInstruct.textquestion-answering10K<n<100K25 likes305 downloads2y agoHugging Face03manifesta /verified-math-code-17k Verified Math & Code, 17,000 rows A math and code instruction dataset where every single row was mechanically checked before it was allowed in. Not filtered by a heuristic, not scored by a model. Checked. Two layers of verification, one per domain: Every math answer was compared against an independent gold answer by exact, numeric and symbolic (SymPy) comparison. If the worked solution did not arrive at the gold answer, the row was dropped.… See the full description on the dataset page: https://huggingface.co/datasets/manifesta/verified-math-code-17k.texttext-generation10K<n<100K0 likes202 downloads1mo agoHugging Face04MathLLMs /MathCodeInstruct-Plus MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning Paper: https://arxiv.org/pdf/2310.03731.pdf Repo: https://github.com/mathllm/MathCoder Introduction We introduce MathCoder, a series of open-source large language models (LLMs) specifically tailored for general math problem-solving. Base Model: Llama-2 Base Model: Code Llama MathCoder-L-7B MathCoder-CL-7B MathCoder-L-13B MathCoder-CL-34B Training Data The models… See the full description on the dataset page: https://huggingface.co/datasets/MathLLMs/MathCodeInstruct-Plus.textquestion-answering10K<n<100K17 likes199 downloads2y agoHugging Face05theblackcat102 /codex-math-qaSolution by codex-davinci-002 for math_qatexttext-generation10K<n<100K31 likes165 downloads4y agoHugging Face06chhao /AM-DeepSeek-R1-Filtered-Math-Code AM DeepSeek R1 Filtered Math and Code This repository publishes reproducible training subsets derived from a-m-team/AM-DeepSeek-R1-Distilled-1.4M at pinned revision 53531c06634904118a2dcd83961918c4d69d1cdf. The source dataset and this derived release use CC BY-NC 4.0. Commercial use is not permitted by that license. Preserve attribution and review the upstream dataset card before use. Contents Config / split Records Bytes SHA-256 math / train 111,657 2… See the full description on the dataset page: https://huggingface.co/datasets/chhao/AM-DeepSeek-R1-Filtered-Math-Code.texttext-generation100K<n<1M0 likes160 downloads2mo agoHugging Face07GXLXY /mopd-math-code-mix MOPD math+code mix `train/`: math:code ≈ 1:1 平衡集(`math.parquet` + `code_*.parquet` shards) `val/mopd_val_mix.parquet`: AIME24 全量 + MATH-500 子集 + Eurus code_validation 子集 路由字段:`ability ∈ {math, code}` texttext-generation10K<n<100K0 likes145 downloads1mo agoHugging Face08albertge /mix60k-math-code-sft mix60k: 50/50 OpenMathInstruct-2 + OpenCodeInstruct SFT mixture This dataset is from the paper Register Tokens for Bounded-State Reasoning in Diffusion Language Models. The exact 60,000-trace SFT mixture used to train the LLaDA-8B-Base and Dream-7B-Base main triad in the dLLM Registers project. Composition: ~30K math traces from OpenMathInstruct-2 and ~30K code traces from OpenCodeInstruct. License: governed by the upstream licenses of OpenMathInstruct-2 and OpenCodeInstruct.… See the full description on the dataset page: https://huggingface.co/datasets/albertge/mix60k-math-code-sft.texttext-generation10K<n<100K0 likes72 downloads8d agoHugging Face09TheJackBright /verisci-verified-science-math-code VeriSci Verified Science Math Code Verifier-grounded dataset for the Adaption AutoScientist Challenge Part 2, targeting the Science category with secondary Math and Code coverage. Summary VeriSci trains models to solve scientific computations, finite-difference PDE updates, numerical ODE steps, unit-checked mechanics, thermodynamics, circuits, chemistry stoichiometry, molarity, unit conversion, vector decomposition, two-point linear modeling, small Python… See the full description on the dataset page: https://huggingface.co/datasets/TheJackBright/verisci-verified-science-math-code.texttext-generation1K<n<10K0 likes72 downloads2mo agoHugging Face10SMH-DEV-AI /NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54 NOESIS DORA SFT Dataset Multilingual supervised fine-tuning dataset built for the NOESIS QwQ+DeepSeek-R1 MoE pipeline. Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform(framework: DHCF-FNO — Deterministic Hybrid Control Framework for Frozen Neural Operators). Founder: Ilia Bolotnikov Organization: AMAImedia.com X (Twitter): @AMAImediacom LinkedIn: Ilia Bolotnikov Telegram: @djbionicl NOESIS version: v14.8-NT89 Build date: 2026-04… See the full description on the dataset page: https://huggingface.co/datasets/SMH-DEV-AI/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.texttext-generation1M<n<10M0 likes55 downloads5mo agoHugging Face11josephmayo /curated-openbmb-code-math Curated OpenBMB Code/Math Post-Training Data English code/math-focused post-training data derived from curated OpenBMB UltraData rows. Contents Config Rows Schema Purpose sft_no_think 25,891 prompt, response Direct code/math SFT plus necessary technical instruction-following/alignment sft_think 6,018 prompt, response Code/math reasoning SFT with <think>...</think> traces Total rows: 31,909. Curation The SFT split keeps English code… See the full description on the dataset page: https://huggingface.co/datasets/josephmayo/curated-openbmb-code-math.texttext-generation10K<n<100K1 likes37 downloads4mo agoHugging Face12flamiinngo /math-code-qa Math & Code QA — Instruction Dataset Worked mathematical solutions and short code answers, built for the Adaption Labs AutoScientist Challenge (Math & Code category). Rows 5,200 Math 3,600 Code 1,600 Distinct answers 5,199 (100%) Duplicate questions none Nulls none Question length median 27 words Answer length median 58 words (max 89) License CC-BY-4.0 What makes the math rows unusual Every math answer is short worked reasoning… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/math-code-qa.textquestion-answering1K<n<10K1 likes34 downloads2mo agoHugging Face13flamiinngo /math-code-qa-v2 Math & Code QA v2 — Instruction Dataset Worked mathematical solutions and short code answers, spanning arithmetic word problems through to algebra, geometry and combinatorics. Built for the Adaption Labs AutoScientist Challenge (Math & Code category). The model trained on this beats Llama-3.3-70B-Instruct 72 to 28 on the held-out Math category evaluation. Rows 5,297 (4,197 math, 1,100 code) Distinct answers 5,297 (100%) Duplicate questions none Nulls none… See the full description on the dataset page: https://huggingface.co/datasets/flamiinngo/math-code-qa-v2.textquestion-answering1K<n<10K0 likes28 downloads2mo agoHugging Face14gimmy256 /math-code-africa Math & Code — African Context Dataset Instruction-tuning dataset covering math and coding in African contexts: word problems with African currencies (UGX, KES, NGN), names, geography, and real-world scenarios (mobile money, market trading, farming); coding challenges for USSD systems, mobile money APIs, SMS gateways, agricultural data pipelines, and multilingual NLP — grounded via web search, generated with gemini-2.5-flash. Dataset Details Rows: 331 Regions… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/math-code-africa.texttext-generationn<1K0 likes22 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.