CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mathmadness /MathCodertext100M<n<1B4 likes1.3k downloads3y agoHugging Face02Hugodonotexit /math-code-science-deepseek-r1-en R1 Dataset Collection Aggregated high-quality English prompts and model-generated responses from DeepSeek R1 and DeepSeek R1-0528. Dataset Summary The R1 Dataset Collection combines multiple public DeepSeek-generated instruction-response corpora into a single, cleaned, English-only JSONL file. Each example consists of a <|user|> prompt and a <|assistant|> response in one "text" field. This release includes: ~21,000 examples from the DeepSeek-R1-0528 Distilled Custom… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/math-code-science-deepseek-r1-en.textquestion-answering1M<n<10M5 likes1k downloads1y agoHugging Face03TMoC /SYNTH-Swallow-Math-Code-Mix Mixed dataset: SYNTH + SwallowMath-v2 + SwallowCode-v2 This high-signal, all synthetic dataset is a complete shuffled mix of the following four sources: SYNTH ~63.5% SwallowCode-v2 ~15.5% SwallowMath-v2-textbook ~10.5% SwallowMath-v2-qa ~10.0% The motivation to provide this on HF was the need for a convenient, pre-shuffled merge of the highest quality synthetic / augmented datasets for small language model pre-training experiments as of… See the full description on the dataset page: https://huggingface.co/datasets/TMoC/SYNTH-Swallow-Math-Code-Mix.texttext-generation100M<n<1B1 likes728 downloads8mo agoHugging Face04MathGenie /MathCode-Pile MathCode-Pile MathCode-Pile is a dataset for continue pretraining large language models to enhance their mathematical reasoning abilities. It is introduced in the paper MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code. It contains 19.2B tokens, with math-related data covering web pages, textbooks, model-synthesized text, and math related code. Currently, filtered-OpenWebMath, filtered-CC-En-math, and translated mathematical code are… See the full description on the dataset page: https://huggingface.co/datasets/MathGenie/MathCode-Pile.text100K<n<1M25 likes593 downloads2y agoHugging Face05AMAImedia /NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54 ⚡ Each donation funds the next large quant. I host free GGUF or MoE quants as independent research. Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant. 🎉 Boosty🦄 &nbsp;|&nbsp; ☕ Buy Me a Coffee🦄 &nbsp;|&nbsp; ⭐ DonationAlerts🦄 💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.text1M<n<10M2 likes415 downloads9d agoHugging Face06mlfoundations-dev /hero_run_4_math_codetabular1M<n<10M0 likes375 downloads1y agoHugging Face07AMAImedia /NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54 ⚡ Each donation funds the next large quant. I host free GGUF or MoE quants as independent research. Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant. 🎉 Boosty🦄 &nbsp;|&nbsp; ☕ Buy Me a Coffee🦄 &nbsp;|&nbsp; ⭐ DonationAlerts🦄 💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.text10K<n<100K7 likes349 downloads9d agoHugging Face08semran1 /test2_math_code_addedtext10M<n<100M0 likes305 downloads10mo agoHugging Face09MathLLMs /MathCodeInstruct MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning Paper: https://arxiv.org/pdf/2310.03731.pdf Repo: https://github.com/mathllm/MathCoder Introduction We introduce MathCoder, a series of open-source large language models (LLMs) specifically tailored for general math problem-solving. Base Model: Llama-2 Base Model: Code Llama MathCoder-L-7B MathCoder-CL-7B MathCoder-L-13B MathCoder-CL-34B Training Data The models… See the full description on the dataset page: https://huggingface.co/datasets/MathLLMs/MathCodeInstruct.textquestion-answering10K<n<100K25 likes302 downloads2y agoHugging Face10manifesta /verified-math-code-17k Verified Math & Code, 17,000 rows A math and code instruction dataset where every single row was mechanically checked before it was allowed in. Not filtered by a heuristic, not scored by a model. Checked. Two layers of verification, one per domain: Every math answer was compared against an independent gold answer by exact, numeric and symbolic (SymPy) comparison. If the worked solution did not arrive at the gold answer, the row was dropped.… See the full description on the dataset page: https://huggingface.co/datasets/manifesta/verified-math-code-17k.texttext-generation10K<n<100K0 likes202 downloads1mo agoHugging Face11MathLLMs /MathCodeInstruct-Plus MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning Paper: https://arxiv.org/pdf/2310.03731.pdf Repo: https://github.com/mathllm/MathCoder Introduction We introduce MathCoder, a series of open-source large language models (LLMs) specifically tailored for general math problem-solving. Base Model: Llama-2 Base Model: Code Llama MathCoder-L-7B MathCoder-CL-7B MathCoder-L-13B MathCoder-CL-34B Training Data The models… See the full description on the dataset page: https://huggingface.co/datasets/MathLLMs/MathCodeInstruct-Plus.textquestion-answering10K<n<100K17 likes198 downloads2y agoHugging Face12mlfoundations-dev /32b_exploit_seed_math_code_dedup_decontaminatetabular100K<n<1M0 likes170 downloads2y agoHugging Face13theblackcat102 /codex-math-qaSolution by codex-davinci-002 for math_qatexttext-generation10K<n<100K31 likes169 downloads4y agoHugging Face14chhao /AM-DeepSeek-R1-Filtered-Math-Code AM DeepSeek R1 Filtered Math and Code This repository publishes reproducible training subsets derived from a-m-team/AM-DeepSeek-R1-Distilled-1.4M at pinned revision 53531c06634904118a2dcd83961918c4d69d1cdf. The source dataset and this derived release use CC BY-NC 4.0. Commercial use is not permitted by that license. Preserve attribution and review the upstream dataset card before use. Contents Config / split Records Bytes SHA-256 math / train 111,657 2… See the full description on the dataset page: https://huggingface.co/datasets/chhao/AM-DeepSeek-R1-Filtered-Math-Code.texttext-generation100K<n<1M0 likes160 downloads2mo agoHugging Face15GXLXY /mopd-math-code-mix MOPD math+code mix `train/`: math:code ≈ 1:1 平衡集(`math.parquet` + `code_*.parquet` shards) `val/mopd_val_mix.parquet`: AIME24 全量 + MATH-500 子集 + Eurus code_validation 子集 路由字段:`ability ∈ {math, code}` texttext-generation10K<n<100K0 likes145 downloads1mo agoHugging Face16albertge /mix60k-math-code-sft mix60k: 50/50 OpenMathInstruct-2 + OpenCodeInstruct SFT mixture This dataset is from the paper Register Tokens for Bounded-State Reasoning in Diffusion Language Models. The exact 60,000-trace SFT mixture used to train the LLaDA-8B-Base and Dream-7B-Base main triad in the dLLM Registers project. Composition: ~30K math traces from OpenMathInstruct-2 and ~30K code traces from OpenCodeInstruct. License: governed by the upstream licenses of OpenMathInstruct-2 and OpenCodeInstruct.… See the full description on the dataset page: https://huggingface.co/datasets/albertge/mix60k-math-code-sft.texttext-generation10K<n<100K0 likes72 downloads9d agoHugging Face17TheJackBright /verisci-verified-science-math-code VeriSci Verified Science Math Code Verifier-grounded dataset for the Adaption AutoScientist Challenge Part 2, targeting the Science category with secondary Math and Code coverage. Summary VeriSci trains models to solve scientific computations, finite-difference PDE updates, numerical ODE steps, unit-checked mechanics, thermodynamics, circuits, chemistry stoichiometry, molarity, unit conversion, vector decomposition, two-point linear modeling, small Python… See the full description on the dataset page: https://huggingface.co/datasets/TheJackBright/verisci-verified-science-math-code.texttext-generation1K<n<10K0 likes72 downloads2mo agoHugging Face18weqweasdas /preference_dataset_mixture2_and_safe_pku30k_and_argilla_math_and_ultra_code_for_preference_modeltext100K<n<1M0 likes71 downloads2y agoHugging Face19NP235 /MathCode-Pile MathCode-Pile MathCode-Pile is a dataset for continue pretraining large language models to enhance their mathematical reasoning abilities. It is introduced in the paper MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code. It contains 19.2B tokens, with math-related data covering web pages, textbooks, model-synthesized text, and math related code. Currently, filtered-OpenWebMath, filtered-CC-En-math, and translated mathematical code… See the full description on the dataset page: https://huggingface.co/datasets/NP235/MathCode-Pile.text100K<n<1M0 likes68 downloads3mo agoHugging Face20MathCodeBench /linear-programmingtextn<1K2 likes65 downloads2y agoHugging Face21open-llm-leaderboard /Josephgflowers__TinyLlama_v1.1_math_code-world-test-1-detailsgated Dataset Card for Evaluation run of Josephgflowers/TinyLlama_v1.1_math_code-world-test-1 Dataset automatically created during the evaluation run of model Josephgflowers/TinyLlama_v1.1_math_code-world-test-1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Josephgflowers__TinyLlama_v1.1_math_code-world-test-1-details.tabular10K<n<100K0 likes61 downloads2y agoHugging Face22sinatra-rd /math-to-code-gpt4o-finetuning-jsonlThis is a high quality dataset for fine tuning GPT4o and GPT4o mini with a focus on solving problems with mathematical operations using different programming languages ​​in a similar way to the code interpreter. Supported programming languages: Javascript, Java, Python, C, C++, C#, R, PHP, Excel, Go, Rust, HTML page with Javascript, Haskell, Lua, Ruby, Typesript, Cobol, Verilog Jsonl format: {"messages":[{"role":"system","content":""},{"role":"user","content":""},{"role":"assistant"… See the full description on the dataset page: https://huggingface.co/datasets/sinatra-rd/math-to-code-gpt4o-finetuning-jsonl.textn<1K0 likes58 downloads1y agoHugging Face23SMH-DEV-AI /NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54 NOESIS DORA SFT Dataset Multilingual supervised fine-tuning dataset built for the NOESIS QwQ+DeepSeek-R1 MoE pipeline. Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform(framework: DHCF-FNO — Deterministic Hybrid Control Framework for Frozen Neural Operators). Founder: Ilia Bolotnikov Organization: AMAImedia.com X (Twitter): @AMAImediacom LinkedIn: Ilia Bolotnikov Telegram: @djbionicl NOESIS version: v14.8-NT89 Build date: 2026-04… See the full description on the dataset page: https://huggingface.co/datasets/SMH-DEV-AI/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.texttext-generation1M<n<10M0 likes55 downloads5mo agoHugging Face24communityai /aptchat-v2-math-code-general-50ktext10K<n<100K0 likes47 downloads3y agoHugging Face25haowu89 /math-ai-bench-sources-code math-ai-bench-sources-code A code benchmark evaluation dataset with 83,072 solution trajectories generated by state-of-the-art thinking models on coding benchmark problems. Overview Each entry is a long-form solution trajectory (chain-of-thought + final code) produced by a reasoning model on a held-out coding benchmark. Every trajectory carries a verified correct label, and every problem carries a correct_ratio (pass rate over all trajectories for that problem).… See the full description on the dataset page: https://huggingface.co/datasets/haowu89/math-ai-bench-sources-code.tabular10K<n<100K0 likes45 downloads3mo agoHugging Face26GXLXY /mopd-math-code-full-valtext1K<n<10K0 likes41 downloads27d agoHugging Face27happzy2633 /math_codetext10K<n<100K1 likes39 downloads10mo agoHugging Face28Rishabh-sucks-at-code /math_dataset_tinytext1K<n<10K0 likes38 downloads2y agoHugging Face29open-llm-leaderboard /EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math-detailsgated Dataset Card for Evaluation run of EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-math-details.tabular10K<n<100K0 likes37 downloads2y agoHugging Face30open-llm-leaderboard /DeepMount00__Qwen2.5-7B-Instruct-MathCoder-detailsgated Dataset Card for Evaluation run of DeepMount00/Qwen2.5-7B-Instruct-MathCoder Dataset automatically created during the evaluation run of model DeepMount00/Qwen2.5-7B-Instruct-MathCoder The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DeepMount00__Qwen2.5-7B-Instruct-MathCoder-details.tabular10K<n<100K0 likes37 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.