CoolFace
20 results

math-code

t2ance /atlas-35-joint-math-and-code-training-data 35. One training set of mathematics and code: OpenMathReasoning and OpenCodeReasoning-2 together 1. Question and links Read this first. The reading copy of this directory is t2ance/atlas-experiments under 35-joint-math-and-code-training-data/; the saved training steps and the per-token training arrays are on the Hugging Face repository t2ance/atlas-35-joint-math-and-code-training-data only. Can OpenCodeReasoning-2 (OCR-2) give code rows with the properties… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-35-joint-math-and-code-training-data.0 likes2.2k downloads10h agoHugging Facemathmadness /MathCodertext100M<n<1B4 likes1.3k downloads3y agoHugging FaceHugodonotexit /math-code-science-deepseek-r1-en R1 Dataset Collection Aggregated high-quality English prompts and model-generated responses from DeepSeek R1 and DeepSeek R1-0528. Dataset Summary The R1 Dataset Collection combines multiple public DeepSeek-generated instruction-response corpora into a single, cleaned, English-only JSONL file. Each example consists of a <|user|> prompt and a <|assistant|> response in one "text" field. This release includes: ~21,000 examples from the DeepSeek-R1-0528 Distilled Custom… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/math-code-science-deepseek-r1-en.textquestion-answering1M<n<10M5 likes1k downloads1y agoHugging FaceTMoC /SYNTH-Swallow-Math-Code-Mix Mixed dataset: SYNTH + SwallowMath-v2 + SwallowCode-v2 This high-signal, all synthetic dataset is a complete shuffled mix of the following four sources: SYNTH ~63.5% SwallowCode-v2 ~15.5% SwallowMath-v2-textbook ~10.5% SwallowMath-v2-qa ~10.0% The motivation to provide this on HF was the need for a convenient, pre-shuffled merge of the highest quality synthetic / augmented datasets for small language model pre-training experiments as of… See the full description on the dataset page: https://huggingface.co/datasets/TMoC/SYNTH-Swallow-Math-Code-Mix.texttext-generation100M<n<1B1 likes728 downloads8mo agoHugging FaceMathGenie /MathCode-Pile MathCode-Pile MathCode-Pile is a dataset for continue pretraining large language models to enhance their mathematical reasoning abilities. It is introduced in the paper MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code. It contains 19.2B tokens, with math-related data covering web pages, textbooks, model-synthesized text, and math related code. Currently, filtered-OpenWebMath, filtered-CC-En-math, and translated mathematical code are… See the full description on the dataset page: https://huggingface.co/datasets/MathGenie/MathCode-Pile.text100K<n<1M25 likes603 downloads2y agoHugging FaceMathGenie /MathCode-Pile-Full MathAnalystPile MathAnalystPile is a dataset for continue pretraining large language models to enhance their mathematical reasoning abilities. It contains approximately 20 B tokens, with math-related data covering web pages, textbooks, model-synthesized text, and math related code. We open source the full pretrain dataset to facilitate future research in this field. Data Composition MathAnalystPile contains a wide range of math-related data. The number of tokens of each… See the full description on the dataset page: https://huggingface.co/datasets/MathGenie/MathCode-Pile-Full.0 likes522 downloads2y agoHugging Face