math-code
atlas-35-joint-math-and-code-training-data
35. One training set of mathematics and code: OpenMathReasoning and OpenCodeReasoning-2 together
1. Question and links
Read this first. The reading copy of this directory is t2ance/atlas-experiments under 35-joint-math-and-code-training-data/; the saved training steps and the per-token training arrays are on the Hugging Face repository t2ance/atlas-35-joint-math-and-code-training-data only.
Can OpenCodeReasoning-2 (OCR-2) give code rows with the properties… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-35-joint-math-and-code-training-data.MathCodermath-code-science-deepseek-r1-en
R1 Dataset Collection
Aggregated high-quality English prompts and model-generated responses from DeepSeek R1 and DeepSeek R1-0528.
Dataset Summary
The R1 Dataset Collection combines multiple public DeepSeek-generated instruction-response corpora into a single, cleaned, English-only JSONL file. Each example consists of a <|user|> prompt and a <|assistant|> response in one "text" field. This release includes:
~21,000 examples from the DeepSeek-R1-0528 Distilled Custom… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/math-code-science-deepseek-r1-en.SYNTH-Swallow-Math-Code-Mix
Mixed dataset: SYNTH + SwallowMath-v2 + SwallowCode-v2
This high-signal, all synthetic dataset is a complete shuffled mix of the following four sources:
SYNTH ~63.5%
SwallowCode-v2 ~15.5%
SwallowMath-v2-textbook ~10.5%
SwallowMath-v2-qa ~10.0%
The motivation to provide this on HF was the need for a convenient, pre-shuffled merge of the highest quality synthetic / augmented datasets for small language model pre-training experiments as of… See the full description on the dataset page: https://huggingface.co/datasets/TMoC/SYNTH-Swallow-Math-Code-Mix.MathCode-Pile
MathCode-Pile
MathCode-Pile is a dataset for continue pretraining large language models to enhance their mathematical reasoning abilities. It is introduced in the paper MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code. It contains 19.2B tokens, with math-related data covering web pages, textbooks, model-synthesized text, and math related code. Currently, filtered-OpenWebMath, filtered-CC-En-math, and translated mathematical code are… See the full description on the dataset page: https://huggingface.co/datasets/MathGenie/MathCode-Pile.MathCode-Pile-Full
MathAnalystPile
MathAnalystPile is a dataset for continue pretraining large language models to enhance their mathematical reasoning abilities. It contains approximately 20 B tokens, with math-related data covering web pages, textbooks, model-synthesized text, and math related code. We open source the full pretrain dataset to facilitate future research in this field.
Data Composition
MathAnalystPile contains a wide range of math-related data. The number of tokens of each… See the full description on the dataset page: https://huggingface.co/datasets/MathGenie/MathCode-Pile-Full.
