octothinker
MegaMath-Web-Pro-Max
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
The Curation of MegaMath-Web-Pro-Max
Step 1: Uniformly and randomly sample millions of documents from the MegaMath-Web corpus, stratified by publication year;
Step 2: Annotate them using Llama-3.1-70B-instruct with a scoring prompt from FineMath and prepare the seed data;
Step 3: Training a fasttext carefully with proper preprocessing;
Step 4: Filtering documents with a threshold (i.e., 0.4);
Step 5:… See the full description on the dataset page: https://huggingface.co/datasets/OctoThinker/MegaMath-Web-Pro-Max.octothinker_decay_stage2_1Btokensself-play-sft-octothinker-Qwen-Qwen3-32B
Dataset Card for "self-play-sft-octothinker-Qwen-Qwen3-32B"
More Information needed
self-play-sft-octothinker-Qwen-Qwen3-32B-52kOctoThinker-d1MOctoThinker-d60000-v2
