datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openthoughts3-en-ar-midtrain
openthoughts3-en-ar-midtrain
Arabic translation of the OpenThoughts3_1.2M split of smoltalk2 (config Mid): long mathematical reasoning traces with <think> blocks, in a two-message user/assistant format. Translated with google/gemma-4-12B-it (bf16, greedy) on A100s. All 1,135,104 source rows are present, none dropped.
The pipeline segments each message into prose and verbatim blocks (code, LaTeX, tables, and inline non-translatables are masked and never sent to the model)… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/openthoughts3-en-ar-midtrain.open-thoughts-4-math-qwen3-32b-annotated
Dataset Card for Open-Thoughts-4-Math-Qwen3-32B-Annotated
This dataset is the Qwen3-32B annotated version of mlfoundations-dev/hero_run_4_math curated by the
OpenThoughts4 team. We provide the responses from Qwen3-32B in the generated_text column. These samples were generated using temperature = 0.8 and max output tokens = 7,500.
We note that many of the responses are truncated, so use this dataset wisely!
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-math-qwen3-32b-annotated.OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-topk128-normalizedOpenThoughts3openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554
mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
HLE
HMMT
AIME25
LiveCodeBenchv5
Accuracy
34.3
74.5
79.4
49.4
51.0
44.3
53.9
21.5
23.1
12.2
17.0
22.7
40.1
AIME24
Average Accuracy: 34.33% ± 1.89%
Number of Runs: 10
Run
Accuracy
Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/openthoughts3_code_100k_annotated_QwQ-32B_sharegpt_eval_5554.openthoughts-4-code-qwen3-32b-32k-annotatedopenthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16
OpenThoughts-4 Code SDG: Qwen3-30B-A3B-Thinking-2507 (n=16, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-30B-A3B-Thinking-2507
on the Marin OpenThoughts-4 code SDG prompt
set.
Each prompt is sampled n=16 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-30b-a3b-thinking-2507-n16-flattened-logprobs-k16.openthoughts4-code-9168-prompts-qwen3-32b-n16-flattened-logprobs-k16
OpenThoughts-4 Code SDG: Qwen3-32B (n=16, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-32B
on the Marin OpenThoughts-4 code SDG prompt
set.
Each prompt is sampled n=16 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field
Value
Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-qwen3-32b-n16-flattened-logprobs-k16.openthoughts3-unfinished-promptsOpenThoughts3-456k-no-cotQwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554
mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
75.7
98.8
90.4
58.1
73.7
68.2
41.9
46.8
47.2
67.7
13.9
64.3
52.0
AIME24
Average Accuracy: 75.67% ± 1.57%
Number of Runs: 10
Run
Accuracy
Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/QwQ-32B_enable-liger-kernel_False_OpenThoughts3_3k_eval_5554.OpenThoughts3-456k-no-cot-with-olmo-system-promptOpenThoughts-114k-math-correct-qwen3-14b-math-prepared-normalizedOpenThoughts3-456k-gpt4.1-cotopenthoughts4-code-9168-prompts-glm-5.2-n4
OpenThoughts-4 Code — GLM-5.2 n=4
Quality-filtered synthetic responses from
zai-org/GLM-5.2-FP8 for the
9,168 unique instruction_seed values in
mlfoundations-dev/hero_run_4_code.
Each prompt has four accepted responses, for 36,672 rows total.
Generation
Field
Value
Generator
zai-org/GLM-5.2-FP8
Samples per prompt
4
Temperature
1.0
Top-p
0.95
Maximum generated tokens
256,000
Thinking mode
enabled
Inference engine
vLLM on 8 GB200 GPUs… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-code-9168-prompts-glm-5.2-n4.openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16
OpenThoughts-4 Science SDG: Qwen3-30B-A3B-Thinking-2507 (n=8, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-30B-A3B-Thinking-2507
on the Marin OpenThoughts-4 science SDG prompt
set.
Each prompt is sampled n=8 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16.openthoughts_math50k-all-style-pcaopen-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8
Open Thoughts 4 - Math (Qwen3-32B, 32K tokens, n=8)
This dataset contains math reasoning problems with 8 independent responses generated by Qwen3-32B.
Overview
Source: marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens
Model: Qwen/Qwen3-32B
Temperature: 0.8
Max tokens: 32,768
Columns
Column
Description
instruction_seed
The math problem prompt
_source
Source dataset identifier
gpt41_mini_response
Reference response from… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8.openthoughts_math50k-k4-all-style-pcaopenthoughts3_100k_llama3_eval_5554
mlfoundations-dev/openthoughts3_100k_llama3_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
HLE
HMMT
AIME25
LiveCodeBenchv5
Accuracy
37.0
75.2
83.8
11.2
45.2
45.1
44.4
13.8
18.3
9.7
19.3
30.3
31.9
AIME24
Average Accuracy: 37.00% ± 1.29%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/openthoughts3_100k_llama3_eval_5554.openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16
OpenThoughts-4 Science SDG: Qwen3-32B (n=8, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-32B
on the Marin OpenThoughts-4 science SDG prompt
set.
Each prompt is sampled n=8 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field
Value
Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16.open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8
Open Thoughts 4 - Code (Qwen3-30B-A3B-Thinking-2507, 32K tokens, n=8)
This dataset contains code reasoning problems with 8 independent responses generated by Qwen3-30B-A3B-Thinking-2507.
Overview
Source: marin-community/open-thoughts-4-30k-code-qwen3-32b-annotated (prompts only)
Model: Qwen/Qwen3-30B-A3B-Thinking-2507
Temperature: 0.8
Max tokens: 32,768
Columns
Column
Description
instruction_seed
The code problem prompt
_source
Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-code-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8.open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8
Open Thoughts 4 - Math (Qwen3-30B-A3B-Thinking-2507, 32K tokens, n=8)
This dataset contains math reasoning problems with 8 independent responses generated by Qwen3-30B-A3B-Thinking-2507.
Overview
Source: marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated (base prompts)
Model: Qwen/Qwen3-30B-A3B-Thinking-2507
Temperature: 0.8
Max tokens: 32,768
Columns
Column
Description
instruction_seed
The math problem prompt
_source
Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8.open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted
Dataset Card for Open-Thoughts-4-30K-Math-Qwen3-30B-A3B-Thinking-2507-Annotated-32768-Tokens-N8-Reformatted
Overview
This dataset is a reformatted version of marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8. The original dataset contained 29,963 samples, each with 8 responses generated by the same model with different random seeds (stored in generated_text, generated_text2, ..., generated_text8 columns). This reformatted… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-30b-a3B-thinking-2507-annotated-32768-tokens-n8-reformatted.open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8
Open Thoughts 4 - Math (Qwen3-4B, 32K tokens, n=8)
This dataset contains math reasoning problems with 8 independent responses generated by Qwen3-4B.
Overview
Source: marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens (n=1 version with 1 response per prompt)
Model: Qwen/Qwen3-4B
Temperature: 0.8
Max tokens: 32,768
Columns
Column
Description
instruction_seed
The math problem prompt
_source
Source dataset identifier… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens-n8.openthoughts3_300k_annotated_Qwen3-32Bopenthoughts-tblite
NousResearch/openthoughts-tblite
This dataset is a reformatted version of OpenThoughts-TBLite for use with the Hermes Agent Terminal-Bench evaluation framework.
Source
OpenThoughts-TBLite was created by the OpenThoughts Agent team in collaboration with Snorkel AI and Bespoke Labs. It is a difficulty-calibrated subset of Terminal-Bench 2.0 designed for faster iteration when developing terminal agents.
Original dataset: open-thoughts/OpenThoughts-TBLite
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/openthoughts-tblite.OpenThoughts-114k-math-correct-qwen3-14b-math-preparedopen-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens
Dataset Card for Open-Thoughts-4-30K-Math-Qwen3-4B-Annotated-32768-Tokens
Overview
This dataset is a variant of the OpenThoughts-4 30K math subset with responses generated by Qwen/Qwen3-4B using max output tokens = 32768, allowing for longer and more complete chain-of-thought reasoning.
Generation Details
Model: Qwen/Qwen3-4B
Temperature: 0.8
Max Output Tokens: 32768
Dataset Statistics
Number of Samples: 29,963
Split: train
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-30k-math-qwen3-4b-annotated-32768-tokens.open-thoughts-4-6865-math-kimi-k2pt5-annotated-32768-tokens-n8-reformatted
open-thoughts-4-6865-math-kimi-k2pt5-annotated-32768-tokens
Math reasoning responses generated by Kimi K2.5 (moonshotai/Kimi-K2.5) via a Together AI dedicated instance.
Overview
Total rows: 54,920
Unique prompts: 6,865 (each with 8 response annotations)
Source prompts: marin-community/open-thoughts-4-30k-math-qwen3-32b-annotated-32768-tokens-n8-reformatted
Generation model: moonshotai/Kimi-K2.5
Max tokens: 32,768
Temperature: 0.8
Tokenizer used for stats: Qwen/Qwen2.5-3B… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/open-thoughts-4-6865-math-kimi-k2pt5-annotated-32768-tokens-n8-reformatted.
