datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bespoke-Stratos-17k
Bespoke-Stratos-17k
We replicated and improved the Berkeley Sky-T1 data pipeline using SFT distillation data
from DeepSeek-R1 to create Bespoke-Stratos-17k -- a reasoning dataset of questions, reasoning traces, and answers.
This data was used to train:
Bespoke-Stratos-32B, a 32B reasoning model which is a fine-tune of Qwen-2.5-32B-Instruct
Bespoke-Stratos-7B, a 7B reasoning model which is a fine-tune of Qwen-2.5-7B-Instruct.
Metrics for Bespoke-Stratos-32B… See the full description on the dataset page: https://huggingface.co/datasets/bespokelabs/Bespoke-Stratos-17k.Bespoke-Stratos-17k
Dataset card for Bespoke-Stratos-17k
This dataset is a TRL-compatible version of bespokelabs/Bespoke-Stratos-17k. Please refer to the source dataset for details.
bespoke-stratos-es
Bespoke-Stratos-ES
Spanish reasoning traces regenerated from bespokelabs/Bespoke-Stratos-17k -- 16709 rows, natively generated in Spanish (not machine-translated from the English traces).
Models trained on this dataset
axiom-of-choice/qwen3-4b-es-reasoning-qlora (Qwen3-4B) -- also for transformers/peft: qwen3-4b-es-reasoning-peft
axiom-of-choice/qwen3-1.7b-es-reasoning-lora (Qwen3-1.7B) -- also: qwen3-1.7b-es-reasoning-peft… See the full description on the dataset page: https://huggingface.co/datasets/axiom-of-choice/bespoke-stratos-es.Bespoke-Stratos-35kBespoke-Stratos-17k-Train-Posterior-PAtulu-3-sft-Bespoke-Stratos-17kBespoke-Stratos-17k-DeepSeekrized
Bespoke-Stratos-17k-DeepSeekrized
Created by: Seungwoo Ryu
Introduction
This dataset is a modified version of the original HuggingFaceH4/Bespoke-Stratos-17k dataset, reformatted to match the output format of DeepSeek models.
Modifications
The user and assistant fields from the original dataset's messages have been moved to user_modified and agent_modified respectively.
The content in the agent_modified field has been transformed to match the DeepSeek model's… See the full description on the dataset page: https://huggingface.co/datasets/tryumanshow/Bespoke-Stratos-17k-DeepSeekrized.Bespoke-Stratos-7B_eval_3a76
chengfu0118/Bespoke-Stratos-7B_eval_3a76
Precomputed model outputs for evaluation.
Evaluation Results
MATH500
Accuracy: 48.80%
Accuracy
Questions Solved
Total Questions
48.80%
244
500
Bespoke-Stratos-35k-messagesBespoke-Stratos-17k-SourceBespoke-Stratos-17k-qa-onlybespoke_stratos_17k_convertedThis is a converted version of the bespokelabs/Bespoke-Stratos-17k dataset into Tulu SFT training format.
The conversion script can be found in our open-instruct repo.
The conversion took the following parameters:
apply_keyword_filters: False
apply_empty_message_filters: False
push_to_hub: True
hf_entity: natolambert
converted_dataset_name: bespoke_stratos_17k_converted
local_save_dir: None
Please refer to the original dataset for more information about this dataset and the license.
Bespoke-Stratos-7B_1755006372_eval_466d
chengfu0118/Bespoke-Stratos-7B_1755006372_eval_466d
Precomputed model outputs for evaluation.
Evaluation Results
MATH500
Accuracy: 79.60%
Accuracy
Questions Solved
Total Questions
79.60%
398
500
Custom-Bespoke-Stratos-7B_1753961306_eval_466d_math500_skip_attn_1
chengfu0118/Custom-Bespoke-Stratos-7B_1753961306_eval_466d_math500_skip_attn_1
Precomputed model outputs for evaluation.
Evaluation Results
MATH500
Accuracy: 1.80%
Accuracy
Questions Solved
Total Questions
1.80%
9
500
Bespoke-Stratos-12K-kostep3-input-bespoke-stratos-17k-test1
Chain of Thought生成データセット
このデータセットは、問題の解答から説明(Chain of Thought)を生成するためのデータセットです。
概要
処理したサンプル数: 92
有効な説明生成数: 92
生成成功率: 100.00%
使用モデル: Qwen/Qwen3-14B
トークン数統計
最小トークン数: 227
最大トークン数: 2348
平均トークン数: 653.8
トークン数分布
0-100トークン: 0件 (0.0%)
101-500トークン: 47件 (51.1%)
501-1000トークン: 31件 (33.7%)
1001-2000トークン: 13件 (14.1%)
2001-5000トークン: 1件 (1.1%)
5001+トークン: 0件 (0.0%)
データセット構造
system_prompt: モデルに送信されたシステムプロンプト
question_text: 元の問題文
answer_text: 問題の解答… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step3-input-bespoke-stratos-17k-test1.bespoke-stratos-17kBespoke-Stratos-35k-qa-onlyBespoke-Stratos-17k-KoEnKo
Let them think in English.
English system prompt + Korean question + English thinking + Korean answer
System message changes
Your role as an assistant involves thoroughly exploring questions ...(중략)...
<|begin_of_solution|> {final formatted, precise, and clear solution **written in the same language as the question.**} <|end_of_solution|> ...(하략)
Translation
Translated with gemini-2.0-flash
Question
Return your final response within \\boxed{}., Generate an… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/Bespoke-Stratos-17k-KoEnKo.Bespoke-Stratos-17k-Numina-Subsetbespoke_stratos_17kBespoke-Stratos-7B_1755005674_eval_07bc
chengfu0118/Bespoke-Stratos-7B_1755005674_eval_07bc
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 33.30% ± 0.70%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
35.62%
182
511
2
32.29%
165
511
3
30.72%
157
511
4
34.25%
175
511
5
34.05%
174
511
6
32.88%
168
511
Bespoke-Stratos-17k-Init-Model-Final-Reinforce-Baseline-Iter1-Strong-Init-Filtered-MergedBespoke-Stratos-17k-iter-2Bespoke-Stratos-17k-ModifiedBespoke-Stratos-17k_tokenizedBespoke-Stratos-50kbespoke_stratos_17k-testBespoke-Stratos-17k-with-valCustom-Bespoke-Stratos-7B_1753940626_eval_104b_aime24_skip_attn_5
chengfu0118/Custom-Bespoke-Stratos-7B_1753940626_eval_104b_aime24_skip_attn_5
Precomputed model outputs for evaluation.
Evaluation Results
AIME24
Average Accuracy: 0.00% ± 0.00%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
0.00%
0
30
2
0.00%
0
30
3
0.00%
0
30
4
0.00%
0
30
5
0.00%
0
30
6
0.00%
0
30
7
0.00%
0
30
8
0.00%
0
30
9
0.00%
0
30
10
0.00%
0
30
