datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenThoughts3-456k-no-cot-with-olmo-system-promptsystem-prompt-reasoning-traces
System-Prompt Reasoning Traces
A novel dataset combining system prompt adherence with structured internal reasoning traces, built on findings from 14+ research papers.
🔬 Research Foundation
This dataset is the first to systematically combine system prompt diversity with structured reasoning traces. It incorporates findings from:
Paper
Key Finding
How We Use It
Sky-T1 (Berkeley, 2025)
Structure > content in reasoning traces — wrong answers with good structure… See the full description on the dataset page: https://huggingface.co/datasets/Michael-Kozu/system-prompt-reasoning-traces.wildchat_perturbed_6000_replaced_no_keyword-with-olmo-system-promptfine-preferences-magpie-generated-system-prompt-v0
Dataset Card for fine-preferences-magpie-generated-system-prompt-v0
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/fine-preferences-magpie-generated-system-prompt-v0/raw/main/pipeline.yaml"
or explore the… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/fine-preferences-magpie-generated-system-prompt-v0.fine-preferences-magpie-generated-system-prompt-v3
Dataset Card for fine-preferences-magpie-generated-system-prompt-v3
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/fine-preferences-magpie-generated-system-prompt-v3/raw/main/pipeline.yaml"
or explore the… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/fine-preferences-magpie-generated-system-prompt-v3.system-promptssystem-prompt-dataset
System Prompt Compilation Dataset
System prompts and associated plausible user messages, generated for training the
operator basis regression (prompt compilation) framework.
Seed prompts are drawn from reshabhs/SPML_Chatbot_Prompt_Injection;
synthetic prompts are generated by an LLM conditioned on seed style/structure.
Schema
Column
Type
Description
id
string
Unique record identifier
system_prompt
string
The system prompt text
source
string
seed… See the full description on the dataset page: https://huggingface.co/datasets/ChapAF/system-prompt-dataset.fine-preferences-magpie-generated-system-prompt-v2
Dataset Card for fine-preferences-magpie-generated-system-prompt-v2
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/fine-preferences-magpie-generated-system-prompt-v2/raw/main/pipeline.yaml"
or explore the… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/fine-preferences-magpie-generated-system-prompt-v2.Unroll-Qwen2.5-7B-Instruct-attn-3-ffn-0-nemo-bespoke-w-system-prompt-seq16k_1755772285_eval_104bsemdpo_pref_pairs_system_prompt_23_testfine-preferences-magpie-generated-system-prompt-v1
Dataset Card for fine-preferences-magpie-generated-system-prompt-v1
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/fine-preferences-magpie-generated-system-prompt-v1/raw/main/pipeline.yaml"
or explore the… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/fine-preferences-magpie-generated-system-prompt-v1.Unroll-Qwen2.5-7B-Instruct-attn-3-ffn-0-nemo-bespoke-w-system-prompt-seq16k_1755766840_eval_104bUnroll-Qwen2.5-7B-Instruct-attn-5-ffn-0-nemo-bespoke-w-system-prompt-seq16k_1755766614_eval_07bcUnroll-Qwen2.5-7B-Instruct-attn-4-ffn-0-nemo-bespoke-w-system-prompt-seq16k_1755766568_eval_07bcUnroll-Qwen2.5-7B-Instruct-attn-0-ffn-1-nemo-bespoke-w-system-prompt-seq16k_1755774247_eval_104bqwen25-7b-it-bespoke-w-system-prompt-seq32k-bs128-steps400_1753823761_eval_2e04
chengfu0118/qwen25-7b-it-bespoke-w-system-prompt-seq32k-bs128-steps400_1753823761_eval_2e04
Precomputed model outputs for evaluation.
Evaluation Results
AIME24
Average Accuracy: 18.67% ± 0.84%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
16.67%
5
30
2
16.67%
5
30
3
20.00%
6
30
4
16.67%
5
30
5
16.67%
5
30
6
23.33%
7
30
7
16.67%
5
30
8
23.33%
7
30
9
16.67%
5
30
10
20.00%
6
30
qwen25-7b-it-bespoke-w-system-prompt-seq32k-bs128-steps400_1753825945_eval_7a7d
chengfu0118/qwen25-7b-it-bespoke-w-system-prompt-seq32k-bs128-steps400_1753825945_eval_7a7d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
MATH500
GPQADiamond
Accuracy
71.0
40.2
MATH500
Accuracy: 71.00%
Accuracy
Questions Solved
Total Questions
71.00%
355
500
GPQADiamond
Average Accuracy: 40.24% ± 2.48%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/chengfu0118/qwen25-7b-it-bespoke-w-system-prompt-seq32k-bs128-steps400_1753825945_eval_7a7d.qwen25-7b-it-bespoke-w-system-prompt-seq32k-bs128-steps400_1753887675_eval_104b
chengfu0118/qwen25-7b-it-bespoke-w-system-prompt-seq32k-bs128-steps400_1753887675_eval_104b
Precomputed model outputs for evaluation.
Evaluation Results
AIME24
Average Accuracy: 20.33% ± 1.37%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
23.33%
7
30
2
16.67%
5
30
3
23.33%
7
30
4
16.67%
5
30
5
20.00%
6
30
6
26.67%
8
30
7
26.67%
8
30
8
13.33%
4
30
9
20.00%
6
30
10
16.67%
5
30
Unroll-Qwen2.5-7B-Instruct-attn-0-ffn-0-nemo-bespoke-w-system-prompt-seq16k_1755766658_eval_104bUnroll-Qwen2.5-7B-Instruct-attn-0-ffn-0-nemo-bespoke-w-system-prompt-seq16k_1755767012_eval_466dUnroll-Qwen2.5-7B-Instruct-attn-1-ffn-0-nemo-bespoke-w-system-prompt-seq16k_1755767100_eval_466dUnroll-Qwen2.5-7B-Instruct-attn-0-ffn-1-nemo-bespoke-w-system-prompt-seq16k_1755767055_eval_466dUnroll-Qwen2.5-7B-Instruct-attn-2-ffn-0-nemo-bespoke-w-system-prompt-seq16k_1755767143_eval_466dUnroll-Qwen2.5-7B-Instruct-attn-3-ffn-0-nemo-bespoke-w-system-prompt-seq16k_1755767190_eval_466dUnroll-Qwen2.5-7B-Instruct-attn-0-ffn-0-nemo-bespoke-w-system-prompt-seq16k_1755767358_eval_e0d7Unroll-Qwen2.5-7B-Instruct-attn-1-ffn-0-nemo-bespoke-w-system-prompt-seq16k_1755767446_eval_e0d7Unroll-Qwen2.5-7B-Instruct-attn-3-ffn-0-nemo-bespoke-w-system-prompt-seq16k_1755767533_eval_e0d7Unroll-Qwen2.5-7B-Instruct-attn-3-ffn-0-nemo-bespoke-w-system-prompt-seq16k_1755767575_eval_e0d7Unroll-Qwen2.5-7B-Instruct-attn-4-ffn-0-nemo-bespoke-w-system-prompt-seq16k_1755767619_eval_e0d7Unroll-Qwen2.5-7B-Instruct-attn-5-ffn-0-nemo-bespoke-w-system-prompt-seq16k_1755767663_eval_e0d7
