datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dclm-pool-7b-2xdclm-pool-7b-1xdolma3_mix-6T-1025-7B
⚠️ WARNING: This dataset is intended ONLY for reproducing Olmo 3 7B ⚠️
For all other training use cases, including training from scratch, please utilize our primary dolma 3 data mix: https://huggingface.co/datasets/allenai/dolma3_mix-6T.
Note: Some olmOCR science PDFs in the current dataset have been redacted following the training of Olmo 3 7B. These texts are indicated with [REMOVED] in the text field. This will affect reproducibility of Olmo 3 7B.
For this reason, please use… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-6T-1025-7B.Transmem_ecsd_qwen2_5_7b_hotpotqa_n4_n8Eurus-2-7B-SFT_eval_2e29
mlfoundations-dev/Eurus-2-7B-SFT_eval_2e29
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
Accuracy
2.3
21.0
30.6
11.0
11.4
10.4
6.8
1.5
2.1
1.3
4.1
4.4
AIME24
Average Accuracy: 2.33% ± 0.67%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
0.00%
0
30
2
3.33%
1… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Eurus-2-7B-SFT_eval_2e29.CerebRM-olmo-3-7b-instruct-sft-list_em-so1_completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the… See the full description on the dataset page: https://huggingface.co/datasets/wetsoledrysoul/CerebRM-olmo-3-7b-instruct-sft-list_em-so1_completions.Llama-2-7b-KronQ-HG
Llama-2-7b — KronQ H_G (output-side gradient covariance)
Paper: arXiv:2607.07964 · Code: GitHub
Pre-computed H_G for Llama-2-7b, the output-side curvature factor used by KronQ under the K-FAC factorization H ≈ H_X ⊗ H_G. H_G is the per-sublayer sampled-Fisher gradient covariance (labels drawn from the model distribution) (E[g gᵀ] over the layer output), distinct from the standard input-side Hessian H_X (which GPTQ/GPTAQ build online during calibration).
Publishing this lets you… See the full description on the dataset page: https://huggingface.co/datasets/donghyunli/Llama-2-7b-KronQ-HG.details_CohereForAI__c4ai-command-r7b-arabic-02-2025_v2
Dataset Card for Evaluation run of CohereForAI/c4ai-command-r7b-arabic-02-2025
Dataset automatically created during the evaluation run of model CohereForAI/c4ai-command-r7b-arabic-02-2025.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_CohereForAI__c4ai-command-r7b-arabic-02-2025_v2.distill_qwen_7b_math_trainDolci-Think-SFT-7B
Dolci-Think-SFT
Sources include a mixture of existing reasoning traces:
OpenThoughts 3 (Apache 2.0): Extended to 32K context length and downsampled code prompts to 16X multiple, to 941,166 total prompts. Access our version, Dolci OpenThoughts 3 here.
SYNTHETIC-2 (Apache 2.0) via the SFT-Verified split, 104,569 prompts.
Nemotron Post-training dataset (CC BY 4), code split only, 113,777 prompts.
New prompts and new reasoning traces from us (all ODC-BY-1.0):
Dolci Think Persona… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Think-SFT-7B.details_meta-llama__Llama-2-7b-hf
Dataset Card for Evaluation run of meta-llama/Llama-2-7b-hf
Dataset Summary
Dataset automatically created during the evaluation run of model meta-llama/Llama-2-7b-hf on the Open LLM Leaderboard.
The dataset is composed of 127 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 16 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_meta-llama__Llama-2-7b-hf.Qwen2.5-7B-BaldEagle-Ultrachatdistill_qwen_7b_aime_verifications_7b_ft_verifierllama2_7b_chat-boolq-results
Dataset Card for "llama2_7b_chat-boolq-results"
More Information needed
llama2_7b_chat-piqa-resultsotto-taxonomy-sdg-mistral-7b-instruct-v0.3details_declare-lab__starling-7B
Dataset Card for Evaluation run of declare-lab/starling-7B
Dataset automatically created during the evaluation run of model declare-lab/starling-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_declare-lab__starling-7B.DeepSeek-R1-Distill-Qwen-7B_eval_d81a
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_d81a
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
MMLUPro
HMMT
HLE
AIME25
LiveCodeBenchv5
Accuracy
43.4
25.0
12.4
36.0
34.5
MMLUPro
Accuracy: 43.38%
Accuracy
Questions Solved
Total Questions
43.38%
N/A
N/A
HMMT
Average Accuracy: 25.00% ± 1.72%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_d81a.OpenReasoning-Nemotron-7B_eval_8179
mlfoundations-dev/OpenReasoning-Nemotron-7B_eval_8179
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
79.0
98.8
89.0
81.7
60.1
62.5
50.6
46.8
68.7
13.3
49.6
59.7
AIME24
Average Accuracy: 79.00% ± 1.42%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
70.00%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/OpenReasoning-Nemotron-7B_eval_8179.details_PocketDoc__Dans-TotSirocco-7b
Dataset Card for Evaluation run of PocketDoc/Dans-TotSirocco-7b
Dataset Summary
Dataset automatically created during the evaluation run of model PocketDoc/Dans-TotSirocco-7b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_PocketDoc__Dans-TotSirocco-7b.details_togethercomputer__RedPajama-INCITE-7B-Base
Dataset Card for Evaluation run of togethercomputer/RedPajama-INCITE-7B-Base
Dataset Summary
Dataset automatically created during the evaluation run of model togethercomputer/RedPajama-INCITE-7B-Base on the Open LLM Leaderboard.
The dataset is composed of 122 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_togethercomputer__RedPajama-INCITE-7B-Base.distill_qwen_7b_math_verifications_7b_ft_verifierdetails_Intel__neural-chat-7b-v3-1
Dataset Card for Evaluation run of Intel/neural-chat-7b-v3-1
Dataset Summary
Dataset automatically created during the evaluation run of model Intel/neural-chat-7b-v3-1 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Intel__neural-chat-7b-v3-1.obelics_100k-tokenized-4image_llava_vicuna-7B_4096rollouts-olmo7b-cue-search
rollouts-olmo7b-cue-search
Model: allenai/Olmo-3-1025-7B (snapshot a81bae42).
Tokenizer: allenai/Olmo-3-1025-7B (snapshot a81bae42).
Protocol: RL-Zero prompt, MATH-500 x 4 rollouts, budget 31,744, T 0.6, top-p 0.95, seed 20260819 (depth-2 exhaustive and n-gram chain: seed 20260821); the top-20 beam nominee screen, ten random-opener arms, every depth-2 opener (84 shards, arm names unique across shards) and the n-gram chain arms.
Rollouts generated on the CSAIL cluster for the… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-cues/rollouts-olmo7b-cue-search.theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results
MO_evals results
Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per
upload; nothing here is aggregated — the Parquet and the .eval logs are the primary
evidence, the scorecard is a summary of them.
<upload>/ results tree, as uploaded
<persona>__<method>__scale<n>__<fam>/ one organism
<spec_hash>/ one seed of it
spec.json spec_hash -> persona… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results.DeepSeek-R1-Distill-Qwen-7B_eval_118b
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_118b
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBenchv5_official
Average Accuracy: 31.18% ± nan%
Number of Runs: 1
Run
Accuracy
Questions Solved
Total Questions
1
31.18%
87
279
Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179
mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
60.7
90.8
89.4
63.2
52.4
48.5
27.4
26.2
48.3
12.0
34.3
34.7
AIME24
Average Accuracy: 60.67% ± 2.25%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179.d614ed7bdetails_maywell__Synatra-7B-v0.3-RP
Dataset Card for Evaluation run of maywell/Synatra-7B-v0.3-RP
Dataset automatically created during the evaluation run of model maywell/Synatra-7B-v0.3-RP.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_maywell__Synatra-7B-v0.3-RP.
