datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eval-Qwen3-235B-A22B-reasoning
qwen-235b-a22-thinking Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.782
math_pass@1:64_samples
64
0.5%
aime25
0.718
math_pass@1:64_samples
64
0.1%
arenahard
0.939
eval/overall_winrate
500
0.0%
bbh_generative
0.884
extractive_match
1
0.0%
creative-writing-v3
0.775
creative_writing_score
96
0.0%
drop_generative_nous
0.903
drop_acc
1
0.0%
eqbench3
0.800
eqbench_score
135
0.0%
gpqa_diamond
0.697… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Qwen3-235B-A22B-reasoning.eval-Qwen3-14B-reasoning
14b-reasoning Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.776
math_pass@1:64_samples
64
0.6%
aime25
0.685
math_pass@1:64_samples
64
1.2%
arenahard
0.878
eval/overall_winrate
500
0.0%
bbh_generative
0.866
extractive_match
1
0.0%
creative-writing-v3
0.666
creative_writing_score
96
0.0%
drop_generative_nous
0.894
drop_acc
1
0.0%
eqbench3
0.748
eqbench_score
135
0.0%
gpqa_diamond
0.620
gpqa_pass@1:8_samples8
0.2%… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Qwen3-14B-reasoning.terminal-bench-2eval-Hermes-4-14B-reasoning
h4-14b-more-stage1-reasoning Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.554
math_pass@1:64_samples
64
0.1%
aime25
0.468
math_pass@1:64_samples
64
0.1%
arenahard
0.830
eval/overall_winrate
500
0.0%
bbh_generative
0.844
extractive_match
1
0.0%
creative-writing-v3
0.616
creative_writing_score
96
0.0%
drop_generative_nous
0.845
drop_acc
1
0.0%
eqbench3
0.772
eqbench_score
135
0.0%
gpqa_diamond
0.602… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-14B-reasoning.eval-Hermes-4-70B-nonreasoning
hermes-70b-nonreasoning Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.095
math_pass@1:64_samples
64
99.4%
aime25
0.073
math_pass@1:64_samples
64
98.2%
arenahard
0.568
eval/overall_winrate
500
0.0%
bbh_generative
0.805
extractive_match
1
100.0%
creative-writing-v3
0.491
creative_writing_score
96
0.0%
drop_generative_nous
0.784
drop_acc
1
100.0%
eqbench3
0.739
eqbench_score
135
0.0%
gpqa_diamond
0.333… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-70B-nonreasoning.eval-Hermes-4-405B-reasoning
405b-e3-40k-reasoning Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.819
math_pass@1:64_samples
64
5.6%
aime25
0.781
math_pass@1:64_samples
64
5.3%
arenahard
0.937
eval/overall_winrate
500
0.0%
bbh_generative
0.863
extractive_match
1
4.7%
creative-writing-v3
0.793
creative_writing_score
96
0.0%
drop_generative_nous
0.835
drop_acc
1
1.6%
eqbench3
0.855
eqbench_score
135
0.0%
gpqa_diamond
0.706… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-405B-reasoning.eval-DeepSeek-R1-0528
r1-0528 Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.865
math_pass@1:64_samples
64
0.0%
aime25
0.831
math_pass@1:64_samples
64
0.0%
arenahard
0.951
eval/overall_winrate
500
0.0%
bbh_generative
0.894
extractive_match
1
0.0%
creative-writing-v3
0.803
creative_writing_score
96
0.0%
drop_generative_nous
0.865
drop_acc
1
0.0%
eqbench3
0.865
eqbench_score
135
0.0%
gpqa_diamond
0.781
gpqa_pass@1:8_samples8
0.1%… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-DeepSeek-R1-0528.eval-Cogito-v2-preview-405B-nonreasoning
cogito-405b-thinking Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.177
math_pass@1:64_samples
64
100.0%
aime25
0.098
math_pass@1:64_samples
64
100.0%
arenahard
0.829
eval/overall_winrate
500
0.0%
bbh_generative
0.880
extractive_match
1
100.0%
creative-writing-v3
0.679
creative_writing_score
96
0.0%
drop_generative_nous
0.856
drop_acc
1
100.0%
eqbench3
0.695
eqbench_score
135
0.0%
gpqa_diamond
0.562… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Cogito-v2-preview-405B-nonreasoning.eval-Cogito-v2-preview-405B-reasoning
cogito-405b-thinking Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.408
math_pass@1:64_samples
64
20.3%
aime25
0.327
math_pass@1:64_samples
64
15.5%
arenahard
0.910
eval/overall_winrate
500
0.0%
bbh_generative
0.893
extractive_match
1
1.3%
creative-writing-v3
0.674
creative_writing_score
96
0.0%
drop_generative_nous
0.871
drop_acc
1
0.3%
eqbench3
0.672
eqbench_score
135
0.0%
gpqa_diamond
0.682… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Cogito-v2-preview-405B-reasoning.eval-Qwen3-14B-nonreasoning
qwen3-14b-reasoning Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.285
math_pass@1:64_samples
64
0.0%
aime25
0.222
math_pass@1:64_samples
64
0.0%
arenahard
0.796
eval/overall_winrate
500
0.0%
bbh_generative
0.825
extractive_match
1
0.0%
creative-writing-v3
0.516
creative_writing_score
96
0.0%
drop_generative_nous
0.750
drop_acc
1
0.0%
eqbench3
0.697
eqbench_score
135
0.0%
gpqa_diamond
0.535
gpqa_pass@1:8_samples8… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Qwen3-14B-nonreasoning.eval-Cogito-v2-preview-70B-reasoning
cogito-thinking Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.322
math_pass@1:64_samples
64
35.2%
aime25
0.221
math_pass@1:64_samples
64
33.3%
arenahard
0.869
eval/overall_winrate
500
0.0%
bbh_generative
0.893
extractive_match
1
2.9%
creative-writing-v3
0.636
creative_writing_score
96
0.0%
drop_generative_nous
0.860
drop_acc
1
0.8%
eqbench3
0.657
eqbench_score
135
0.0%
gpqa_diamond
0.591
gpqa_pass@1:8_samples8… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Cogito-v2-preview-70B-reasoning.eval-Hermes-4-405B-nonreasoning
h4-405b-e3-nonthinking Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.114
math_pass@1:64_samples
64
100.0%
aime25
0.106
math_pass@1:64_samples
64
100.0%
arenahard
0.535
eval/overall_winrate
500
0.0%
bbh_generative
0.687
extractive_match
1
100.0%
creative-writing-v3
0.506
creative_writing_score
96
0.0%
drop_generative_nous
0.776
drop_acc
1
100.0%
eqbench3
0.746
eqbench_score
135
0.0%
gpqa_diamond
0.394… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-405B-nonreasoning.eval-Cogito-v2-preview-70B-nonreasoning
cogito-70b-nonthinking Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.122
math_pass@1:64_samples
64
100.0%
aime25
0.060
math_pass@1:64_samples
64
100.0%
arenahard
0.819
eval/overall_winrate
500
0.0%
bbh_generative
0.876
extractive_match
1
100.0%
creative-writing-v3
0.655
creative_writing_score
96
0.0%
drop_generative_nous
0.841
drop_acc
1
100.0%
eqbench3
0.681
eqbench_score
135
0.0%
gpqa_diamond
0.528… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Cogito-v2-preview-70B-nonreasoning.eval-DeepSeek-V3-0324
dsv3 Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.506
math_pass@1:64_samples
64
100.0%
aime25
0.422
math_pass@1:64_samples
64
100.0%
arenahard
0.926
eval/overall_winrate
500
0.0%
bbh_generative
0.868
extractive_match
1
100.0%
creative-writing-v3
0.767
creative_writing_score
96
0.0%
drop_generative_nous
0.829
drop_acc
1
100.0%
eqbench3
0.831
eqbench_score
135
0.0%
gpqa_diamond
0.680
gpqa_pass@1:8_samples8
100.0%… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-DeepSeek-V3-0324.eval-Qwen3-235B-A22B-nonreasoning
qwen-235b-22a-nonreasoning Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.341
math_pass@1:64_samples
64
0.0%
aime25
0.251
math_pass@1:64_samples
64
0.0%
arenahard
0.917
eval/overall_winrate
500
0.0%
bbh_generative
0.860
extractive_match
1
0.0%
creative-writing-v3
0.741
creative_writing_score
96
0.0%
drop_generative_nous
0.794
drop_acc
1
0.0%
eqbench3
0.811
eqbench_score
135
0.0%
gpqa_diamond
0.577… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Qwen3-235B-A22B-nonreasoning.eval-Hermes-4-14B-reasoning-old
h4-e3-overlong-masked-30k-rerun Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.527
math_pass@1:64_samples
64
6.6%
aime25
0.414
math_pass@1:64_samples
64
8.1%
arenahard
0.782
eval/overall_winrate
500
0.0%
bbh_generative
0.844
extractive_match
1
5.8%
creative-writing-v3
0.617
creative_writing_score
96
0.0%
drop_generative_nous
0.827
drop_acc
1
2.4%
eqbench3
0.805
eqbench_score
135
0.0%
gpqa_diamond
0.556… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-14B-reasoning-old.eval-Hermes-4.3-36B
36bpsychev2 Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.719
math_pass@1:64_samples
64
17.6%
aime25
0.693
math_pass@1:64_samples
64
18.8%
bbh_generative
0.864
extractive_match
1
4.8%
drop_generative_nous
0.835
drop_acc
1
2.7%
gpqa_diamond
0.655
gpqa_pass@1:8_samples
8
2.2%
ifeval
0.779
inst_level_loose_acc
1
7.8%
math_500
0.938
math_pass@1:4_samples
4
1.8%
mmlu_generative
0.877
extractive_match1
0.1%
mmlu_pro… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4.3-36B.eval-Hermes-4-14B-nonreasoning-old
h4-14b-nonreasoning-30k-cot Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.105
math_pass@1:64_samples
64
99.7%
aime25
0.066
math_pass@1:64_samples
64
100.0%
arenahard
0.498
eval/overall_winrate
500
0.0%
bbh_generative
0.632
extractive_match
1
100.0%
creative-writing-v3
0.405
creative_writing_score
96
0.0%
drop_generative_nous
0.714
drop_acc
1
100.0%
eqbench3
0.690
eqbench_score
135
0.0%
gpqa_diamond
0.450… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-14B-nonreasoning-old.eval-Hermes-4-70B-reasoning
hermes-4-70b-reasoning-40k Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.735
math_pass@1:64_samples
64
8.4%
aime25
0.674
math_pass@1:64_samples
64
9.6%
arenahard
0.901
eval/overall_winrate
500
0.0%
bbh_generative
0.878
extractive_match
1
4.8%
creative-writing-v3
0.775
creative_writing_score
96
0.0%
drop_generative_nous
0.850
drop_acc
1
1.4%
eqbench3
0.847
eqbench_score
135
0.0%
gpqa_diamond
0.661… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-70B-reasoning.eval-Hermes-4-14B-nonreasoning
h4-14b-more-stage1-nonreasoning Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.110
math_pass@1:64_samples
64
99.9%
aime25
0.069
math_pass@1:64_samples
64
99.6%
arenahard
0.502
eval/overall_winrate
500
0.0%
bbh_generative
0.740
extractive_match
1
100.0%
creative-writing-v3
0.355
creative_writing_score
96
0.0%
drop_generative_nous
0.739
drop_acc
1
100.0%
eqbench3
0.580
eqbench_score
135
0.0%
gpqa_diamond
0.390… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-14B-nonreasoning.openthoughts-tblite
NousResearch/openthoughts-tblite
This dataset is a reformatted version of OpenThoughts-TBLite for use with the Hermes Agent Terminal-Bench evaluation framework.
Source
OpenThoughts-TBLite was created by the OpenThoughts Agent team in collaboration with Snorkel AI and Bespoke Labs. It is a difficulty-calibrated subset of Terminal-Bench 2.0 designed for faster iteration when developing terminal agents.
Original dataset: open-thoughts/OpenThoughts-TBLite
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/openthoughts-tblite.eval-Hermes-4.3-36B-centralized
36btorchtitan Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.706
math_pass@1:64_samples
64
24.6%
aime25
0.669
math_pass@1:64_samples
64
26.8%
bbh_generative
0.847
extractive_match
1
9.3%
drop_generative_nous
0.817
drop_acc
1
6.9%
gpqa_diamond
0.649
gpqa_pass@1:8_samples
8
4.2%
ifeval
0.740
inst_level_loose_acc
1
13.1%
math_500
0.923
math_pass@1:4_samples
4
3.1%
mmlu_generative
0.866
extractive_match1
1.4%… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4.3-36B-centralized.lm-eval-results-NousResearch-Hermes-2-Pro-Llama-3-8B-private
Dataset Card for Evaluation run of NousResearch/Hermes-2-Pro-Llama-3-8B
Dataset automatically created during the evaluation run of model NousResearch/Hermes-2-Pro-Llama-3-8B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-NousResearch-Hermes-2-Pro-Llama-3-8B-private.NousResearch__Nous-Hermes-2-Mistral-7B-DPO-details
Dataset Card for Evaluation run of NousResearch/Nous-Hermes-2-Mistral-7B-DPO
Dataset automatically created during the evaluation run of model NousResearch/Nous-Hermes-2-Mistral-7B-DPO
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Nous-Hermes-2-Mistral-7B-DPO-details.NousResearch__Hermes-3-Llama-3.2-3B-details
Dataset Card for Evaluation run of NousResearch/Hermes-3-Llama-3.2-3B
Dataset automatically created during the evaluation run of model NousResearch/Hermes-3-Llama-3.2-3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Hermes-3-Llama-3.2-3B-details.NousResearch__DeepHermes-3-Mistral-24B-Preview-details
Dataset Card for Evaluation run of NousResearch/DeepHermes-3-Mistral-24B-Preview
Dataset automatically created during the evaluation run of model NousResearch/DeepHermes-3-Mistral-24B-Preview
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__DeepHermes-3-Mistral-24B-Preview-details.NousResearch__Yarn-Mistral-7b-128k-details
Dataset Card for Evaluation run of NousResearch/Yarn-Mistral-7b-128k
Dataset automatically created during the evaluation run of model NousResearch/Yarn-Mistral-7b-128k
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Yarn-Mistral-7b-128k-details.NousResearch__Yarn-Solar-10b-64k-details
Dataset Card for Evaluation run of NousResearch/Yarn-Solar-10b-64k
Dataset automatically created during the evaluation run of model NousResearch/Yarn-Solar-10b-64k
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Yarn-Solar-10b-64k-details.NousResearch__Nous-Hermes-2-Mixtral-8x7B-DPO-details
Dataset Card for Evaluation run of NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO
Dataset automatically created during the evaluation run of model NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO
The dataset is composed of 39 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Nous-Hermes-2-Mixtral-8x7B-DPO-details.NousResearch__Hermes-2-Pro-Mistral-7B-details
Dataset Card for Evaluation run of NousResearch/Hermes-2-Pro-Mistral-7B
Dataset automatically created during the evaluation run of model NousResearch/Hermes-2-Pro-Mistral-7B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Hermes-2-Pro-Mistral-7B-details.
