CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NousResearch /eval-Qwen3-235B-A22B-reasoning qwen-235b-a22-thinking Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.782 math_pass@1:64_samples 64 0.5% aime25 0.718 math_pass@1:64_samples 64 0.1% arenahard 0.939 eval/overall_winrate 500 0.0% bbh_generative 0.884 extractive_match 1 0.0% creative-writing-v3 0.775 creative_writing_score 96 0.0% drop_generative_nous 0.903 drop_acc 1 0.0% eqbench3 0.800 eqbench_score 135 0.0% gpqa_diamond 0.697… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Qwen3-235B-A22B-reasoning.tabular100K<n<1M3 likes465 downloads1y agoHugging Face02NousResearch /eval-Qwen3-14B-reasoning 14b-reasoning Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.776 math_pass@1:64_samples 64 0.6% aime25 0.685 math_pass@1:64_samples 64 1.2% arenahard 0.878 eval/overall_winrate 500 0.0% bbh_generative 0.866 extractive_match 1 0.0% creative-writing-v3 0.666 creative_writing_score 96 0.0% drop_generative_nous 0.894 drop_acc 1 0.0% eqbench3 0.748 eqbench_score 135 0.0% gpqa_diamond 0.620 gpqa_pass@1:8_samples8 0.2%… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Qwen3-14B-reasoning.tabular100K<n<1M2 likes425 downloads1y agoHugging Face03NousResearch /terminal-bench-2tabularn<1K4 likes402 downloads8mo agoHugging Face04NousResearch /eval-Hermes-4-14B-reasoning h4-14b-more-stage1-reasoning Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.554 math_pass@1:64_samples 64 0.1% aime25 0.468 math_pass@1:64_samples 64 0.1% arenahard 0.830 eval/overall_winrate 500 0.0% bbh_generative 0.844 extractive_match 1 0.0% creative-writing-v3 0.616 creative_writing_score 96 0.0% drop_generative_nous 0.845 drop_acc 1 0.0% eqbench3 0.772 eqbench_score 135 0.0% gpqa_diamond 0.602… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-14B-reasoning.tabular100K<n<1M2 likes379 downloads1y agoHugging Face05NousResearch /eval-Hermes-4-70B-nonreasoning hermes-70b-nonreasoning Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.095 math_pass@1:64_samples 64 99.4% aime25 0.073 math_pass@1:64_samples 64 98.2% arenahard 0.568 eval/overall_winrate 500 0.0% bbh_generative 0.805 extractive_match 1 100.0% creative-writing-v3 0.491 creative_writing_score 96 0.0% drop_generative_nous 0.784 drop_acc 1 100.0% eqbench3 0.739 eqbench_score 135 0.0% gpqa_diamond 0.333… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-70B-nonreasoning.tabular100K<n<1M3 likes374 downloads1y agoHugging Face06NousResearch /eval-Hermes-4-405B-reasoning 405b-e3-40k-reasoning Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.819 math_pass@1:64_samples 64 5.6% aime25 0.781 math_pass@1:64_samples 64 5.3% arenahard 0.937 eval/overall_winrate 500 0.0% bbh_generative 0.863 extractive_match 1 4.7% creative-writing-v3 0.793 creative_writing_score 96 0.0% drop_generative_nous 0.835 drop_acc 1 1.6% eqbench3 0.855 eqbench_score 135 0.0% gpqa_diamond 0.706… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-405B-reasoning.tabular100K<n<1M9 likes358 downloads1y agoHugging Face07NousResearch /eval-DeepSeek-R1-0528 r1-0528 Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.865 math_pass@1:64_samples 64 0.0% aime25 0.831 math_pass@1:64_samples 64 0.0% arenahard 0.951 eval/overall_winrate 500 0.0% bbh_generative 0.894 extractive_match 1 0.0% creative-writing-v3 0.803 creative_writing_score 96 0.0% drop_generative_nous 0.865 drop_acc 1 0.0% eqbench3 0.865 eqbench_score 135 0.0% gpqa_diamond 0.781 gpqa_pass@1:8_samples8 0.1%… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-DeepSeek-R1-0528.tabular100K<n<1M2 likes344 downloads1y agoHugging Face08NousResearch /eval-Cogito-v2-preview-405B-nonreasoning cogito-405b-thinking Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.177 math_pass@1:64_samples 64 100.0% aime25 0.098 math_pass@1:64_samples 64 100.0% arenahard 0.829 eval/overall_winrate 500 0.0% bbh_generative 0.880 extractive_match 1 100.0% creative-writing-v3 0.679 creative_writing_score 96 0.0% drop_generative_nous 0.856 drop_acc 1 100.0% eqbench3 0.695 eqbench_score 135 0.0% gpqa_diamond 0.562… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Cogito-v2-preview-405B-nonreasoning.tabular100K<n<1M2 likes309 downloads1y agoHugging Face09NousResearch /eval-Cogito-v2-preview-405B-reasoning cogito-405b-thinking Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.408 math_pass@1:64_samples 64 20.3% aime25 0.327 math_pass@1:64_samples 64 15.5% arenahard 0.910 eval/overall_winrate 500 0.0% bbh_generative 0.893 extractive_match 1 1.3% creative-writing-v3 0.674 creative_writing_score 96 0.0% drop_generative_nous 0.871 drop_acc 1 0.3% eqbench3 0.672 eqbench_score 135 0.0% gpqa_diamond 0.682… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Cogito-v2-preview-405B-reasoning.tabular100K<n<1M3 likes304 downloads1y agoHugging Face10NousResearch /eval-Qwen3-14B-nonreasoning qwen3-14b-reasoning Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.285 math_pass@1:64_samples 64 0.0% aime25 0.222 math_pass@1:64_samples 64 0.0% arenahard 0.796 eval/overall_winrate 500 0.0% bbh_generative 0.825 extractive_match 1 0.0% creative-writing-v3 0.516 creative_writing_score 96 0.0% drop_generative_nous 0.750 drop_acc 1 0.0% eqbench3 0.697 eqbench_score 135 0.0% gpqa_diamond 0.535 gpqa_pass@1:8_samples8… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Qwen3-14B-nonreasoning.tabular100K<n<1M3 likes279 downloads1y agoHugging Face11NousResearch /eval-Cogito-v2-preview-70B-reasoning cogito-thinking Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.322 math_pass@1:64_samples 64 35.2% aime25 0.221 math_pass@1:64_samples 64 33.3% arenahard 0.869 eval/overall_winrate 500 0.0% bbh_generative 0.893 extractive_match 1 2.9% creative-writing-v3 0.636 creative_writing_score 96 0.0% drop_generative_nous 0.860 drop_acc 1 0.8% eqbench3 0.657 eqbench_score 135 0.0% gpqa_diamond 0.591 gpqa_pass@1:8_samples8… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Cogito-v2-preview-70B-reasoning.tabular100K<n<1M2 likes278 downloads1y agoHugging Face12NousResearch /eval-Hermes-4-405B-nonreasoning h4-405b-e3-nonthinking Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.114 math_pass@1:64_samples 64 100.0% aime25 0.106 math_pass@1:64_samples 64 100.0% arenahard 0.535 eval/overall_winrate 500 0.0% bbh_generative 0.687 extractive_match 1 100.0% creative-writing-v3 0.506 creative_writing_score 96 0.0% drop_generative_nous 0.776 drop_acc 1 100.0% eqbench3 0.746 eqbench_score 135 0.0% gpqa_diamond 0.394… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-405B-nonreasoning.tabular100K<n<1M4 likes264 downloads1y agoHugging Face13NousResearch /eval-Cogito-v2-preview-70B-nonreasoning cogito-70b-nonthinking Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.122 math_pass@1:64_samples 64 100.0% aime25 0.060 math_pass@1:64_samples 64 100.0% arenahard 0.819 eval/overall_winrate 500 0.0% bbh_generative 0.876 extractive_match 1 100.0% creative-writing-v3 0.655 creative_writing_score 96 0.0% drop_generative_nous 0.841 drop_acc 1 100.0% eqbench3 0.681 eqbench_score 135 0.0% gpqa_diamond 0.528… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Cogito-v2-preview-70B-nonreasoning.tabular100K<n<1M2 likes251 downloads1y agoHugging Face14NousResearch /eval-DeepSeek-V3-0324 dsv3 Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.506 math_pass@1:64_samples 64 100.0% aime25 0.422 math_pass@1:64_samples 64 100.0% arenahard 0.926 eval/overall_winrate 500 0.0% bbh_generative 0.868 extractive_match 1 100.0% creative-writing-v3 0.767 creative_writing_score 96 0.0% drop_generative_nous 0.829 drop_acc 1 100.0% eqbench3 0.831 eqbench_score 135 0.0% gpqa_diamond 0.680 gpqa_pass@1:8_samples8 100.0%… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-DeepSeek-V3-0324.tabular100K<n<1M2 likes230 downloads1y agoHugging Face15NousResearch /eval-Qwen3-235B-A22B-nonreasoning qwen-235b-22a-nonreasoning Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.341 math_pass@1:64_samples 64 0.0% aime25 0.251 math_pass@1:64_samples 64 0.0% arenahard 0.917 eval/overall_winrate 500 0.0% bbh_generative 0.860 extractive_match 1 0.0% creative-writing-v3 0.741 creative_writing_score 96 0.0% drop_generative_nous 0.794 drop_acc 1 0.0% eqbench3 0.811 eqbench_score 135 0.0% gpqa_diamond 0.577… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Qwen3-235B-A22B-nonreasoning.tabular100K<n<1M2 likes224 downloads1y agoHugging Face16NousResearch /eval-Hermes-4-14B-reasoning-old h4-e3-overlong-masked-30k-rerun Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.527 math_pass@1:64_samples 64 6.6% aime25 0.414 math_pass@1:64_samples 64 8.1% arenahard 0.782 eval/overall_winrate 500 0.0% bbh_generative 0.844 extractive_match 1 5.8% creative-writing-v3 0.617 creative_writing_score 96 0.0% drop_generative_nous 0.827 drop_acc 1 2.4% eqbench3 0.805 eqbench_score 135 0.0% gpqa_diamond 0.556… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-14B-reasoning-old.tabular100K<n<1M3 likes218 downloads1y agoHugging Face17NousResearch /eval-Hermes-4.3-36B 36bpsychev2 Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.719 math_pass@1:64_samples 64 17.6% aime25 0.693 math_pass@1:64_samples 64 18.8% bbh_generative 0.864 extractive_match 1 4.8% drop_generative_nous 0.835 drop_acc 1 2.7% gpqa_diamond 0.655 gpqa_pass@1:8_samples 8 2.2% ifeval 0.779 inst_level_loose_acc 1 7.8% math_500 0.938 math_pass@1:4_samples 4 1.8% mmlu_generative 0.877 extractive_match1 0.1% mmlu_pro… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4.3-36B.tabular100K<n<1M5 likes192 downloads10mo agoHugging Face18NousResearch /eval-Hermes-4-14B-nonreasoning-old h4-14b-nonreasoning-30k-cot Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.105 math_pass@1:64_samples 64 99.7% aime25 0.066 math_pass@1:64_samples 64 100.0% arenahard 0.498 eval/overall_winrate 500 0.0% bbh_generative 0.632 extractive_match 1 100.0% creative-writing-v3 0.405 creative_writing_score 96 0.0% drop_generative_nous 0.714 drop_acc 1 100.0% eqbench3 0.690 eqbench_score 135 0.0% gpqa_diamond 0.450… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-14B-nonreasoning-old.tabular100K<n<1M3 likes191 downloads1y agoHugging Face19NousResearch /eval-Hermes-4-70B-reasoning hermes-4-70b-reasoning-40k Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.735 math_pass@1:64_samples 64 8.4% aime25 0.674 math_pass@1:64_samples 64 9.6% arenahard 0.901 eval/overall_winrate 500 0.0% bbh_generative 0.878 extractive_match 1 4.8% creative-writing-v3 0.775 creative_writing_score 96 0.0% drop_generative_nous 0.850 drop_acc 1 1.4% eqbench3 0.847 eqbench_score 135 0.0% gpqa_diamond 0.661… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-70B-reasoning.tabular100K<n<1M5 likes170 downloads1y agoHugging Face20NousResearch /eval-Hermes-4-14B-nonreasoning h4-14b-more-stage1-nonreasoning Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.110 math_pass@1:64_samples 64 99.9% aime25 0.069 math_pass@1:64_samples 64 99.6% arenahard 0.502 eval/overall_winrate 500 0.0% bbh_generative 0.740 extractive_match 1 100.0% creative-writing-v3 0.355 creative_writing_score 96 0.0% drop_generative_nous 0.739 drop_acc 1 100.0% eqbench3 0.580 eqbench_score 135 0.0% gpqa_diamond 0.390… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-14B-nonreasoning.tabular100K<n<1M2 likes169 downloads1y agoHugging Face21NousResearch /openthoughts-tblite NousResearch/openthoughts-tblite This dataset is a reformatted version of OpenThoughts-TBLite for use with the Hermes Agent Terminal-Bench evaluation framework. Source OpenThoughts-TBLite was created by the OpenThoughts Agent team in collaboration with Snorkel AI and Bespoke Labs. It is a difficulty-calibrated subset of Terminal-Bench 2.0 designed for faster iteration when developing terminal agents. Original dataset: open-thoughts/OpenThoughts-TBLite GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/openthoughts-tblite.tabularn<1K11 likes156 downloads7mo agoHugging Face22NousResearch /eval-Hermes-4.3-36B-centralized 36btorchtitan Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.706 math_pass@1:64_samples 64 24.6% aime25 0.669 math_pass@1:64_samples 64 26.8% bbh_generative 0.847 extractive_match 1 9.3% drop_generative_nous 0.817 drop_acc 1 6.9% gpqa_diamond 0.649 gpqa_pass@1:8_samples 8 4.2% ifeval 0.740 inst_level_loose_acc 1 13.1% math_500 0.923 math_pass@1:4_samples 4 3.1% mmlu_generative 0.866 extractive_match1 1.4%… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4.3-36B-centralized.tabular100K<n<1M3 likes151 downloads10mo agoHugging Face23nyu-dice-lab /lm-eval-results-NousResearch-Hermes-2-Pro-Llama-3-8B-private Dataset Card for Evaluation run of NousResearch/Hermes-2-Pro-Llama-3-8B Dataset automatically created during the evaluation run of model NousResearch/Hermes-2-Pro-Llama-3-8B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-NousResearch-Hermes-2-Pro-Llama-3-8B-private.tabular100K<n<1M0 likes142 downloads2y agoHugging Face24open-llm-leaderboard /NousResearch__Nous-Hermes-2-Mistral-7B-DPO-detailsgated Dataset Card for Evaluation run of NousResearch/Nous-Hermes-2-Mistral-7B-DPO Dataset automatically created during the evaluation run of model NousResearch/Nous-Hermes-2-Mistral-7B-DPO The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Nous-Hermes-2-Mistral-7B-DPO-details.tabular10K<n<100K0 likes62 downloads2y agoHugging Face25open-llm-leaderboard /NousResearch__Hermes-3-Llama-3.2-3B-detailsgated Dataset Card for Evaluation run of NousResearch/Hermes-3-Llama-3.2-3B Dataset automatically created during the evaluation run of model NousResearch/Hermes-3-Llama-3.2-3B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Hermes-3-Llama-3.2-3B-details.tabular10K<n<100K0 likes61 downloads2y agoHugging Face26open-llm-leaderboard /NousResearch__DeepHermes-3-Mistral-24B-Preview-detailsgated Dataset Card for Evaluation run of NousResearch/DeepHermes-3-Mistral-24B-Preview Dataset automatically created during the evaluation run of model NousResearch/DeepHermes-3-Mistral-24B-Preview The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__DeepHermes-3-Mistral-24B-Preview-details.tabular10K<n<100K0 likes55 downloads2y agoHugging Face27open-llm-leaderboard /NousResearch__Yarn-Mistral-7b-128k-detailsgated Dataset Card for Evaluation run of NousResearch/Yarn-Mistral-7b-128k Dataset automatically created during the evaluation run of model NousResearch/Yarn-Mistral-7b-128k The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Yarn-Mistral-7b-128k-details.tabular10K<n<100K0 likes32 downloads2y agoHugging Face28open-llm-leaderboard /NousResearch__Yarn-Solar-10b-64k-detailsgated Dataset Card for Evaluation run of NousResearch/Yarn-Solar-10b-64k Dataset automatically created during the evaluation run of model NousResearch/Yarn-Solar-10b-64k The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Yarn-Solar-10b-64k-details.tabular10K<n<100K0 likes32 downloads2y agoHugging Face29open-llm-leaderboard /NousResearch__Nous-Hermes-2-Mixtral-8x7B-DPO-detailsgated Dataset Card for Evaluation run of NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO Dataset automatically created during the evaluation run of model NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO The dataset is composed of 39 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Nous-Hermes-2-Mixtral-8x7B-DPO-details.tabular10K<n<100K0 likes31 downloads2y agoHugging Face30open-llm-leaderboard /NousResearch__Hermes-2-Pro-Mistral-7B-detailsgated Dataset Card for Evaluation run of NousResearch/Hermes-2-Pro-Mistral-7B Dataset automatically created during the evaluation run of model NousResearch/Hermes-2-Pro-Mistral-7B The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Hermes-2-Pro-Mistral-7B-details.tabular10K<n<100K0 likes30 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.