datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
po_qwen14b_tabular_data
BoLT Prompt Optimization — Tabular Dataset
For prompt optimization tasks in BoLT, an accessible benchmark for black-box optimization on LLM tasks.
Dataset Description
The dataset covers 5,014 evaluated instructions. Each row is a candidate system-prompt instruction paired with its empirically measured MATH-500 (4-shot, non-thinking mode) scores.
Evaluation details:
Model: Qwen/Qwen3-14B
Task: minerva_math500 (4-shot) (from lm-eval library)
System prompt:… See the full description on the dataset page: https://huggingface.co/datasets/chewwt/po_qwen14b_tabular_data.Transmem_ecsd_qwen2_5_14b_hotpotqa_n4_n8Transmem_ecsd_qwen3_14b_hotpotqa_n4_n8loracle-pretrain-v5-qwen14b-tokenslatent-mas-safety-dataset-seq-qwen3-14bKeural-MoE-14B-stage1-Datasettheo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results
MO_evals results
Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per
upload; nothing here is aggregated — the Parquet and the .eval logs are the primary
evidence, the scorecard is a summary of them.
<upload>/ results tree, as uploaded
<persona>__<method>__scale<n>__<fam>/ one organism
<spec_hash>/ one seed of it
spec.json spec_hash -> persona… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results.details_Qwen__Qwen1.5-14B-Chat
Dataset Card for Evaluation run of Qwen/Qwen1.5-14B-Chat
Dataset automatically created during the evaluation run of model Qwen/Qwen1.5-14B-Chat.
The dataset is composed of 117 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/amztheory/details_Qwen__Qwen1.5-14B-Chat.details_Qwen__Qwen3-14B_v2
Dataset Card for Evaluation run of Qwen/Qwen3-14B
Dataset automatically created during the evaluation run of model Qwen/Qwen3-14B.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen3-14B_v2.details_Azure99__blossom-v4-qwen1_5-14b
Dataset Card for Evaluation run of Azure99/blossom-v4-qwen1_5-14b
Dataset automatically created during the evaluation run of model Azure99/blossom-v4-qwen1_5-14b on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Azure99__blossom-v4-qwen1_5-14b.details_Azure99__blossom-v5-14b
Dataset Card for Evaluation run of Azure99/blossom-v5-14b
Dataset automatically created during the evaluation run of model Azure99/blossom-v5-14b on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Azure99__blossom-v5-14b.aime_1983_2023_deepseek-r1-distill-qwen-14b_traces_32768OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-topk128-normalizedc4-rewritten-14b-retok-smollm360mdclm-14b-c4-rewritten-14b-retok-smollm360meval-Qwen3-14B-reasoning
14b-reasoning Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.776
math_pass@1:64_samples
64
0.6%
aime25
0.685
math_pass@1:64_samples
64
1.2%
arenahard
0.878
eval/overall_winrate
500
0.0%
bbh_generative
0.866
extractive_match
1
0.0%
creative-writing-v3
0.666
creative_writing_score
96
0.0%
drop_generative_nous
0.894
drop_acc
1
0.0%
eqbench3
0.748
eqbench_score
135
0.0%
gpqa_diamond
0.620
gpqa_pass@1:8_samples8
0.2%… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Qwen3-14B-reasoning.eval-Hermes-4-14B-reasoning
h4-14b-more-stage1-reasoning Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.554
math_pass@1:64_samples
64
0.1%
aime25
0.468
math_pass@1:64_samples
64
0.1%
arenahard
0.830
eval/overall_winrate
500
0.0%
bbh_generative
0.844
extractive_match
1
0.0%
creative-writing-v3
0.616
creative_writing_score
96
0.0%
drop_generative_nous
0.845
drop_acc
1
0.0%
eqbench3
0.772
eqbench_score
135
0.0%
gpqa_diamond
0.602… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-14B-reasoning.loracle-ia-14b-direction-tokensOpenThoughts-114k-math-correct-qwen3-14b-math-prepared-normalizeddclm-14b-c4-rewrriten-table-prompt-14b-dpoed-retokdclm-14b-c4-rewrriten-table-prompt-14b-retokWan2.2-Animate-14B-COPY
Wan2.2
💜 Wan | 🖥️ GitHub | 🤗 Hugging Face | 🤖 ModelScope | 📑 Paper | 📑 Blog | 💬 Discord
📕 使用指南(中文) | 📘 User Guide(English) | 💬 WeChat(微信)
Wan: Open and Advanced Large-Scale Video Generative Models
We are excited to introduce Wan2.2, a major upgrade to our foundational video models. With Wan2.2, we have focused on incorporating the following innovations:
👍 Effective MoE Architecture: Wan2.2… See the full description on the dataset page: https://huggingface.co/datasets/Bohdanio2408/Wan2.2-Animate-14B-COPY.OpenThoughts-114k-math-correct-qwen3-14b-math-preparedeval-Qwen3-14B-nonreasoning
qwen3-14b-reasoning Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.285
math_pass@1:64_samples
64
0.0%
aime25
0.222
math_pass@1:64_samples
64
0.0%
arenahard
0.796
eval/overall_winrate
500
0.0%
bbh_generative
0.825
extractive_match
1
0.0%
creative-writing-v3
0.516
creative_writing_score
96
0.0%
drop_generative_nous
0.750
drop_acc
1
0.0%
eqbench3
0.697
eqbench_score
135
0.0%
gpqa_diamond
0.535
gpqa_pass@1:8_samples8… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Qwen3-14B-nonreasoning.details_01-ZeroOne__SUHAIL-14B-preview_v2
Dataset Card for Evaluation run of 01-ZeroOne/SUHAIL-14B-preview
Dataset automatically created during the evaluation run of model 01-ZeroOne/SUHAIL-14B-preview.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_01-ZeroOne__SUHAIL-14B-preview_v2.details_Qwen__Qwen3-14B-Base_v2
Dataset Card for Evaluation run of Qwen/Qwen3-14B-Base
Dataset automatically created during the evaluation run of model Qwen/Qwen3-14B-Base.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen3-14B-Base_v2.details_Qwen__Qwen2.5-14B_v2
Dataset Card for Evaluation run of Qwen/Qwen2.5-14B
Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-14B.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen2.5-14B_v2.eval-Hermes-4-14B-reasoning-old
h4-e3-overlong-masked-30k-rerun Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.527
math_pass@1:64_samples
64
6.6%
aime25
0.414
math_pass@1:64_samples
64
8.1%
arenahard
0.782
eval/overall_winrate
500
0.0%
bbh_generative
0.844
extractive_match
1
5.8%
creative-writing-v3
0.617
creative_writing_score
96
0.0%
drop_generative_nous
0.827
drop_acc
1
2.4%
eqbench3
0.805
eqbench_score
135
0.0%
gpqa_diamond
0.556… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-Hermes-4-14B-reasoning-old.dbbench-distilled-qwen3-14b-multiturnappworld-rollouts-recursive-14b-mar06
