datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sorrel-T-qwen3-14b-base-seed0-documentsministral-14b-eval-logs-and-scoresgPRM-14B-test_qwen
Reward of test_qwen split extracted by gPRM-14B: gPRM-14B-test_qwen
Usage
from datasets import load_dataset
# Load entire dataset
dataset = load_dataset("dongboklee/gPRM-14B-test_qwen")
# Load specific domain
law_dataset = load_dataset("dongboklee/gPRM-14B-test_qwen", split="law")
lm-eval-results-TomGrc-FusionNet_7Bx2_MoE_14B-private
Dataset Card for Evaluation run of TomGrc/FusionNet_7Bx2_MoE_14B
Dataset automatically created during the evaluation run of model TomGrc/FusionNet_7Bx2_MoE_14B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-TomGrc-FusionNet_7Bx2_MoE_14B-private.dORM-14B-test_gemma
Reward of test_gemma split extracted by dORM-14B: dORM-14B-test_gemma
Usage
from datasets import load_dataset
# Load entire dataset
dataset = load_dataset("dongboklee/dORM-14B-test_gemma")
# Load specific domain
law_dataset = load_dataset("dongboklee/dORM-14B-test_gemma", split="law")
gPRM-14B-test_llama
Reward of test_llama split extracted by gPRM-14B: gPRM-14B-test_llama
Usage
from datasets import load_dataset
# Load entire dataset
dataset = load_dataset("dongboklee/gPRM-14B-test_llama")
# Load specific domain
law_dataset = load_dataset("dongboklee/gPRM-14B-test_llama", split="law")
dORM-14B-test_qwen
Reward of test_qwen split extracted by dORM-14B: dORM-14B-test_qwen
Usage
from datasets import load_dataset
# Load entire dataset
dataset = load_dataset("dongboklee/dORM-14B-test_qwen")
# Load specific domain
law_dataset = load_dataset("dongboklee/dORM-14B-test_qwen", split="law")
gORM-14B-test_llama
Reward of test_llama split extracted by dORM-14B: dORM-14B-test_llama
Usage
from datasets import load_dataset
# Load entire dataset
dataset = load_dataset("dongboklee/dORM-14B-test_llama")
# Load specific domain
law_dataset = load_dataset("dongboklee/dORM-14B-test_llama", split="law")
suayptalha__Lix-14B-v0.1-details
Dataset Card for Evaluation run of suayptalha/Lix-14B-v0.1
Dataset automatically created during the evaluation run of model suayptalha/Lix-14B-v0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/suayptalha__Lix-14B-v0.1-details.dORM-14B-test_llama
Reward of test_llama split extracted by dORM-14B: dORM-14B-test_llama
Usage
from datasets import load_dataset
# Load entire dataset
dataset = load_dataset("dongboklee/dORM-14B-test_llama")
# Load specific domain
law_dataset = load_dataset("dongboklee/dORM-14B-test_llama", split="law")
dPRM-14B-test
Reward of test split extracted by dPRM-14B: dPRM-14B-test
Usage
from datasets import load_dataset
# Load entire dataset
dataset = load_dataset("dongboklee/dPRM-14B-test")
# Load specific domain
law_dataset = load_dataset("dongboklee/dPRM-14B-test", split="law")
JungZoona__T3Q-qwen2.5-14b-v1.0-e3-details
Dataset Card for Evaluation run of JungZoona/T3Q-qwen2.5-14b-v1.0-e3
Dataset automatically created during the evaluation run of model JungZoona/T3Q-qwen2.5-14b-v1.0-e3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/JungZoona__T3Q-qwen2.5-14b-v1.0-e3-details.prithivMLmods__Gaea-Opus-14B-Exp-details
Dataset Card for Evaluation run of prithivMLmods/Gaea-Opus-14B-Exp
Dataset automatically created during the evaluation run of model prithivMLmods/Gaea-Opus-14B-Exp
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Gaea-Opus-14B-Exp-details.Info_Wan_Video_2.2_T2V-A14B
Model Index by Creator
423748 Page
Model
Base Model
Full Model Page
Archive Link
wan2.2,t2v,low,zzzyixuan.
Wan Video 2.2 T2V-A14B
View
View
Version Links
Model
Version
Base Model
Version Link
wan2.2,t2v,low,zzzyixuan.
v1.0
Wan Video 2.2 T2V-A14B
View
Aaron_PP Page
Model
Base Model
Full Model Page
Archive Link
NSFW WAN 2.2 T2V Bunny girl, red patent leather tights, black high stockings, red high heels
Wan Video 2.2… See the full description on the dataset page: https://huggingface.co/datasets/ApacheOne/Info_Wan_Video_2.2_T2V-A14B.dORM-14B-test
Reward of test split extracted by dORM-14B: dORM-14B-test
Usage
from datasets import load_dataset
# Load entire dataset
dataset = load_dataset("dongboklee/dORM-14B-test")
# Load specific domain
law_dataset = load_dataset("dongboklee/dORM-14B-test", split="law")
Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8-details
Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8
Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8-details.Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.9-details
Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.9
Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.9
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.9-details.CultriX__SeQwence-14B-details
Dataset Card for Evaluation run of CultriX/SeQwence-14B
Dataset automatically created during the evaluation run of model CultriX/SeQwence-14B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CultriX__SeQwence-14B-details.mrm8488__phi-4-14B-grpo-gsm8k-3e-details
Dataset Card for Evaluation run of mrm8488/phi-4-14B-grpo-gsm8k-3e
Dataset automatically created during the evaluation run of model mrm8488/phi-4-14B-grpo-gsm8k-3e
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mrm8488__phi-4-14B-grpo-gsm8k-3e-details.YOYO-AI__Qwen2.5-14B-YOYO-V4-details
Dataset Card for Evaluation run of YOYO-AI/Qwen2.5-14B-YOYO-V4
Dataset automatically created during the evaluation run of model YOYO-AI/Qwen2.5-14B-YOYO-V4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/YOYO-AI__Qwen2.5-14B-YOYO-V4-details.ja_conv_wikipedia_orion14B_100K
Abstruct
This is a multi-turn conversation dataset generated from the Japanese Wikipedia dataset using Orion14B-Chat. Commercial use is possible, but the license is complicated, so please read it carefully before using it.
I generated V100x4 on 200 machines in about half a week.
License
【Orion-14B Series】 Models Community License Agreement
https://huggingface.co/OrionStarAI/Orion-14B-Chat/blob/main/ModelsCommunityLicenseAgreement
Computing
ABCI… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/ja_conv_wikipedia_orion14B_100K.gORM-14B-test
Reward of test split extracted by dORM-14B: dORM-14B-test
Usage
from datasets import load_dataset
# Load entire dataset
dataset = load_dataset("dongboklee/dORM-14B-test")
# Load specific domain
law_dataset = load_dataset("dongboklee/dORM-14B-test", split="law")
sometimesanotion__Qwenvergence-14B-v12-Prose-DS-details
Dataset Card for Evaluation run of sometimesanotion/Qwenvergence-14B-v12-Prose-DS
Dataset automatically created during the evaluation run of model sometimesanotion/Qwenvergence-14B-v12-Prose-DS
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwenvergence-14B-v12-Prose-DS-details.TSD-KD-Qwen2.5-14B-Instruct-Gen
TSD-KD-Qwen2.5-1.5B-Instruct-Gen
This dataset contains student-generated examples used for Token-Selective Dual Knowledge Distillation (TSD-KD), introduced in our ICLR 2026 paper:
"Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation"
Paper: https://arxiv.org/abs/2603.13260
Github: https://github.com/kmswin1/TSD-KD
Dataset Description
This dataset contains teacher-generated instruction-response examples from… See the full description on the dataset page: https://huggingface.co/datasets/Minsang/TSD-KD-Qwen2.5-14B-Instruct-Gen.HeraiHench__Phi-4-slerp-ReasoningRP-14B-details
Dataset Card for Evaluation run of HeraiHench/Phi-4-slerp-ReasoningRP-14B
Dataset automatically created during the evaluation run of model HeraiHench/Phi-4-slerp-ReasoningRP-14B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HeraiHench__Phi-4-slerp-ReasoningRP-14B-details.YOYO-AI__Qwen2.5-14B-it-restore-details
Dataset Card for Evaluation run of YOYO-AI/Qwen2.5-14B-it-restore
Dataset automatically created during the evaluation run of model YOYO-AI/Qwen2.5-14B-it-restore
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/YOYO-AI__Qwen2.5-14B-it-restore-details.gORM-14B-test_smollm
Reward of test_smollm split extracted by dORM-14B: dORM-14B-test_smollm
Usage
from datasets import load_dataset
# Load entire dataset
dataset = load_dataset("dongboklee/dORM-14B-test_smollm")
# Load specific domain
law_dataset = load_dataset("dongboklee/dORM-14B-test_smollm", split="law")
gORM-14B-test_qwen
Reward of test_qwen split extracted by dORM-14B: dORM-14B-test_qwen
Usage
from datasets import load_dataset
# Load entire dataset
dataset = load_dataset("dongboklee/dORM-14B-test_qwen")
# Load specific domain
law_dataset = load_dataset("dongboklee/dORM-14B-test_qwen", split="law")
gPRM-14B-test_smollm
Reward of test_smollm split extracted by gPRM-14B: gPRM-14B-test_smollm
Usage
from datasets import load_dataset
# Load entire dataset
dataset = load_dataset("dongboklee/gPRM-14B-test_smollm")
# Load specific domain
law_dataset = load_dataset("dongboklee/gPRM-14B-test_smollm", split="law")
sometimesanotion__Qwen-2.5-14B-Virmarckeoso-details
Dataset Card for Evaluation run of sometimesanotion/Qwen-2.5-14B-Virmarckeoso
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen-2.5-14B-Virmarckeoso
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen-2.5-14B-Virmarckeoso-details.
