datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llm-cipher-reasoning
llm-cipher-reasoning — data, eval results and full research ledger
Everything except the weights from a research run asking: can an LLM be trained to reason in a
more compact "language" than English, and does that actually save tokens?
Two linked lines of work on Qwen/Qwen3-4B-Instruct-2507:
Cipher invention / cross-model communication — cold-decoding tests, negotiated cipher
collusion between model pairs, a cipher-hardening arms race, and GEPA prompt optimization to get
a… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/llm-cipher-reasoning.NLP-Course-LLM-Reasoning-Eval-May2025
Overview of LLM Reasoning Eval Dataset
This dataset contains evaluation of multiple large language models (LLMs) over 918 MCQ reasoning questions created by 184 students.
Each question was used to test 3 LLMs (each 3 times): GPT-4o, Claude Sonnet 3.x (3.5 or 3.7), and Deepseek R1.
The questions target various reasoning areas (i.e., Math, Logic, Temporal, Commonsense) and are included only if 3 seperate attempts (in a new session) by ChatGPT (GPT-4o) fail at giving the correct… See the full description on the dataset page: https://huggingface.co/datasets/nlpllmeval/NLP-Course-LLM-Reasoning-Eval-May2025.llm-medical-reasoning-steps-benchmark
LLM Medical Reasoning Steps Benchmark
This dataset contains 1,170 medical reasoning benchmark questions with final answers, reference reasoning steps, and reference key points.
Dataset Files
data/all.jsonl: all 1,170 examples.
data/mcq.jsonl: 592 multiple-choice examples.
data/oeq.jsonl: 578 open-ended examples.
No model prediction outputs are included in this release.
Schema
Each JSONL row has the following fields:
{
"id": "mcq_0001",
"task_type":… See the full description on the dataset page: https://huggingface.co/datasets/medreason/llm-medical-reasoning-steps-benchmark.KingNish__Reasoning-0.5b-details
Dataset Card for Evaluation run of KingNish/Reasoning-0.5b
Dataset automatically created during the evaluation run of model KingNish/Reasoning-0.5b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/KingNish__Reasoning-0.5b-details.llm-complex-reasoning-train-qwen2-72b-instruct-correct
Note
Data Seed from 基于封闭世界假设的复杂逻辑推理
Generate from Qwen2-72B-Instruct with prompt
train.jsonl for 推理答案和题目答案一致, no_train.jsonl推理答案和题目答案不一致
注: 题目答案不一定正确
HeraiHench__Phi-4-slerp-ReasoningRP-14B-details
Dataset Card for Evaluation run of HeraiHench/Phi-4-slerp-ReasoningRP-14B
Dataset automatically created during the evaluation run of model HeraiHench/Phi-4-slerp-ReasoningRP-14B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HeraiHench__Phi-4-slerp-ReasoningRP-14B-details.EpistemeAI__Reasoning-Llama-3.2-3B-Math-Instruct-RE1-details
Dataset Card for Evaluation run of EpistemeAI/Reasoning-Llama-3.2-3B-Math-Instruct-RE1
Dataset automatically created during the evaluation run of model EpistemeAI/Reasoning-Llama-3.2-3B-Math-Instruct-RE1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Reasoning-Llama-3.2-3B-Math-Instruct-RE1-details.EpistemeAI__ReasoningCore-3B-RE1-V2A-details
Dataset Card for Evaluation run of EpistemeAI/ReasoningCore-3B-RE1-V2A
Dataset automatically created during the evaluation run of model EpistemeAI/ReasoningCore-3B-RE1-V2A
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__ReasoningCore-3B-RE1-V2A-details.EpistemeAI__ReasoningCore-3B-Instruct-r01-Reflect-details
Dataset Card for Evaluation run of EpistemeAI/ReasoningCore-3B-Instruct-r01-Reflect
Dataset automatically created during the evaluation run of model EpistemeAI/ReasoningCore-3B-Instruct-r01-Reflect
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__ReasoningCore-3B-Instruct-r01-Reflect-details.Dans-DiscountModels__12b-mn-dans-reasoning-test-2-details
Dataset Card for Evaluation run of Dans-DiscountModels/12b-mn-dans-reasoning-test-2
Dataset automatically created during the evaluation run of model Dans-DiscountModels/12b-mn-dans-reasoning-test-2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Dans-DiscountModels__12b-mn-dans-reasoning-test-2-details.EpistemeAI__ReasoningCore-3B-RE1-V2-details
Dataset Card for Evaluation run of EpistemeAI/ReasoningCore-3B-RE1-V2
Dataset automatically created during the evaluation run of model EpistemeAI/ReasoningCore-3B-RE1-V2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__ReasoningCore-3B-RE1-V2-details.EpistemeAI__ReasoningCore-1.0-3B-Instruct-r01-Reflect-Math-details
Dataset Card for Evaluation run of EpistemeAI/ReasoningCore-1.0-3B-Instruct-r01-Reflect-Math
Dataset automatically created during the evaluation run of model EpistemeAI/ReasoningCore-1.0-3B-Instruct-r01-Reflect-Math
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__ReasoningCore-1.0-3B-Instruct-r01-Reflect-Math-details.EpistemeAI__Reasoning-Llama-3.1-CoT-RE1-NMT-details
Dataset Card for Evaluation run of EpistemeAI/Reasoning-Llama-3.1-CoT-RE1-NMT
Dataset automatically created during the evaluation run of model EpistemeAI/Reasoning-Llama-3.1-CoT-RE1-NMT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Reasoning-Llama-3.1-CoT-RE1-NMT-details.saishshinde15__TethysAI_Base_Reasoning-details
Dataset Card for Evaluation run of saishshinde15/TethysAI_Base_Reasoning
Dataset automatically created during the evaluation run of model saishshinde15/TethysAI_Base_Reasoning
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/saishshinde15__TethysAI_Base_Reasoning-details.EpistemeAI__ReasoningCore-3B-RE1-V2C-details
Dataset Card for Evaluation run of EpistemeAI/ReasoningCore-3B-RE1-V2C
Dataset automatically created during the evaluation run of model EpistemeAI/ReasoningCore-3B-RE1-V2C
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__ReasoningCore-3B-RE1-V2C-details.Dans-DiscountModels__12b-mn-dans-reasoning-test-3-details
Dataset Card for Evaluation run of Dans-DiscountModels/12b-mn-dans-reasoning-test-3
Dataset automatically created during the evaluation run of model Dans-DiscountModels/12b-mn-dans-reasoning-test-3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Dans-DiscountModels__12b-mn-dans-reasoning-test-3-details.saishshinde15__TethysAI_Vortex_Reasoning-details
Dataset Card for Evaluation run of saishshinde15/TethysAI_Vortex_Reasoning
Dataset automatically created during the evaluation run of model saishshinde15/TethysAI_Vortex_Reasoning
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/saishshinde15__TethysAI_Vortex_Reasoning-details.sometimesanotion__IF-reasoning-experiment-80-details
Dataset Card for Evaluation run of sometimesanotion/IF-reasoning-experiment-80
Dataset automatically created during the evaluation run of model sometimesanotion/IF-reasoning-experiment-80
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__IF-reasoning-experiment-80-details.prithivMLmods__Deepthink-Reasoning-7B-details
Dataset Card for Evaluation run of prithivMLmods/Deepthink-Reasoning-7B
Dataset automatically created during the evaluation run of model prithivMLmods/Deepthink-Reasoning-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Deepthink-Reasoning-7B-details.EpistemeAI__ReasoningCore-3B-R01-details
Dataset Card for Evaluation run of EpistemeAI/ReasoningCore-3B-R01
Dataset automatically created during the evaluation run of model EpistemeAI/ReasoningCore-3B-R01
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__ReasoningCore-3B-R01-details.ewre324__Thinker-SmolLM2-135M-Instruct-Reasoning-details
Dataset Card for Evaluation run of ewre324/Thinker-SmolLM2-135M-Instruct-Reasoning
Dataset automatically created during the evaluation run of model ewre324/Thinker-SmolLM2-135M-Instruct-Reasoning
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ewre324__Thinker-SmolLM2-135M-Instruct-Reasoning-details.EpistemeAI__ReasoningCore-3B-RE1-V2B-details
Dataset Card for Evaluation run of EpistemeAI/ReasoningCore-3B-RE1-V2B
Dataset automatically created during the evaluation run of model EpistemeAI/ReasoningCore-3B-RE1-V2B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__ReasoningCore-3B-RE1-V2B-details.EpistemeAI__ReasoningCore-3B-T1_1-details
Dataset Card for Evaluation run of EpistemeAI/ReasoningCore-3B-T1_1
Dataset automatically created during the evaluation run of model EpistemeAI/ReasoningCore-3B-T1_1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__ReasoningCore-3B-T1_1-details.DavidAU__Qwen2.5-MOE-6x1.5B-DeepSeek-Reasoning-e32-details
Dataset Card for Evaluation run of DavidAU/Qwen2.5-MOE-6x1.5B-DeepSeek-Reasoning-e32
Dataset automatically created during the evaluation run of model DavidAU/Qwen2.5-MOE-6x1.5B-DeepSeek-Reasoning-e32
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DavidAU__Qwen2.5-MOE-6x1.5B-DeepSeek-Reasoning-e32-details.bunnycore__Phi-4-ReasoningRP-details
Dataset Card for Evaluation run of bunnycore/Phi-4-ReasoningRP
Dataset automatically created during the evaluation run of model bunnycore/Phi-4-ReasoningRP
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__Phi-4-ReasoningRP-details.ewre324__Thinker-Llama-3.2-3B-Instruct-Reasoning-details
Dataset Card for Evaluation run of ewre324/Thinker-Llama-3.2-3B-Instruct-Reasoning
Dataset automatically created during the evaluation run of model ewre324/Thinker-Llama-3.2-3B-Instruct-Reasoning
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ewre324__Thinker-Llama-3.2-3B-Instruct-Reasoning-details.ewre324__Thinker-Qwen2.5-0.5B-Instruct-Reasoning-details
Dataset Card for Evaluation run of ewre324/Thinker-Qwen2.5-0.5B-Instruct-Reasoning
Dataset automatically created during the evaluation run of model ewre324/Thinker-Qwen2.5-0.5B-Instruct-Reasoning
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ewre324__Thinker-Qwen2.5-0.5B-Instruct-Reasoning-details.EpistemeAI__Reasoning-Llama-3.2-3B-Math-Instruct-RE1-ORPO-details
Dataset Card for Evaluation run of EpistemeAI/Reasoning-Llama-3.2-3B-Math-Instruct-RE1-ORPO
Dataset automatically created during the evaluation run of model EpistemeAI/Reasoning-Llama-3.2-3B-Math-Instruct-RE1-ORPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Reasoning-Llama-3.2-3B-Math-Instruct-RE1-ORPO-details.EpistemeAI__Reasoning-Llama-3.2-1B-Instruct-v1.2-details
Dataset Card for Evaluation run of EpistemeAI/Reasoning-Llama-3.2-1B-Instruct-v1.2
Dataset automatically created during the evaluation run of model EpistemeAI/Reasoning-Llama-3.2-1B-Instruct-v1.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Reasoning-Llama-3.2-1B-Instruct-v1.2-details.DavidAU__DeepThought-MOE-8X3B-R1-Llama-3.2-Reasoning-18B-details
Dataset Card for Evaluation run of DavidAU/DeepThought-MOE-8X3B-R1-Llama-3.2-Reasoning-18B
Dataset automatically created during the evaluation run of model DavidAU/DeepThought-MOE-8X3B-R1-Llama-3.2-Reasoning-18B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DavidAU__DeepThought-MOE-8X3B-R1-Llama-3.2-Reasoning-18B-details.
