datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
M-Thinker-SFT-dataKAG-Thinker-training-datasetDavidAU__DeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B-details
Dataset Card for Evaluation run of DavidAU/DeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B
Dataset automatically created during the evaluation run of model DavidAU/DeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DavidAU__DeepSeek-MOE-4X8B-R1-Distill-Llama-3.1-Deep-Thinker-Uncensored-24B-details.Detailed-Thinker
Detailed Thinker (snapshot)
This is an early snapshot of Detailed Thinker, started 12 August 2026. The set is still in active development. Treat this release as a work-in-progress cut (around 300 examples), not a finished corpus.
What this is for
Shallow build requests often get shallow answers. If you say "build a barbershop simulator," a model can sketch something but it usually skips real depth on functionality, constraints, and follow-through.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/Detailed-Thinker.ewre324__Thinker-SmolLM2-135M-Instruct-Reasoning-details
Dataset Card for Evaluation run of ewre324/Thinker-SmolLM2-135M-Instruct-Reasoning
Dataset automatically created during the evaluation run of model ewre324/Thinker-SmolLM2-135M-Instruct-Reasoning
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ewre324__Thinker-SmolLM2-135M-Instruct-Reasoning-details.olmo_msgs_thinkerunfiltered-thinker
Unfiltered-Thinker: A Dataset for Intermediate Cognitive Reasoning
A corpus of 1,908 samples designed to showcase intermediate thinking, cognitive processes, and structured emotional reasoning.
Source: UnfilteredAI/unfiltered-thinker on Hugging Face
⚠️ Content Warning: This dataset contains content that will be considered offensive, disturbing, or explicit. This includes discussions of dark humor, profanity, criminal activity, violence, substance use, and psychological distress.… See the full description on the dataset page: https://huggingface.co/datasets/sonic-coder/unfiltered-thinker.fhai50032__Unaligned-Thinker-PHI-4-details
Dataset Card for Evaluation run of fhai50032/Unaligned-Thinker-PHI-4
Dataset automatically created during the evaluation run of model fhai50032/Unaligned-Thinker-PHI-4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fhai50032__Unaligned-Thinker-PHI-4-details.ewre324__Thinker-Llama-3.2-3B-Instruct-Reasoning-details
Dataset Card for Evaluation run of ewre324/Thinker-Llama-3.2-3B-Instruct-Reasoning
Dataset automatically created during the evaluation run of model ewre324/Thinker-Llama-3.2-3B-Instruct-Reasoning
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ewre324__Thinker-Llama-3.2-3B-Instruct-Reasoning-details.ewre324__Thinker-Qwen2.5-0.5B-Instruct-Reasoning-details
Dataset Card for Evaluation run of ewre324/Thinker-Qwen2.5-0.5B-Instruct-Reasoning
Dataset automatically created during the evaluation run of model ewre324/Thinker-Qwen2.5-0.5B-Instruct-Reasoning
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ewre324__Thinker-Qwen2.5-0.5B-Instruct-Reasoning-details.nbeerbower-Purpura-DPO-thinker-rawUnfiltered
System promt for creating a dataset:
You are an expert AI assistant specializing in text generation. Your task is to reverse-engineer the thought process that leads to a given textual `response`.
Based on the user's `prompt` and the final `response` text, generate a plausible, detailed reasoning process of an LLM.
This reasoning should cover:
1. **Analysis of the User's Prompt:** Deconstruct the user's request, identifying explicit constraints (like length, format) and implicit… See the full description on the dataset page: https://huggingface.co/datasets/Disya/nbeerbower-Purpura-DPO-thinker-raw.bunnycore__Qwen2.5-3B-RP-Thinker-details
Dataset Card for Evaluation run of bunnycore/Qwen2.5-3B-RP-Thinker
Dataset automatically created during the evaluation run of model bunnycore/Qwen2.5-3B-RP-Thinker
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__Qwen2.5-3B-RP-Thinker-details.bunnycore__Qwen2.5-3B-RP-Thinker-V2-details
Dataset Card for Evaluation run of bunnycore/Qwen2.5-3B-RP-Thinker-V2
Dataset automatically created during the evaluation run of model bunnycore/Qwen2.5-3B-RP-Thinker-V2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__Qwen2.5-3B-RP-Thinker-V2-details.bunnycore__Qwen2.5-7B-RRP-1M-Thinker-details
Dataset Card for Evaluation run of bunnycore/Qwen2.5-7B-RRP-1M-Thinker
Dataset automatically created during the evaluation run of model bunnycore/Qwen2.5-7B-RRP-1M-Thinker
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__Qwen2.5-7B-RRP-1M-Thinker-details.thinkertulu_3_thinker_test
