datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Intermediate-Thinking-130k
Intermediate-Thinking-130k
A comprehensive dataset of 135,000 high-quality samples designed to advance language model reasoning capabilities through structured intermediate thinking processes. This dataset enables training and evaluation of models with sophisticated self-correction and iterative reasoning abilities across 42 languages.
OG Link
Overview
Intermediate-Thinking-130k addresses a fundamental limitation in current language models: their inability to pause… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/Intermediate-Thinking-130k.Intermediate-Thinking-130k
Intermediate-Thinking-130k
A comprehensive dataset of 135,000 high-quality samples designed to advance language model reasoning capabilities through structured intermediate thinking processes. This dataset enables training and evaluation of models with sophisticated self-correction and iterative reasoning abilities across 42 languages.
Overview
Intermediate-Thinking-130k addresses a fundamental limitation in current language models: their inability to pause, reflect, and… See the full description on the dataset page: https://huggingface.co/datasets/HelpingAI/Intermediate-Thinking-130k.Intermediate-Thinking-130k
Qwen3 Tokenizer
Total tokens in dataset: 118.991.025
Max token length: 11.406
Max token sample index: 119.408
Using System Prompt (308 Tokens):
J.O.S.I.E.-I.R.-1 (Just One Super Intelligent Entity - Intermediate Reasoning - Version 1), an advanced AI assistant designed for strong intermediate reasoning and high-quality problem solving. You refer to yourself as Josie.
You solve problems by reasoning through them in stages. When reasoning is required, you naturally perform an… See the full description on the dataset page: https://huggingface.co/datasets/Goekdeniz-Guelmez/Intermediate-Thinking-130k.
