datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenThought2_length_bucketedOpenThought-144k-Backfill-0.2OpenThought3-Qwen3-4BOpenThought3-Qwen3-4B
OpenThought3-Qwen3-4B is a math reasoning supervised fine-tuning dataset in chat-message JSONL format.
Data Creation and Cleaning
This dataset was generated by Qwen3-4B (Non-thinking) from math-domain prompts selected from OpenThoughts3-1.2M. The generated responses were cleaned through deduplication, removal of degenerate repetition/repeater-style outputs, and template checks on the assistant… See the full description on the dataset page: https://huggingface.co/datasets/Thinking-Space/OpenThought3-Qwen3-4B.OpenThoughts-TR-18k
OpenThoughts-TR-18k: Turkish Synthetic Reasoning Dataset
OpenThoughts-TR-18k is a Turkish translation of a subset of the original Open-Thoughts-114k dataset. It contains ~18k high-quality synthetic reasoning examples covering mathematics, science, coding problems, and puzzles, all translated into Turkish. This dataset is designed to support reasoning task fine tuning for Turkish language models.
Dataset Details
~18k translated reasoning examples
Covers multiple domains:… See the full description on the dataset page: https://huggingface.co/datasets/selimc/OpenThoughts-TR-18k.Pashto-OpenThoughts-15K-Reasoning
Pashto-OpenThoughts-15K-Reasoning
Pashto reasoning dataset based on OpenThoughts-114k, filtered to samples up to approximately 15K characters and translated into natural Pashto.
📌 Dataset Description
Pashto-OpenThoughts-15K-Reasoning is a Pashto reasoning dataset created from the OpenThoughts-114k dataset.
The dataset focuses on translating and preserving reasoning-oriented examples into Pashto while maintaining important technical structures such as:
Python and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-OpenThoughts-15K-Reasoning.sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13
Nemotron Terminal SFT reproduction evaluation artifacts
This repository contains the complete Harbor artifact tree for the 300-trial
OpenThoughts-TBLite evaluation of
laion/sft-repro-thinking-step630-nemotron-terminal-step1888.
The checkpoint was trained from the Grug stage-2 thinking checkpoint on the
Nemotron Terminal corpus for 1,888 steps.
Result
Measure
Value
Attempted / completed
300 / 300
Verifier-scoreable
259 (86.33%)
Aggregate reward, all… See the full description on the dataset page: https://huggingface.co/datasets/laion/sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13.Da-Ploshi-OpenThoughts_Cache
Da-Ploshi OpenThoughts Cache
Da-Ploshi OpenThoughts Cache is a large-scale English→Pashto translation dataset focused on programming terminology, algorithmic instructions, code comments, and technical micro‑phrases.
The dataset is provided exclusively in JSONL format due to its size (3GB+), making it efficient for streaming, sharding, and training Pashto LLMs.
Dataset Structure
Each line in the dataset is a standalone JSON object containing an English source… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Da-Ploshi-OpenThoughts_Cache.OpenThoughts-114k-Normalizedprefixes = [
"Generate an executable Python function generated from the given prompt. The function should take stdin as input and print the output. Simply call the function after the definition.",
"Return your final response within \\boxed{}. ",
"Generate an executable Python function generated from the given prompt. Return the function body without invoking it at the final solution.",
]
// -1 if None
hypergraph_openthoughts30k
Hypergraph OpenThoughts Math 30K
Reasoning hypergraphs generated for the 29,434 examples in
siyanzhao/Openthoughts_math_30k_opsd.
Generation
Model: Qwen/Qwen3.6-35B-A3B-FP8
Thinking mode: disabled
Construction: semantic-step segmentation followed by primary-support DAG induction
Graph constraint: at most one earlier-step parent per semantic step
Processing order: source dataset row order
Schema
Each JSONL record contains:
row_index: source… See the full description on the dataset page: https://huggingface.co/datasets/dvtiendat/hypergraph_openthoughts30k.OpenThoughts-114k-JSONLopenthoughts_merged_think_39k
openthoughts_merged_think_39k
Merged think-format SFT dataset (ShareGPT-style: system + conversations with
from/value), 39,874 examples, for OLMo SFT.
Composition (concatenation of two decontaminated think-format sources):
open-thoughts114k_math_20k_decontam_think — 20,000 examples sampled from
OpenThoughts-114k math, decontaminated against the OpenThoughts3 set below.
openthoughts3_math_decontam_resp_lt8192_think — 19,874 examples from… See the full description on the dataset page: https://huggingface.co/datasets/pre-to-post-olmo/openthoughts_merged_think_39k.OpenThoughts3_science_qwen7binst_sft_2048qwen3_4b_openthoughts4_code9K_instill_n8_valredundancy5_factualcorrectness_round1openthoughts3_math_10k8_answerqwen3_4b_openthoughts4_code9K_instill_n8_valredundancy5_round1openthoughts3_numinamath-1.5-pro_mixtureqwen3_4b_openthoughts4_code9K_instill_n4_valredundancy5_round1deepseek_r1_distill_llama_8b_openthoughts4_code9K_instill_n8_valredundancy5_round1qwen3_4b_openthoughts4_code9K_instill_n8_valredundancy5_cyclecorrectness_round1open-thoughts__OpenThinker-7B-details
Dataset Card for Evaluation run of open-thoughts/OpenThinker-7B
Dataset automatically created during the evaluation run of model open-thoughts/OpenThinker-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/open-thoughts__OpenThinker-7B-details.merged_raw_openthought2_math_unfiltered_split2openthoughts3-math-50k8qwen3_32b_openthoughts3_math53K_instill_n8_valredundancy5_round1cleand_openthought312_dif9_tiny元データ: https://huggingface.co/datasets/LLMTeamAkiyama/clean_openthought312_difficulty_9_filterd
データ件数: 1,456
平均トークン数: 5,894
最大トークン数: 8,186
合計トークン数: 8,581,562
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 33.2 MB
加工内容:
元データに対して、token数を8912以下に制限したテスト用tiny版
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/openthoughts3/clean_openthoughts3_tiny_pickup.ipynb
temp_openthoughts4_math_qwen3-4b_round1qwen3_32b_openthoughts4_code9K_instill_n8_valredundancy5_round1openthoughts-114k-linkedopen-thoughts-deepseekr1OpenThoughts
OpenThoughts Long Chain-Of-Thought Collection
This dataset is an unofficial, curated compilation of long-form reasoning traces derived from the OpenThoughts organization's datasets. It is designed to provide high-quality, valid reasoning chains for training and fine-tuning large language models.
Dataset Overview
This collection aggregates reasoning data generated by state-of-the-art models. The data has been cleaned to ensure that only rows containing both a valid answer… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/OpenThoughts.merged_raw_openthought2_math
