datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenThought3-Qwen3-4BOpenThought3-Qwen3-4B
OpenThought3-Qwen3-4B is a math reasoning supervised fine-tuning dataset in chat-message JSONL format.
Data Creation and Cleaning
This dataset was generated by Qwen3-4B (Non-thinking) from math-domain prompts selected from OpenThoughts3-1.2M. The generated responses were cleaned through deduplication, removal of degenerate repetition/repeater-style outputs, and template checks on the assistant… See the full description on the dataset page: https://huggingface.co/datasets/Thinking-Space/OpenThought3-Qwen3-4B.Pashto-OpenThoughts-15K-Reasoning
Pashto-OpenThoughts-15K-Reasoning
Pashto reasoning dataset based on OpenThoughts-114k, filtered to samples up to approximately 15K characters and translated into natural Pashto.
📌 Dataset Description
Pashto-OpenThoughts-15K-Reasoning is a Pashto reasoning dataset created from the OpenThoughts-114k dataset.
The dataset focuses on translating and preserving reasoning-oriented examples into Pashto while maintaining important technical structures such as:
Python and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-OpenThoughts-15K-Reasoning.sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13
Nemotron Terminal SFT reproduction evaluation artifacts
This repository contains the complete Harbor artifact tree for the 300-trial
OpenThoughts-TBLite evaluation of
laion/sft-repro-thinking-step630-nemotron-terminal-step1888.
The checkpoint was trained from the Grug stage-2 thinking checkpoint on the
Nemotron Terminal corpus for 1,888 steps.
Result
Measure
Value
Attempted / completed
300 / 300
Verifier-scoreable
259 (86.33%)
Aggregate reward, all… See the full description on the dataset page: https://huggingface.co/datasets/laion/sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13.Da-Ploshi-OpenThoughts_Cache
Da-Ploshi OpenThoughts Cache
Da-Ploshi OpenThoughts Cache is a large-scale English→Pashto translation dataset focused on programming terminology, algorithmic instructions, code comments, and technical micro‑phrases.
The dataset is provided exclusively in JSONL format due to its size (3GB+), making it efficient for streaming, sharding, and training Pashto LLMs.
Dataset Structure
Each line in the dataset is a standalone JSON object containing an English source… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Da-Ploshi-OpenThoughts_Cache.openthoughts_merged_think_39k
openthoughts_merged_think_39k
Merged think-format SFT dataset (ShareGPT-style: system + conversations with
from/value), 39,874 examples, for OLMo SFT.
Composition (concatenation of two decontaminated think-format sources):
open-thoughts114k_math_20k_decontam_think — 20,000 examples sampled from
OpenThoughts-114k math, decontaminated against the OpenThoughts3 set below.
openthoughts3_math_decontam_resp_lt8192_think — 19,874 examples from… See the full description on the dataset page: https://huggingface.co/datasets/pre-to-post-olmo/openthoughts_merged_think_39k.openthoughts3_numinamath-1.5-pro_mixtureOpenThoughts
OpenThoughts Long Chain-Of-Thought Collection
This dataset is an unofficial, curated compilation of long-form reasoning traces derived from the OpenThoughts organization's datasets. It is designed to provide high-quality, valid reasoning chains for training and fine-tuning large language models.
Dataset Overview
This collection aggregates reasoning data generated by state-of-the-art models. The data has been cleaned to ensure that only rows containing both a valid answer… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/OpenThoughts.
