CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01JacobiForcing /OpenThought2_length_bucketedtext1M<n<10M0 likes114 downloads1y agoHugging Face02joeyzero /OpenThought-144k-Backfill-0.2text100K<n<1M0 likes107 downloads11mo agoHugging Face03Thinking-Space /OpenThought3-Qwen3-4BOpenThought3-Qwen3-4B OpenThought3-Qwen3-4B is a math reasoning supervised fine-tuning dataset in chat-message JSONL format. Data Creation and Cleaning This dataset was generated by Qwen3-4B (Non-thinking) from math-domain prompts selected from OpenThoughts3-1.2M. The generated responses were cleaned through deduplication, removal of degenerate repetition/repeater-style outputs, and template checks on the assistant… See the full description on the dataset page: https://huggingface.co/datasets/Thinking-Space/OpenThought3-Qwen3-4B.texttext-generation100K<n<1M3 likes107 downloads5mo agoHugging Face04selimc /OpenThoughts-TR-18k OpenThoughts-TR-18k: Turkish Synthetic Reasoning Dataset OpenThoughts-TR-18k is a Turkish translation of a subset of the original Open-Thoughts-114k dataset. It contains ~18k high-quality synthetic reasoning examples covering mathematics, science, coding problems, and puzzles, all translated into Turkish. This dataset is designed to support reasoning task fine tuning for Turkish language models. Dataset Details ~18k translated reasoning examples Covers multiple domains:… See the full description on the dataset page: https://huggingface.co/datasets/selimc/OpenThoughts-TR-18k.text10K<n<100K8 likes99 downloads2y agoHugging Face05nassimjp /Pashto-OpenThoughts-15K-Reasoning Pashto-OpenThoughts-15K-Reasoning Pashto reasoning dataset based on OpenThoughts-114k, filtered to samples up to approximately 15K characters and translated into natural Pashto. 📌 Dataset Description Pashto-OpenThoughts-15K-Reasoning is a Pashto reasoning dataset created from the OpenThoughts-114k dataset. The dataset focuses on translating and preserving reasoning-oriented examples into Pashto while maintaining important technical structures such as: Python and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-OpenThoughts-15K-Reasoning.texttext-generation1K<n<10K0 likes55 downloads1d agoHugging Face06laion /sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13 Nemotron Terminal SFT reproduction evaluation artifacts This repository contains the complete Harbor artifact tree for the 300-trial OpenThoughts-TBLite evaluation of laion/sft-repro-thinking-step630-nemotron-terminal-step1888. The checkpoint was trained from the Grug stage-2 thinking checkpoint on the Nemotron Terminal corpus for 1,888 steps. Result Measure Value Attempted / completed 300 / 300 Verifier-scoreable 259 (86.33%) Aggregate reward, all… See the full description on the dataset page: https://huggingface.co/datasets/laion/sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13.texttext-generationn<1K0 likes51 downloads1mo agoHugging Face07nassimjp /Da-Ploshi-OpenThoughts_Cache Da-Ploshi OpenThoughts Cache Da-Ploshi OpenThoughts Cache is a large-scale English→Pashto translation dataset focused on programming terminology, algorithmic instructions, code comments, and technical micro‑phrases. The dataset is provided exclusively in JSONL format due to its size (3GB+), making it efficient for streaming, sharding, and training Pashto LLMs. Dataset Structure Each line in the dataset is a standalone JSON object containing an English source… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Da-Ploshi-OpenThoughts_Cache.texttext-generation100K<n<1M0 likes50 downloads1d agoHugging Face08lemon-mint /OpenThoughts-114k-Normalizedprefixes = [ "Generate an executable Python function generated from the given prompt. The function should take stdin as input and print the output. Simply call the function after the definition.", "Return your final response within \\boxed{}. ", "Generate an executable Python function generated from the given prompt. Return the function body without invoking it at the final solution.", ] // -1 if None text100K<n<1M1 likes47 downloads2y agoHugging Face09dvtiendat /hypergraph_openthoughts30k Hypergraph OpenThoughts Math 30K Reasoning hypergraphs generated for the 29,434 examples in siyanzhao/Openthoughts_math_30k_opsd. Generation Model: Qwen/Qwen3.6-35B-A3B-FP8 Thinking mode: disabled Construction: semantic-step segmentation followed by primary-support DAG induction Graph constraint: at most one earlier-step parent per semantic step Processing order: source dataset row order Schema Each JSONL record contains: row_index: source… See the full description on the dataset page: https://huggingface.co/datasets/dvtiendat/hypergraph_openthoughts30k.tabular10K<n<100K0 likes47 downloads10d agoHugging Face10lemon-mint /OpenThoughts-114k-JSONLtext100K<n<1M1 likes42 downloads2y agoHugging Face11pre-to-post-olmo /openthoughts_merged_think_39k openthoughts_merged_think_39k Merged think-format SFT dataset (ShareGPT-style: system + conversations with from/value), 39,874 examples, for OLMo SFT. Composition (concatenation of two decontaminated think-format sources): open-thoughts114k_math_20k_decontam_think — 20,000 examples sampled from OpenThoughts-114k math, decontaminated against the OpenThoughts3 set below. openthoughts3_math_decontam_resp_lt8192_think — 19,874 examples from… See the full description on the dataset page: https://huggingface.co/datasets/pre-to-post-olmo/openthoughts_merged_think_39k.texttext-generation10K<n<100K0 likes42 downloads3mo agoHugging Face12melhoushi /OpenThoughts3_science_qwen7binst_sft_2048text10K<n<100K0 likes36 downloads5mo agoHugging Face13teetone /qwen3_4b_openthoughts4_code9K_instill_n8_valredundancy5_factualcorrectness_round1text1K<n<10K0 likes33 downloads2mo agoHugging Face14XinnanZhang /openthoughts3_math_10k8_answertext10K<n<100K0 likes30 downloads5mo agoHugging Face15teetone /qwen3_4b_openthoughts4_code9K_instill_n8_valredundancy5_round1text1K<n<10K0 likes29 downloads6mo agoHugging Face16codex-master /openthoughts3_numinamath-1.5-pro_mixturetexttext-generation10K<n<100K0 likes29 downloads1mo agoHugging Face17teetone /qwen3_4b_openthoughts4_code9K_instill_n4_valredundancy5_round1text1K<n<10K1 likes26 downloads6mo agoHugging Face18teetone /deepseek_r1_distill_llama_8b_openthoughts4_code9K_instill_n8_valredundancy5_round1text1K<n<10K0 likes26 downloads2mo agoHugging Face19teetone /qwen3_4b_openthoughts4_code9K_instill_n8_valredundancy5_cyclecorrectness_round1text1K<n<10K0 likes25 downloads2mo agoHugging Face20open-llm-leaderboard /open-thoughts__OpenThinker-7B-detailsgated Dataset Card for Evaluation run of open-thoughts/OpenThinker-7B Dataset automatically created during the evaluation run of model open-thoughts/OpenThinker-7B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/open-thoughts__OpenThinker-7B-details.tabular10K<n<100K0 likes23 downloads2y agoHugging Face21Snyhlx /merged_raw_openthought2_math_unfiltered_split2text100K<n<1M0 likes23 downloads1y agoHugging Face22XinnanZhang /openthoughts3-math-50k8text100K<n<1M0 likes22 downloads5mo agoHugging Face23teetone /qwen3_32b_openthoughts3_math53K_instill_n8_valredundancy5_round1text10K<n<100K0 likes22 downloads4mo agoHugging Face24LLMTeamAkiyama /cleand_openthought312_dif9_tiny元データ: https://huggingface.co/datasets/LLMTeamAkiyama/clean_openthought312_difficulty_9_filterd データ件数: 1,456 平均トークン数: 5,894 最大トークン数: 8,186 合計トークン数: 8,581,562 ファイル形式: JSONL ファイル分割数: 1 合計ファイルサイズ: 33.2 MB 加工内容: 元データに対して、token数を8912以下に制限したテスト用tiny版 使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/openthoughts3/clean_openthoughts3_tiny_pickup.ipynb tabularquestion-answering1K<n<10K0 likes21 downloads1y agoHugging Face25teetone /temp_openthoughts4_math_qwen3-4b_round1text10K<n<100K0 likes21 downloads8mo agoHugging Face26teetone /qwen3_32b_openthoughts4_code9K_instill_n8_valredundancy5_round1text1K<n<10K0 likes21 downloads5mo agoHugging Face27Axolotl-Partners /openthoughts-114k-linkedtext100K<n<1M0 likes21 downloads2mo agoHugging Face28gauravjain14 /open-thoughts-deepseekr1textn<1K0 likes20 downloads2y agoHugging Face29agentlans /OpenThoughts OpenThoughts Long Chain-Of-Thought Collection This dataset is an unofficial, curated compilation of long-form reasoning traces derived from the OpenThoughts organization's datasets. It is designed to provide high-quality, valid reasoning chains for training and fine-tuning large language models. Dataset Overview This collection aggregates reasoning data generated by state-of-the-art models. The data has been cleaned to ensure that only rows containing both a valid answer… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/OpenThoughts.texttext-generation100K<n<1M0 likes20 downloads5mo agoHugging Face30Snyhlx /merged_raw_openthought2_mathtext10K<n<100K0 likes19 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.