CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01open-r1 /codeforces-cots Dataset Card for CodeForces-CoTs Dataset description CodeForces-CoTs is a large-scale dataset for training reasoning models on competitive programming tasks. It consists of 10k CodeForces problems with up to five reasoning traces generated by DeepSeek R1. We did not filter the traces for correctness, but found that around 84% of the Python ones pass the public tests. The dataset consists of several subsets: solutions: we prompt R1 to solve the problem and produce code.… See the full description on the dataset page: https://huggingface.co/datasets/open-r1/codeforces-cots.tabular100K<n<1M227 likes6k downloads1y agoHugging Face02Mumon /mmlu-pro-self-cot-deepseek-r1Use deepseek-r1 to generate COT in few-shot examples. tabular10K<n<100K1 likes1.6k downloads2y agoHugging Face03Chainticks /cftc-cot Chainticks CFTC COT Normalized CFTC Commitments of Traders legacy futures rows from public-domain CFTC archives. import pandas as pd DATE = "YYYY-MM-DD" URL = "https://huggingface.co/datasets/Chainticks/cftc-cot/resolve/main/legacy_futures/date={DATE}/part-0000.parquet" df = pd.read_parquet(URL) print(df.head()) Layout legacy_futures/date=YYYY-MM-DD/part-0000.parquet _schema.json _manifest.json LATEST_DATE.txt Provenance Rows must have… See the full description on the dataset page: https://huggingface.co/datasets/Chainticks/cftc-cot.tabular10K<n<100K0 likes1.2k downloads7h agoHugging Face04beyoru /Aesir-Character-CoT-roleplay Overview Think with your role. Most reasoning datasets teach models to think like an AI. This one teaches them to think like the character. Continue updating until money run out, I will try to update this dataset in near future Stats 1,973 high-quality conversations (filtered from 2,000 distilled — 27 dropped: prohibited content + missing-review + empty-content) ~14,349 assistant turns, each with full character-POV reasoning Teacher: deepseek-v4-pro… See the full description on the dataset page: https://huggingface.co/datasets/beyoru/Aesir-Character-CoT-roleplay.tabulartext-generation1K<n<10K32 likes1.1k downloads5mo agoHugging Face05RLAIF /numina-math-llama-3.1-8b-bon-meta-cottabular100K<n<1M0 likes635 downloads2y agoHugging Face06videron /gen3_cotrainingThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "right_joint_1.pos", "right_joint_2.pos", "right_joint_3.pos", "right_joint_4.pos", "right_joint_5.pos", "right_joint_6.pos", "right_joint_7.pos"… See the full description on the dataset page: https://huggingface.co/datasets/videron/gen3_cotraining.tabularrobotics1M<n<10M0 likes635 downloads28d agoHugging Face07OpenSakura /OpenSakura-DS-260220-LN-ja-zh-COT-Lilith OpenSakura Lilith LN COT Dataset OpenSakura-DS-260220-LN-ja-zh-COT-Lilith is the COT/segment-level derivative built from the same LN source stream, with reasoning_content preserved. Stats below are computed from the actual generated parquet files. Dataset Summary Metric Value Dataset ID OpenSakura/OpenSakura-DS-260220-LN-ja-zh-COT-Lilith Total rows 692,587 Total parquet files 233 (train: 162, arena: 12, reserve: 12, validation: 24, test: 23) Total size 8… See the full description on the dataset page: https://huggingface.co/datasets/OpenSakura/OpenSakura-DS-260220-LN-ja-zh-COT-Lilith.tabulartranslation100K<n<1M2 likes622 downloads4mo agoHugging Face08lfaviate /China-K12-STEM-10K-CoT-Reasoning K12-STEM-CoT-Chinese 1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams. The largest structured Chinese math/physics/chemistry reasoning dataset. This is a curated sample (10,000 problems) of the full 1.54M dataset available via API. Full Dataset Access Access the full 1,540,000+ problems via API → This Sample Full API Total problems 10,025 1,540,000+ With CoT solutions 10,025 1,490,000+ With diagrams 6,093 740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.tabularquestion-answering10K<n<100K3 likes604 downloads7mo agoHugging Face09ankner /mmlu-pro-CoTtabular10K<n<100K0 likes592 downloads2y agoHugging Face10Arimancy /cftc-cot-weekly CFTC Commitments of Traders weekly panel Every CFTC Commitments of Traders report family in one tidy, model ready weekly panel: harmonized positions, net positioning and COT index features, a market reference map, and documented release provenance, from 1986 to last Friday, in Parquet and CSV. Dataset structure Four tables, each its own named config (different schemas, never concatenated): cot_panel_long: the tidy long panel, one row per (report_date, contract… See the full description on the dataset page: https://huggingface.co/datasets/Arimancy/cftc-cot-weekly.tabular1M<n<10M1 likes374 downloads2d agoHugging Face11jacobmorrison /OpenThoughts3-456k-no-cot-with-olmo-system-prompttabular100K<n<1M0 likes365 downloads1y agoHugging Face12llamastack /gpqa_0shot_cottabular1K<n<10K0 likes364 downloads6mo agoHugging Face13HayatoHongoEveryonesAI /qa_verify_cot_new_6M_unfiltered_v7dataset_names = [ "HayatoHongoEveryonesAI/qa_verify_1m_cot_1", "HayatoHongoEveryonesAI/qa_verify_1m_cot_2", "HayatoHongoEveryonesAI/qa_verify_1m_cot_3", "HayatoHongoEveryonesAI/qa_verify_1m_cot_4", "HayatoHongoEveryonesAI/qa_verify_1m_cot_5", "HayatoHongoEveryonesAI/qa_verify_2m_cot_2", "HayatoHongoEveryonesAI/qa_verify_2m_cot_3", ] https://colab.research.google.com/drive/1272DRwGt02zokQiHHOl4HpoKezdyw59O?usp=sharing tabular1M<n<10M0 likes360 downloads8mo agoHugging Face14jacobmorrison /OpenThoughts3-456k-gpt4.1-cottabular100K<n<1M1 likes336 downloads1y agoHugging Face15domofon /Domofon-Cot-Conversations-700k Domofon-Cot-Conversations-700k Synthetic XML conversation data for training small language models on reasoning, instruction following, XML formatting, and tool-use traces. Repository: domofon/Domofon-Cot-Conversations-700k What is inside The dataset contains cleaned generated XML conversations from six families: conv: multi-turn factual conversations with tool-use traces. instruct: text-processing instructions, including deterministic count tool calls. ds:… See the full description on the dataset page: https://huggingface.co/datasets/domofon/Domofon-Cot-Conversations-700k.tabulartext-generation1M<n<10M1 likes335 downloads4mo agoHugging Face16HayatoHongoEveryonesAI /qa_verify_cot_new_5.1M_v7HayatoHongoEveryonesAI/qa_verify_cot_new_6M_unfiltered_v7 https://colab.research.google.com/drive/1tjJ14xLa0UZ0slYqnRuQR8ngk1sPyY8v?usp=sharing tabular1M<n<10M0 likes309 downloads8mo agoHugging Face17jacobmorrison /OpenThoughts3-456k-no-cottabular100K<n<1M0 likes253 downloads1y agoHugging Face18zhoudoe23 /chess-reasoning-cot-evalstabular1M<n<10M0 likes222 downloads29d agoHugging Face19teddyyyy123 /gpqa_0shot_cottabular1K<n<10K0 likes217 downloads2y agoHugging Face20Qipei /Task_Picupgloves_50fps_cotrain1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "trossen_subversion": "v1.0", "robot_type": "trossen_ai_mobile", "total_episodes": 51, "total_frames": 16676, "total_tasks": 1, "total_videos": 153, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:51" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/Qipei/Task_Picupgloves_50fps_cotrain1.tabularrobotics10K<n<100K0 likes216 downloads1y agoHugging Face21Crownelius /GLM-5.2-CoT-Library GLM-5.2 — CoT Library A maintained mirror of publicly-available GLM-5.2 chain-of-thought datasets on Hugging Face — content-verified, deduplicated, and attributed to their original authors. Dataset Viewer | Parquet // what this is A maintained library — a community mirror of publicly-available GLM-5.2 CoT datasets, aggregated, validity-filtered and content-verified, with per-row source attribution in first_source_dataset. It is not Crownelius' own data — every… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/GLM-5.2-CoT-Library.tabulartext-generation10K<n<100K2 likes216 downloads2mo agoHugging Face22gjoelbye /cot-hidden-state-trajectories CoT Hidden-State Trajectories Chain-of-thought traces and generation-time hidden-state activations from 11 open-weight language models, on Codeforces (competitive programming), Hendrycks MATH, and SATBench (Boolean satisfiability). This dataset accompanies the paper Reasoning Models Don't Just Think Longer, They Move Differently (arXiv:2605.15454). The paper asks whether reasoning-trained models follow different hidden-state paths than matched instruction-tuned baselines, after… See the full description on the dataset page: https://huggingface.co/datasets/gjoelbye/cot-hidden-state-trajectories.tabulartext-generation10K<n<100K0 likes192 downloads4mo agoHugging Face23Magpie-Align /Magpie-Reasoning-V2-250K-CoT-Llama3 Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Reasoning-V2-250K-CoT-Llama3.tabulartext-generation100K<n<1M11 likes190 downloads2y agoHugging Face24llamastack /mmlu_pro_cottabular10K<n<100K0 likes172 downloads1y agoHugging Face25HydraLM /CoT-Collection-standardized Dataset Card for "CoT-Collection-standardized" More Information needed tabular1M<n<10M3 likes168 downloads3y agoHugging Face26justicedao /ipfs_cotedivoire_laws_ir Cotedivoire legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_cotedivoire_laws (revision 67ef43d5e4b3053bccc68b9a6bebbbc1f4e98bd3) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of Cotedivoire prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No legal… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_cotedivoire_laws_ir.tabulartext-retrieval10K<n<100K0 likes166 downloads1d agoHugging Face27XumengWen /AIME24-25_CoT_Verification Dataset for ICLR 2026 Paper: Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs 📌 Dataset Summary This dataset contains the rollouts (reasoning traces) and verification results used in our ICLR 2026 paper. The data allows for the analysis of how Reinforcement Learning with Verifiable Rewards (RLVR) incentivizes the correct reasoning of Large Language Models (LLMs) on challenging mathematics benchmarks. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/XumengWen/AIME24-25_CoT_Verification.tabulartext-generation100K<n<1M1 likes159 downloads8mo agoHugging Face28ameek /measuring_cot_monitorability_transcripts Measuring Chain-of-Thought Monitorability Transcripts This dataset contains model transcripts from language models evaluated on MMLU, BIG-Bench Hard (BBH), and GPQA Diamond. Each sample group includes a baseline response (no cue) paired with five adaptive variations where different cues were injected to test chain-of-thought faithfulness. We use this dataset to measure how faithfully models represent their reasoning processes in their chain-of-thought outputs. By comparing baseline… See the full description on the dataset page: https://huggingface.co/datasets/ameek/measuring_cot_monitorability_transcripts.tabularquestion-answering100K<n<1M1 likes144 downloads10mo agoHugging Face29Crownelius /Kimi-K3-CoT-Library Kimi K3 — CoT Library A maintained mirror of publicly-available Kimi K3 chain-of-thought datasets on Hugging Face — content-verified, deduplicated, and attributed to their original authors. Dataset Viewer | Parquet // what this is A maintained library — a community mirror of publicly-available Kimi K3 CoT datasets, aggregated, validity-filtered and content-verified, with per-row source attribution in first_source_dataset. It is not Crownelius' own data — every… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Kimi-K3-CoT-Library.tabulartext-generation1K<n<10K2 likes143 downloads2mo agoHugging Face30cds-jb /cot-gemma4-26b-a4b Gemma-4-26B-A4B-it Chain-of-Thought Oracle Corpus Chain-of-thought rollouts generated with google/gemma-4-26B-A4B-it (MoE, 25.2B total / 3.8B active), in its native thinking mode, across a diverse suite of reasoning tasks. Structure follows ceselder/cot-oracle-corpus-v5 (CoT-only subset of the columns), built for chain-of-thought monitoring / activation-oracle research. 2,121,354 rollouts over 212,161 unique problems (10 sampled thinking rollouts per problem, temperature 0.8).… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/cot-gemma4-26b-a4b.tabulartext-generation1M<n<10M0 likes137 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.