CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lfaviate /China-K12-STEM-10K-CoT-Reasoning K12-STEM-CoT-Chinese 1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams. The largest structured Chinese math/physics/chemistry reasoning dataset. This is a curated sample (10,000 problems) of the full 1.54M dataset available via API. Full Dataset Access Access the full 1,540,000+ problems via API → This Sample Full API Total problems 10,025 1,540,000+ With CoT solutions 10,025 1,490,000+ With diagrams 6,093 740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.tabularquestion-answering10K<n<100K3 likes605 downloads7mo agoHugging Face02zhoudoe23 /chess-reasoning-cot-evalstabular1M<n<10M0 likes218 downloads28d agoHugging Face03Magpie-Align /Magpie-Reasoning-V2-250K-CoT-Llama3 Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Reasoning-V2-250K-CoT-Llama3.tabulartext-generation100K<n<1M11 likes189 downloads2y agoHugging Face04isaiahbjork /cot-logic-reasoningtabular10K<n<100K18 likes111 downloads2y agoHugging Face05miscovery /Math_CoT_Arabic_English_Reasoning Math CoT Arabic English Dataset A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI. Overview Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.tabularquestion-answering1K<n<10K17 likes73 downloads1y agoHugging Face06expertdata-factory /cybersecurity-reasoning-cot-v1 🛡️ Expert Cybersecurity Reasoning Dataset (CoT) This dataset contains 89 high-fidelity, expert-verified reasoning records focusing on complex cybersecurity attack vectors. It is designed specifically for fine-tuning Large Language Models (LLMs) on sophisticated security analysis and threat logic. 💎 Key Highlights Niche Rarity 1.0: Covers rare and emerging threats with zero prior representation in open-source datasets. Advanced Vectors: Includes detailed reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/expertdata-factory/cybersecurity-reasoning-cot-v1.tabulartext-generationn<1K2 likes63 downloads7mo agoHugging Face07a13905873166 /China-K12-STEM-10K-CoT-Reasoning K12-STEM-CoT-Chinese 1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams. The largest structured Chinese math/physics/chemistry reasoning dataset. This is a curated sample (10,000 problems) of the full 1.54M dataset available via API. Full Dataset Access Access the full 1,540,000+ problems via API → This Sample Full API Total problems 10,025 1,540,000+ With CoT solutions 10,025 1,490,000+ With diagrams 6,093 740… See the full description on the dataset page: https://huggingface.co/datasets/a13905873166/China-K12-STEM-10K-CoT-Reasoning.tabularquestion-answering10K<n<100K1 likes53 downloads9d agoHugging Face08reasoningMIA /aime25_amc23_gpqa_olymp_cottabular10K<n<100K0 likes40 downloads1y agoHugging Face09open-llm-leaderboard /EpistemeAI__Reasoning-Llama-3.1-CoT-RE1-NMT-detailsgated Dataset Card for Evaluation run of EpistemeAI/Reasoning-Llama-3.1-CoT-RE1-NMT Dataset automatically created during the evaluation run of model EpistemeAI/Reasoning-Llama-3.1-CoT-RE1-NMT The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Reasoning-Llama-3.1-CoT-RE1-NMT-details.tabular10K<n<100K0 likes29 downloads2y agoHugging Face10francescortu /cot-oracle-reasoning-termination cot-oracle-reasoning-termination Mirror of japhba/cot-oracle-reasoning-termination for reproducibility (orig org fragile). tabular10K<n<100K0 likes26 downloads3mo agoHugging Face11ceselder /cot-oracle-eval-reasoning-termination-riya CoT Oracle Eval: reasoning_termination_riya Reasoning termination prediction — given a CoT prefix, predict whether the model will emit within the next 100 tokens. Labels are resampled (50 continuations per prefix): will_terminate if >=45/50 end within 20-60 tokens, will_continue if >=45/50 continue beyond 200 tokens. Includes Wilson CIs on resample counts. 50/50 balanced. Source: AI-MO/aimo-validation-aime + AI-MO/aimo-validation-amc (no overlap with GSM8K/MATH training data). Part… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/cot-oracle-eval-reasoning-termination-riya.tabularn<1K0 likes19 downloads7mo agoHugging Face12LLMTeamAkiyama /cleand_moremilk_CoT_Reasoning_Quantom_Physics_And_Computing元データ: https://huggingface.co/datasets/moremilk/CoT_Reasoning_Quantom_Physics_And_Computing 使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT_Reasoning_Quantom_Physics_And_Computing データ件数: 2,862 平均トークン数: 1,110 最大トークン数: 2,334 合計トークン数: 3,175,666 ファイル形式: JSONL ファイル分割数: 1 合計ファイルサイズ: 15.5 MB 加工内容: メタデータ列の解析と新列生成: metadata列(辞書型)を解析し、その中のreasoningをthought列に、difficultyをdifficulty列に展開しました。解析に失敗した行は除外されました。また、元のmetadata列は削除されました。 難易度によるフィルタリング:… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_moremilk_CoT_Reasoning_Quantom_Physics_And_Computing.tabularquestion-answering1K<n<10K0 likes13 downloads1y agoHugging Face13ceselder /cot-oracle-reasoning-termination-balancedtabular10K<n<100K0 likes13 downloads7mo agoHugging Face14LLMTeamAkiyama /cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research元データ: https://huggingface.co/datasets/moremilk/CoT_Reasoning_Scientific_Discovery_and_Research 使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT_Reasoning_Scientific_Discovery_and_Research データ件数: 3,733 平均トークン数: 1,193 最大トークン数: 2,489 合計トークン数: 4,453,517 ファイル形式: JSONL ファイル分割数: 1 合計ファイルサイズ: 23.2 MB 加工内容: メタデータ列の解析と新列生成: metadata列(辞書型)を解析し、その中のreasoningをthought列に、difficultyをdifficulty列に展開しました。解析に失敗した行は除外されました。また、元のmetadata列は削除されました。 難易度によるフィルタリング:… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research.tabularquestion-answering1K<n<10K0 likes11 downloads1y agoHugging Face15open-llm-leaderboard /EpistemeAI__Reasoning-Llama-3.1-CoT-RE1-NMT-V2-ORPO-detailsgated Dataset Card for Evaluation run of EpistemeAI/Reasoning-Llama-3.1-CoT-RE1-NMT-V2-ORPO Dataset automatically created during the evaluation run of model EpistemeAI/Reasoning-Llama-3.1-CoT-RE1-NMT-V2-ORPO The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Reasoning-Llama-3.1-CoT-RE1-NMT-V2-ORPO-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face161Happy-neuron /Quant-CoT-Factor-Reasoning-PreviewgatedQuantitative Factor Generation: Chain-of-Thought (CoT) Trajectories Dataset Description This is a 100-episode preview of a proprietary Reinforcement Learning from Environment Feedback (RLEF) dataset. It is designed to fine-tune Large Language Models (LLMs) for institutional quantitative finance, specifically systematic factor discovery and vectorized Python execution. The Architecture The data captures multi-turn agentic loops where the LLM: Formulates a cross-sectional equity factor… See the full description on the dataset page: https://huggingface.co/datasets/1Happy-neuron/Quant-CoT-Factor-Reasoning-Preview.tabularn<1K0 likes8 downloads1mo agoHugging Face17japhba /cot-oracle-reasoning-terminationtabular10K<n<100K0 likes7 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.