CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01stindardlogic /math-reasoning-sft-100k Math Reasoning SFT (100K) 100,000 math problems with detailed step-by-step solutions — ready for supervised fine-tuning of math reasoning models. Dataset Description 100,000 problems across 8 mathematical categories and 3 difficulty levels: Categories Category Examples Topics word_problems ~23,100 Rate/time/distance, work problems, mixture, meeting/catch-up arithmetic ~15,400 Percentages, profit/loss, ratios geometry ~15,400 Area… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/math-reasoning-sft-100k.texttext-generation100K<n<1M1 likes654 downloads2mo agoHugging Face02mast-benchmark /100k-corpus-2026 MAST 100K Corpus 2026 This dataset contains the fixed English document corpus used for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages by retrieving English evidence and producing short, correct English answers. This corpus is copied from BrowseComp-Plus, a benchmark for Deep-Research systems that isolates the effect of the retriever and the LLM agent to… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/100k-corpus-2026.textquestion-answering100K<n<1M0 likes482 downloads2mo agoHugging Face03weijiezz /math-datasets-100k Merged Math Datasets (100k Subset) This dataset combines multiple mathematical datasets for training and evaluation purposes. This version contains a shuffled 100k subset of the training data for faster experimentation. Dataset Description A comprehensive collection of mathematical problems and solutions from various sources, organized into training and multiple test subsets. Dataset Structure Training Set Size: 100000 examples Fields: source… See the full description on the dataset page: https://huggingface.co/datasets/weijiezz/math-datasets-100k.textquestion-answering100K<n<1M0 likes354 downloads1y agoHugging Face04stindardlogic /medical-clinical-reasoning-sft-100k Medical Clinical Reasoning SFT 100K A synthetic supervised fine-tuning dataset of 100,000 high-quality medical and clinical reasoning conversations designed to train AI assistants capable of supporting clinical decision-making, documentation, and medical education. Dataset Description This dataset covers a broad spectrum of clinical practice scenarios across 10 medical specialty categories. Each record follows the ShareGPT conversation format with a detailed human… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/medical-clinical-reasoning-sft-100k.texttext-generation100K<n<1M0 likes340 downloads2mo agoHugging Face05stindardlogic /science-qa-sft-100k Science QA SFT (100K) 100,000 science Q&A examples with step-by-step explanations for SFT fine-tuning. Covers physics, chemistry, biology, astronomy, and earth science at beginner through advanced difficulty. Motivation Models trained on general text often give superficially plausible but mechanistically wrong answers to science questions — stating the right conclusion without understanding the underlying reasoning. This dataset trains models to explain why an… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/science-qa-sft-100k.texttext-generation100K<n<1M0 likes236 downloads2mo agoHugging Face06khoaliamle /MedDialog-EN-100ktextquestion-answering100K<n<1M0 likes129 downloads1y agoHugging Face07DJLougen /ornstein-curated-100k Ornstein Curated 100K A curriculum-sorted reasoning dataset for SFT and post-training experiments. Ornstein Curated 100K is a multi-domain instruction dataset built around explicit reasoning traces, difficulty progression, and curriculum-style ordering. The dataset contains 100,000 samples across mathematics, programming, conversational reasoning, and cognitive-science-inspired tasks. It is sorted from easier to harder examples so users can train with the provided order, compare… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/ornstein-curated-100k.texttext-generation100K<n<1M10 likes126 downloads3mo agoHugging Face08stindardlogic /cybersecurity-sft-100k Cybersecurity SFT 100K A synthetic supervised fine-tuning dataset of 100,000 high-quality cybersecurity conversations designed to train AI assistants for security operations, threat analysis, incident response, and defensive security engineering. Dataset Description This dataset covers real-world security scenarios across 9 cybersecurity domains. Each record follows the ShareGPT conversation format with a practitioner-level query and a detailed, structured… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/cybersecurity-sft-100k.texttext-generation100K<n<1M0 likes125 downloads2mo agoHugging Face09Jackrong /Chinese-Qwen3-235B-Thinking-2507-Distill-100k 📌 Note: The English translation of this dataset card is provided below. Chinese-Qwen3-235B-Thinking-2507-Distill-100k Dataset Summary Chinese-Qwen3-235B-Thinking-2507-Distill-100k 是一个包含约 100k 条高质量中文推理与指令数据的数据集,由 Qwen-3-235B-A22B-Thinking-2507(官方 Thinking 模式,上下文长度 32K)蒸馏生成。 该数据集覆盖了多个重要领域: 数学与工程任务(Mathematics, Applied Math, Advanced Math) 通用知识与写作(General Knowledge, Language & Writing) 技术与编程(Technology & Programming) 商业与经济(Business & Economics)… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100k.tabulartext-classification100K<n<1M19 likes120 downloads1y agoHugging Face10sujet-ai /Sujet-Finance-QA-Vision-100k Dataset Description 📊🔍 The Sujet-Finance-QA-Vision-100k is a comprehensive dataset containing over 100,000 question-answer pairs derived from more than 9,800 financial document images. This dataset is designed to support research and development in the field of financial document analysis and visual question answering. Key Features: 🖼️ 9,801 unique financial document images ❓ 107,050 question-answer pairs 🇬🇧 English language 📄 Diverse financial document types… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Finance-QA-Vision-100k.imagequestion-answering1K<n<10K39 likes102 downloads2y agoHugging Face11Hatman /plot-palette-100k Empowering Writers with a Universe of Ideas Plot Palette DataSet HuggingFace » Plot Palette was created to fine-tune large language models for creative writing, generating diverse outputs through iterative loops and seed data. It is designed to be run on a Linux system with systemctl for managing services. Included is the service structure, specific category prompts and ~100k data entries. The dataset is available here or… See the full description on the dataset page: https://huggingface.co/datasets/Hatman/plot-palette-100k.textquestion-answering10K<n<100K4 likes97 downloads2y agoHugging Face12stindardlogic /devops-kubernetes-sft-100k DevOps and Kubernetes SFT 100K A synthetic supervised fine-tuning dataset of 100,000 high-quality DevOps and Kubernetes conversations designed to train AI assistants capable of supporting platform engineers, SREs, and DevOps practitioners. Dataset Description This dataset covers production-grade Kubernetes operations, cloud infrastructure, CI/CD pipelines, GitOps workflows, and platform engineering across 13 specialized categories. Each record follows the ShareGPT… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/devops-kubernetes-sft-100k.texttext-generation100K<n<1M1 likes94 downloads2mo agoHugging Face1311-47 /AgentAngel_100k Within Us AI — AgentAngel_100k (Agentic Coding 2026) AgentAngel is a master-scholar, evidence-backed dataset family for training and evaluating agentic coding models that plan, patch, run checks, and iterate with tests-as-truth. This release contains 100,000 examples per split (500,000 JSONL rows total): Q&A (facts + rights/wrongs) Instruct (messages) Thinking (concise rationales) Reasoning (constraints + verification checks) Chat (multi-turn) Evidence discipline… See the full description on the dataset page: https://huggingface.co/datasets/11-47/AgentAngel_100k.text-generation0 likes85 downloads9mo agoHugging Face14weijiezz /NuminaMath-100k Merged Math Datasets (100k Subset) This dataset combines multiple mathematical datasets for training and evaluation purposes. This version contains a shuffled 100k subset of the training data for faster experimentation. Dataset Description A comprehensive collection of mathematical problems and solutions from various sources, organized into training and multiple test subsets. Dataset Structure Training Set Size: 100000 examples Fields: source… See the full description on the dataset page: https://huggingface.co/datasets/weijiezz/NuminaMath-100k.textquestion-answering100K<n<1M0 likes84 downloads1y agoHugging Face15lossisnotanumber /browsecomp-plus-100k-corpus-as-local-folder BrowseComp-Plus 100K Corpus — as local folder The 100K-document subset of the BrowseComp-Plus benchmark corpus, processed into a plain document-directory form using the browsecomp-plus processing code from DCI-Agent-Lite. This repository stores the corpus exactly as it is expected on local disk: a flat tree of <domain>/<title>.txt files, ready to be pointed at by --corpus-dir. It is the form consumed by the RARG / DCI-Agent retrieval-augmented agent during the BrowseComp-Plus… See the full description on the dataset page: https://huggingface.co/datasets/lossisnotanumber/browsecomp-plus-100k-corpus-as-local-folder.text-retrieval100K<n<1M0 likes77 downloads2mo agoHugging Face16stindardlogic /mlops-deployment-sft-100k MLOps Deployment SFT 100K A synthetic supervised fine-tuning dataset of 100,000 high-quality conversations covering MLOps and ML model deployment — from model serving and inference optimization to monitoring, CI/CD, and production scaling. Designed to train AI assistants that can help ML engineers deploy and operate models at scale. Dataset Description This dataset covers the full MLOps lifecycle across 13 specialized categories. Each record follows the ShareGPT… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/mlops-deployment-sft-100k.texttext-generation100K<n<1M2 likes71 downloads2mo agoHugging Face17lcw99 /DistilQwen_100k_korean DistilQwen 100k Korean This dataset is a Korean translation of the original alibaba-pai/DistilQwen_100k dataset. Dataset Structure The dataset contains both English and Korean versions of instruction-response pairs: { "instruction": "Original English instruction text", "output": "Original English response/answer", "instruction_kr": "Korean translation of the instruction", "output_kr": "Korean translation of the response/answer", "_dataset_index": 30000 } Each… See the full description on the dataset page: https://huggingface.co/datasets/lcw99/DistilQwen_100k_korean.texttext-generation10K<n<100K1 likes63 downloads1y agoHugging Face18davidfoss /bitcoin-security-reasoning-100k Dataset Card for Bitcoin Security Reasoning 100K 100,000 high-quality synthetic training samples for fine-tuning LLMs on Bitcoin protocol security analysis. Teaches models to analyze vulnerability clusters, form security hypotheses, and generate differential testing code. Dataset Details Dataset Description This dataset contains structured security reasoning chains for Bitcoin protocol vulnerabilities. Each sample presents a cluster of causal… See the full description on the dataset page: https://huggingface.co/datasets/davidfoss/bitcoin-security-reasoning-100k.texttext-generation100K<n<1M0 likes63 downloads8mo agoHugging Face19sfd-anonymous /sefd-archive-100k-analysis-sample-qwen3-20260524 SEFD Archive 100k Analysis Sample Qwen3 20260524 Retained artifacts for the completed archive-wide 100,000-filing Stanford EDGAR Filings Dataset (SEFD) analysis sample used in the arXiv paper update. The sample contains 2,971,490,909 final SEFD tokens, counted with the Qwen3-1.7B tokenizer. This repository is a new versioned artifact and intentionally does not replace the earlier sfd-archive-100k-analysis-sample repository used for the original conference submission. Included:… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/sefd-archive-100k-analysis-sample-qwen3-20260524.text-generation0 likes61 downloads4mo agoHugging Face20alfredplpl /wikipedia-qa-ja-100k Dataset Card for "wikipedia-qa-ja-100k" Original Dataset hpprc/wikipedia-20240101 Procedure Extract the first line of the title from the dataset. Generate the answer by summizing the line using LLM: Input RAG-like prompt to CALM 2 7B Chat. Format the response. RAG-like Prompt f"""USER: {title}とはなんですか?次の文章を参考に一言でまとめてください。{text} ASSISTANT: """ textquestion-answering100K<n<1M3 likes59 downloads3y agoHugging Face21rlhn /rlhn-100K Dataset Card for RLHN-100K Dataset Description Repository | Paper | ArXiv RLHN is a cascading LLM framework designed to accurately relabel hard negatives in existing IR/RAG training datasets, such as MS MARCO and HotpotQA. This Tevatron dataset (100K training pairs) contains the queries, positives + relabeled hard negatives, remaining hard negatives for 7 datasets in the BGE training collection. This repository contains the training pairs that can be used to fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/rlhn/rlhn-100K.textquestion-answering10K<n<100K1 likes58 downloads1y agoHugging Face22WithinUsAI /Genesis_AI_Code_100k Genesis AI Code 100K (Frontier) Developed by: Within Us AI Frontier dataset with tool-call traces, self-grading, budgets, and audit orientation. Splits train: 98,000 validation: 2,000 Highlights Tests-as-truth supervision patterns Diff-first patching Agentic loops (plan→edit→test→reflect) with bounded budgets Tool-call trace supervision (where present) Governance/audit & policy-gate awareness Storage format Parquet unavailable (No… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_100k.texttext-generation10K<n<100K1 likes51 downloads9mo agoHugging Face23bysismo /100k_Tdk_zurriyet_dna_v6.jsonl 🌟 DESTEK & TOPLULUK ÇAĞRISI (SUPPORT & LIKE):Açık kaynak ve ücretsiz olarak sunduğum bu devasa çalışmayı faydalı bulduysanız, projenin sürdürülebilirliğine ve açık kaynak ekosisteminin görünürlüğüne katkı sağlamak için lütfen sayfanın sağ üstündeki Like (❤️ Beğeni) butonuna basarak destek olmayı unutmayın!(If you find this open-source dataset valuable for your research or models, please consider leaving a ❤️ Like at the top-right to support future updates and maintenance). 🇹🇷… See the full description on the dataset page: https://huggingface.co/datasets/bysismo/100k_Tdk_zurriyet_dna_v6.jsonl.textquestion-answering10K<n<100K1 likes50 downloads1mo agoHugging Face24gravermistakes /Genesis_AI_Code_100k Genesis AI Code 100K (Frontier) Developed by: Within Us AI Frontier dataset with tool-call traces, self-grading, budgets, and audit orientation. Splits train: 98,000 validation: 2,000 Highlights Tests-as-truth supervision patterns Diff-first patching Agentic loops (plan→edit→test→reflect) with bounded budgets Tool-call trace supervision (where present) Governance/audit & policy-gate awareness Storage format Parquet unavailable (No module named… See the full description on the dataset page: https://huggingface.co/datasets/gravermistakes/Genesis_AI_Code_100k.texttext-generation10K<n<100K0 likes50 downloads7mo agoHugging Face25deep-div /HealthTalks-100ktexttable-question-answering100K<n<1M2 likes49 downloads1y agoHugging Face26philschmid /slimorca-dedup-chatml-100k Copy of Open-Orca/SlimOrca-Dedup in ChatML format downsample to 100k "SlimOrca Dedup" is a deduplicated, unfiltered subset of the SlimOrca dataset, excluding RLHF instances, resulting in 363k unique examples. Key Features Removal of RLHF instances. Deduplication using minhash and Jaccard similarity techniques. Demo Models Note: These models were trained on the full SlimOrca dataset, not the deduplicated, unfiltered version. *… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/slimorca-dedup-chatml-100k.texttext-classification100K<n<1M2 likes45 downloads3y agoHugging Face27stindardlogic /rag-systems-sft-100k RAG Systems SFT 100K A synthetic supervised fine-tuning dataset of 100,000 high-quality conversations covering Retrieval-Augmented Generation (RAG) systems — from basic pipelines to advanced multi-hop retrieval, evaluation, and production optimization. Designed to train AI assistants that can help engineers build, debug, and scale RAG applications. Dataset Description This dataset covers the full spectrum of RAG system development across 12 specialized categories.… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/rag-systems-sft-100k.texttext-generation100K<n<1M0 likes45 downloads2mo agoHugging Face28rlhn /hn-remove-100K Dataset Card for HN-Remove 100K Dataset Description Repository | Paper | ArXiv RLHN is a cascading LLM framework designed to accurately relabel hard negatives in existing IR/RAG training datasets, such as MS MARCO and HotpotQA. This Tevatron dataset (100K training pairs) contains the queries, positives, hard negatives (with dropped false negatives) for 7 datasets in the BGE training collection. This repository contains the training pairs that can be used to fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/rlhn/hn-remove-100K.textquestion-answering10K<n<100K0 likes41 downloads1y agoHugging Face29PinkPixel /Reasoning-Mix-100k 🧠 Reasoning Mix 100k ✨ A high-quality, balanced reasoning dataset consisting of 99,999 samples extracted from three reasoning datasets. This dataset is specifically formatted for models to utilize a /think block for step-by-step reasoning. 📊 Dataset Summary The Reasoning Mix 100k is a curated collection of reasoning tasks, primarily focused on mathematics, logic, and general problem-solving. It combines the strengths of three high-performing datasets into a unified… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/Reasoning-Mix-100k.texttext-generation10K<n<100K1 likes39 downloads5mo agoHugging Face30junaid008 /Pashto-100k-Pairs Qehwa AI - Pashto 100K Fine-Tuning Dataset Overview Qehwa AI presents a large-scale Pashto instruction tuning dataset containing 100,000+ high-quality instruction-response pairs designed for supervised fine-tuning, conversational AI, and downstream NLP tasks. This dataset was created to advance AI research for the Pashto language, a significantly underrepresented low-resource language spoken by millions worldwide. The dataset covers more than 20 diverse domains and is… See the full description on the dataset page: https://huggingface.co/datasets/junaid008/Pashto-100k-Pairs.texttext-generation100K<n<1M6 likes38 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.