CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Xkev /LLaVA-CoT-100k Dataset Card for LLaVA-CoT The LLaVA-CoT-100k dataset is introduced in the paper LLaVA-CoT: Let Vision Language Models Reason Step-by-Step. This dataset is designed to enable Vision-Language Models (VLMs) to perform autonomous multistage reasoning, integrating samples from various visual question-answering sources with structured reasoning annotations. It aims to address the challenges VLMs face in systematic and structured reasoning for complex visual question-answering tasks.… See the full description on the dataset page: https://huggingface.co/datasets/Xkev/LLaVA-CoT-100k.textvisual-question-answering10K<n<100K106 likes1.9k downloads9mo agoHugging Face02wangrongsheng /HealthCareMagic-100k-entext100K<n<1M22 likes880 downloads3y agoHugging Face03MiG-NJU /OmniVideo-100K OmniVideo-100K Official repository for OmniVideo-100K, an instruction-tuning dataset introduced in our paper: "OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains". This repository includes: videos.tar.part_xx: Raw video files. train_oe_70k.jsonl: Original Open-Ended (OE) training samples. train_mcq_30k.jsonl: Original Multiple-Choice (MCQ) training samples. train_oe_70k_formatted.jsonl: Instruction-formatted OE samples (ready for… See the full description on the dataset page: https://huggingface.co/datasets/MiG-NJU/OmniVideo-100K.textvideo-text-to-text100K<n<1M18 likes794 downloads3mo agoHugging Face04stindardlogic /math-reasoning-sft-100k Math Reasoning SFT (100K) 100,000 math problems with detailed step-by-step solutions — ready for supervised fine-tuning of math reasoning models. Dataset Description 100,000 problems across 8 mathematical categories and 3 difficulty levels: Categories Category Examples Topics word_problems ~23,100 Rate/time/distance, work problems, mixture, meeting/catch-up arithmetic ~15,400 Percentages, profit/loss, ratios geometry ~15,400 Area… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/math-reasoning-sft-100k.texttext-generation100K<n<1M1 likes652 downloads2mo agoHugging Face05Mohammed-Altaf /medical-instruction-100k What is the Dataset About?🤷🏼‍♂️ The dataset is useful for training a Generative Language Model for the Medical application and instruction purposes, the dataset consists of various thoughs proposed by the people [mentioned as the Human ] and there responses including Medical Terminologies not limited to but including names of the drugs, prescriptions, yogic exercise suggessions, breathing exercise suggessions and few natural home made prescriptions. How the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Mohammed-Altaf/medical-instruction-100k.text100K<n<1M17 likes566 downloads3y agoHugging Face06maplebb /UniREdit-Data-100KUniREditBench: A Unified Reasoning-based Image Editing Benchmark text10K<n<100K3 likes507 downloads10mo agoHugging Face07stindardlogic /tool-calling-english-100k Tool Calling English (100K) 100,000 tool-calling conversations in OpenAI function calling format — the largest general English tool-use dataset for fine-tuning. Motivation Models trained without tool-calling examples struggle in agentic deployments. This dataset trains the full cycle: deciding when to call a tool, calling it with correct arguments, interpreting the result, and producing a grounded final response. Dataset Description 100,000… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/tool-calling-english-100k.texttext-generation100K<n<1M1 likes386 downloads2mo agoHugging Face08Eurolingua /DCLM-200-100k-exact-deduptext10M<n<100M0 likes361 downloads10mo agoHugging Face09Alignment-Lab-AI /Expert-Sudoku-100ktabular100K<n<1M0 likes348 downloads2y agoHugging Face10RafaelMPereira /HealthCareMagic-100k-Chat-Format-entext100K<n<1M8 likes324 downloads3y agoHugging Face11stindardlogic /medical-clinical-reasoning-sft-100k Medical Clinical Reasoning SFT 100K A synthetic supervised fine-tuning dataset of 100,000 high-quality medical and clinical reasoning conversations designed to train AI assistants capable of supporting clinical decision-making, documentation, and medical education. Dataset Description This dataset covers a broad spectrum of clinical practice scenarios across 10 medical specialty categories. Each record follows the ShareGPT conversation format with a detailed human… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/medical-clinical-reasoning-sft-100k.texttext-generation100K<n<1M0 likes321 downloads2mo agoHugging Face12MBZUAI /VideoInstruct-100KVideoInstruct100K is a high-quality video conversation dataset generated using human-assisted and semi-automatic annotation techniques. The question answers in the dataset are related to, Video Summariazation Description-based question-answers (exploring spatial, temporal, relationships, and reasoning concepts) Creative/generative question-answers For mored details, please visit Oryx/VideoChatGPT/video-instruction-data-generation. If you find this dataset useful, please consider citing the… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/VideoInstruct-100K.text100K<n<1M49 likes315 downloads3y agoHugging Face13stindardlogic /writing-quality-dpo-100k Writing Quality DPO (100K) 100,000 DPO preference pairs training models to write with clarity, concision, structure, and impact. Each chosen response demonstrates high-quality prose; each rejected response contains exactly one identified writing defect. Motivation Writing assistance is the #1 use case for LLMs, yet most training data optimizes for factual correctness rather than writing craft. This dataset trains models to distinguish genuinely good writing from… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/writing-quality-dpo-100k.texttext-generation100K<n<1M0 likes307 downloads2mo agoHugging Face14stindardlogic /science-qa-sft-100k Science QA SFT (100K) 100,000 science Q&A examples with step-by-step explanations for SFT fine-tuning. Covers physics, chemistry, biology, astronomy, and earth science at beginner through advanced difficulty. Motivation Models trained on general text often give superficially plausible but mechanistically wrong answers to science questions — stating the right conclusion without understanding the underlying reasoning. This dataset trains models to explain why an… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/science-qa-sft-100k.texttext-generation100K<n<1M0 likes234 downloads2mo agoHugging Face15stindardlogic /email-writing-sft-100k Email Writing SFT (100K) 100,000 ShareGPT conversations demonstrating professional email writing across 22 business contexts. Each example shows how to draft clear, purposeful emails that achieve their communication goal — from cold outreach to salary negotiations to apology emails. Motivation Email is the primary communication channel for most professional work, yet LLMs often produce emails that are: Too long: Including unnecessary preamble, excessive context… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/email-writing-sft-100k.texttext-generation100K<n<1M3 likes229 downloads2mo agoHugging Face16iecjsu /lavita-ChatDoctor-HealthCareMagic-100ktext10K<n<100K1 likes215 downloads2y agoHugging Face17jjjlimaus /sn38-quality-gold-100kgated SN38 quality prompt + gold continuation Synthetic incomplete-sentence prompts with gold continuations for Bittensor subnet 38. 13 categories, 13 items per category per call, temperature 1.0. Each row: category: reading_comprehension, language_understanding, world_knowledge, commonsense_reasoning, language_modeling, causal_reasoning, logical_inference, temporal_reasoning, math_reasoning, truthfulness, pronoun_resolution, paraphrase_detection, word_sense_disambiguation prompt:… See the full description on the dataset page: https://huggingface.co/datasets/jjjlimaus/sn38-quality-gold-100k.texttext-generation1M<n<10M0 likes207 downloads13d agoHugging Face18ruihangxu /IMIG-100K IMIG-100K: A Large-Scale Synthetic Dataset for Multi-Instance Image Generation with Detailed Annotation Ruihang Xu, Dewei Zhou, Fan Ma†, Yi Yang ReLER Lab, CCAI, Zhejiang University 📄 Dataset Overview The IMIG-100K dataset is a large-scale synthetic dataset designed for multi-instance image generation tasks. It contains more than 100,000 high-quality image samples, each annotated with masks and layout information. The dataset is organized into several sub-datasets… See the full description on the dataset page: https://huggingface.co/datasets/ruihangxu/IMIG-100K.texttext-to-image100K<n<1M6 likes195 downloads7mo agoHugging Face19stindardlogic /technical-writing-sft-100k Technical Writing SFT (100K) 100,000 ShareGPT conversations demonstrating high-quality technical writing across 20 document types. Each example produces a complete, professional technical document — from API reference to architecture decision records to runbooks — written in the style that experienced technical writers and senior engineers actually use. Motivation Technical writing is one of the most underserved capabilities in LLMs. Common model failures: Wrong… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/technical-writing-sft-100k.texttext-generation100K<n<1M0 likes187 downloads2mo agoHugging Face20PKU-Alignment /Align-Anything-Instruction-100K-zh Dataset Card for Align-Anything-Instruction-100K-zh [🏠 Homepage] [🤗 Instruction-Dataset-100K(en)] [🤗 Instruction-Dataset-100K(zh)] [🤗 Align-Anything Datasets] Instruction-Dataset-100K(zh) Highlights Data sources: Firefly (47.8%), COIG (2.9%), and our meticulously constructed QA pairs (49.3%). 100K QA pairs (zh): 104,550 meticulously crafted instructions, selected and polished from various Chinese datasets… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Align-Anything-Instruction-100K-zh.texttext-generation100K<n<1M10 likes162 downloads2y agoHugging Face21barc0 /100k-gpt4omini-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds. We generate the dataset with the following steps: Generate 120k descriptions by GPT4o-mini. Generate 120k codes follow each description by GPT4o-mini. Run the 120k codes and do auto-filtering. Get the final 100k legitimate ARC-like tasks with examples. texttext-generation100K<n<1M1 likes158 downloads2y agoHugging Face22stindardlogic /instruction-following-hard-sft-100k Hard Instruction Following SFT (100K) 100,000 ShareGPT conversations where the assistant correctly satisfies multiple simultaneous explicit constraints in a single response. Each example pairs a multi-constraint prompt with a response that honors every constraint without dropping any. Targets the instruction-following capability measured by IFEval and similar benchmarks. Motivation A key failure mode in deployed LLMs is dropping constraints under load — responding… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/instruction-following-hard-sft-100k.texttext-generation100K<n<1M0 likes142 downloads2mo agoHugging Face23barc0 /100k-gpt4-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds. We generate the dataset with the following steps: Generate 120k descriptions by GPT4. Generate 120k codes follow each description by GPT4o-mini. Run the 120k codes and do auto-filtering. Get the final 100k legitimate ARC-like tasks with examples. texttext-generation100K<n<1M1 likes135 downloads2y agoHugging Face24jumplander /JumpLander-PMB-100K 🚀 JumpLander-PMB-100K Persian Model Behavior Dataset for Intent, Constraint, and Safe Response Evaluation جامپ‌لندر PMB-100K | مجموعه‌داده فارسی برای سنجش رفتار مدل، فهم نیت، رعایت محدودیت و پاسخ امن Built by JumpLander Official Website: jumplander.orgPersian Website: jumplander.org/faDocumentation: jumplander.org/fa/docsAbout JumpLander: jumplander.org/fa/aboutSupport JumpLander: jumplander.org/fa/rateHugging Face Organization:… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JumpLander-PMB-100K.texttext-generation100K<n<1M6 likes133 downloads3mo agoHugging Face25stindardlogic /cybersecurity-sft-100k Cybersecurity SFT 100K A synthetic supervised fine-tuning dataset of 100,000 high-quality cybersecurity conversations designed to train AI assistants for security operations, threat analysis, incident response, and defensive security engineering. Dataset Description This dataset covers real-world security scenarios across 9 cybersecurity domains. Each record follows the ShareGPT conversation format with a practitioner-level query and a detailed, structured… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/cybersecurity-sft-100k.texttext-generation100K<n<1M0 likes131 downloads2mo agoHugging Face26Jackrong /Chinese-Qwen3-235B-Thinking-2507-Distill-100k 📌 Note: The English translation of this dataset card is provided below. Chinese-Qwen3-235B-Thinking-2507-Distill-100k Dataset Summary Chinese-Qwen3-235B-Thinking-2507-Distill-100k 是一个包含约 100k 条高质量中文推理与指令数据的数据集,由 Qwen-3-235B-A22B-Thinking-2507(官方 Thinking 模式,上下文长度 32K)蒸馏生成。 该数据集覆盖了多个重要领域: 数学与工程任务(Mathematics, Applied Math, Advanced Math) 通用知识与写作(General Knowledge, Language & Writing) 技术与编程(Technology & Programming) 商业与经济(Business & Economics)… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100k.tabulartext-classification100K<n<1M19 likes128 downloads1y agoHugging Face27Nithish2410 /recommendations-ml-100k MovieLens Leave-One-Out Five chronological interactions predict the next interaction. One final test target per user; no rating filter. Histories are audit-only, not wholesale model inputs. Actors are supplementary; see actor_sources.json. { "schema": "movie-fields-v1", "source": "official MovieLens 100K", "sample_policy": "leave-one-out-windows", "past_order": "oldest-first", "history_length": 5, "stride": 1, "shuffle_seed": 42, "timestamp_policy": "rating… See the full description on the dataset page: https://huggingface.co/datasets/Nithish2410/recommendations-ml-100k.text10K<n<100K0 likes127 downloads1d agoHugging Face28PKU-Alignment /Align-Anything-Instruction-100K Dataset Card for Align-Anything-Instruction-100K [🏠 Homepage] [🤗 Instruction-Dataset-100K(en)] [🤗 Instruction-Dataset-100K(zh)] [🤗 Align-Anything Datasets] Highlights Data sources: PKU-SafeRLHF QA , DialogSum, Empathetic, Instruction-Wild, and Alpaca. 100K QA pairs: By leveraging GPT-4 to annotate meticulously refined instructions, we obtain 105,333 QA pairs.… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Align-Anything-Instruction-100K.texttext-generation100K<n<1M9 likes118 downloads2y agoHugging Face29NilanE /ParallelFiction-Ja_En-100k Dataset details: Each entry in this dataset is a sentence-aligned Japanese web novel chapter and English fan translation. The intended use-case is for document translation tasks. Dataset format: { 'src': 'JAPANESE WEB NOVEL CHAPTER', 'trg': 'CORRESPONDING ENGLISH TRANSLATION', 'meta': { 'general': { 'series_title_eng': 'ENGLISH SERIES TITLE', 'series_title_jap': 'JAPANESE SERIES TITLE', 'sentence_alignment_score':… See the full description on the dataset page: https://huggingface.co/datasets/NilanE/ParallelFiction-Ja_En-100k.texttranslation100K<n<1M82 likes116 downloads2y agoHugging Face30khoaliamle /MedDialog-EN-100ktextquestion-answering100K<n<1M0 likes116 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.