CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Trendyol /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K134 likes5.7k downloads1y agoHugging Face02FreedomIntelligence /TCM-Instruction-Tuning-ShizhenGPT 📚 Introduction This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced fine-tuning dataset consists of three parts: Modality Data Quantity TCM Text Instructions 📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT.textquestion-answering100K<n<1M13 likes420 downloads1y agoHugging Face03botcoinmoney /domain-agnostic-causal-reasoning-tuning Domain-Agnostic Causal Reasoning Tuning Dataset Training data for fine-tuning language models on multi-hop document reasoning. Each example is a graded reasoning trace produced by a frontier AI agent solving a procedurally generated challenge from the Botcoin proof-of-inference network. The traces contain no real domain knowledge. Entities are fictional, numbers are random, and documents are generated deterministically from 128-bit seeds. The reasoning structure is what matters:… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/domain-agnostic-causal-reasoning-tuning.textquestion-answering10K<n<100K1 likes164 downloads6mo agoHugging Face04SmartQHSE /hse-instruction-tuning SmartQHSE HSE Instruction-Tuning Corpus Alpaca-style instruction-tuning variant of the SmartQHSE HSE Q&A Corpus. Each row has the canonical fine-tuning schema: { "instruction": "What is OSHA Process Safety Management 1910.119?", "input": "", "output": "<authoritative long-form answer with citations>", "category": "us-osha", "source_url": "https://www.smartqhse.com/answers/<slug>" } Suitable for LoRA / SFT training of HSE-domain LLMs and RAG systems. Citation (preferred… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/hse-instruction-tuning.text-generationn<1K0 likes150 downloads4mo agoHugging Face05alxfgh /ChEMBL_Drug_Instruction_Tuning Dataset Card for ChEMBL Drug Instruction Tuning Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/alxfgh/ChEMBL_Drug_Instruction_Tuning.textquestion-answering100K<n<1M15 likes120 downloads3y agoHugging Face06CarsonnnNN /TCM-Instruction-Tuning-ShizhenGPT 📚 Introduction This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced fine-tuning dataset consists of three parts: Modality Data Quantity TCM Text Instructions 📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/CarsonnnNN/TCM-Instruction-Tuning-ShizhenGPT.textquestion-answering100K<n<1M0 likes107 downloads7mo agoHugging Face07markendo /Visual-Extraction-Tuning-382K Visual Extraction Tuning 382K This repository contains the generated visual extraction tuning dataset from the paper Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models. Project page: https://web.stanford.edu/~markendo/projects/downscaling_intelligence Code: https://github.com/markendo/downscaling_intelligence Overview We provide the 382K examples generated using our visual extraction tuning data generation pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/markendo/Visual-Extraction-Tuning-382K.textvisual-question-answering100K<n<1M0 likes85 downloads10mo agoHugging Face08IndoHealth-NLP /NCD_Instruct-Tuning_Medical_QA_Indonesian IndoHealth-NLP Vol. 2: NCD Instruct-Tuning Medical QA (Sample) 📁 VIEW & DOWNLOAD SAMPLE FILES HERE ⚠️ DATASET LIMITATION NOTE: This repository contains a FREE SAMPLE (200 rows) for evaluation purposes. To download the full, production-ready dataset containing 3,497 meticulously curated rows, please visit our official Gumroad page: [https://3929431511879.gumroad.com/l/IndoHealth-NLPVol2NCDInstruct-TuningMedicalQAIndonesian] Dataset Summary Building localized… See the full description on the dataset page: https://huggingface.co/datasets/IndoHealth-NLP/NCD_Instruct-Tuning_Medical_QA_Indonesian.textquestion-answeringn<1K1 likes80 downloads10d agoHugging Face09Star-gazer /medical_instruction_tuningtextquestion-answering10K<n<100K0 likes67 downloads2y agoHugging Face10MohamedAhmedAE /Med_LLaMa3_fine-tuning_dataset Med-LLaMA3 — Medical Instruction Fine-Tuning Dataset A large, unified medical instruction-tuning corpus (~1.65 million examples) compiled, cleaned, and standardized from a diverse set of public medical sources. It is the training corpus used to fine-tune the Med-LLaMA3 family (LLaMA-3.2 1B/3B and LLaMA-3.1 8B) in the paper “Med-LLaMA3: Advancing Medical Question-Answering Through Parameter-Efficient Fine-Tuning of Large Language Models” (Applied Sciences, 2026). All sources were… See the full description on the dataset page: https://huggingface.co/datasets/MohamedAhmedAE/Med_LLaMa3_fine-tuning_dataset.textquestion-answering1M<n<10M1 likes67 downloads3mo agoHugging Face11Fredithefish /Instruction-Tuning-with-GPT-4-RedPajama-Chat Instruction Tuning with GPT 4 RedPajama-Chat This dataset has been converted from the Instruction-Tuning-with-GPT-4 dataset for the purpose of fine-tuning the RedPajama-INCITE-Chat-3B-v1 model. About Instruction-Tuning-with-GPT-4 English Instruction-Following Data generated by GPT-4 using Alpaca prompts for fine-tuning LLMs. Usage and License Notices The data is intended and licensed for research use only. The dataset is CC BY NC 4.0 (allowing only… See the full description on the dataset page: https://huggingface.co/datasets/Fredithefish/Instruction-Tuning-with-GPT-4-RedPajama-Chat.textquestion-answering10K<n<100K6 likes58 downloads3y agoHugging Face12Soban1234 /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/Soban1234/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K3 likes49 downloads3mo agoHugging Face13ChipHolmes /Trendyol-Cybersecurity-Instruction-Tuning-Dataset-archive Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/Trendyol-Cybersecurity-Instruction-Tuning-Dataset-archive.texttext-generation10K<n<100K1 likes45 downloads2mo agoHugging Face14snorkelai /snorkel-curated-instruction-tuningPlease check out our Blog Post - How we built a better GenAI with programmatic data development for more details! Summary snorkel-curated-instruction-tuning is a curated dataset that consists of high-quality instruction-response pairs. These pairs were programmatically filtered with weak supervision from open-source datasets Databricks Dolly-15k, Open Assistant, and Helpful Instructions. To enhance the dataset, we also programmatically classified each instruction based on the… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/snorkel-curated-instruction-tuning.question-answering10K<n<100K11 likes44 downloads3y agoHugging Face15ahmadkaab /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/ahmadkaab/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K2 likes43 downloads9mo agoHugging Face16andycoco1128 /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/andycoco1128/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K1 likes43 downloads4mo agoHugging Face17ukcli /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/ukcli/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K0 likes38 downloads5mo agoHugging Face18invinciblejha01 /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/invinciblejha01/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K1 likes36 downloads5mo agoHugging Face19xTayyub /High-Quality-Synthetic-Python-Dataset-with-Reasoning-Traces-Chain-of-Thought-for-LLM-Fine-Tuning PyReason-7k: Advanced Python Chain-of-Thought Dataset Dataset Description This dataset contains 7,000+ high-quality Python programming examples designed for LLM fine-tuning. Each entry includes a detailed thought_process (Chain-of-Thought) to teach models logical reasoning before coding. Key Features: Chain-of-Thought: Step-by-step reasoning traces. Error Handling: Solutions include try-except blocks and logging. Diverse Tasks: Algorithms, API handling, Data Structures.… See the full description on the dataset page: https://huggingface.co/datasets/xTayyub/High-Quality-Synthetic-Python-Dataset-with-Reasoning-Traces-Chain-of-Thought-for-LLM-Fine-Tuning.text-generation1K<n<10K5 likes33 downloads10mo agoHugging Face20Madhu348 /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Madhu348/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K0 likes29 downloads7mo agoHugging Face21sanjaypantdsd /fine-tuning-socratic-dataset Fine-Tuning Concepts Dataset - Socratic Method A dataset of 100 conversation pairs teaching fine-tuning concepts through Socratic questioning. Dataset Summary Size: 100 conversations Format: Chat format (system, user, assistant) Method: Socratic questioning - guides learning through questions rather than direct answers Topics: Fine-tuning, PEFT methods (LoRA, QLoRA), data quality, troubleshooting Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/sanjaypantdsd/fine-tuning-socratic-dataset.textquestion-answeringn<1K0 likes27 downloads1y agoHugging Face22vector124 /my-recipe-chat-fine-tuning-data Dataset Card for my-recipe-chat-fine-tuning-data This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/vector124/my-recipe-chat-fine-tuning-data/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/vector124/my-recipe-chat-fine-tuning-data.texttext-generationn<1K0 likes17 downloads2y agoHugging Face23JEAPI /Trendyol-Cybersecurity-Instruction-Tuning-Dataset-EDIT Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/JEAPI/Trendyol-Cybersecurity-Instruction-Tuning-Dataset-EDIT.texttext-generation10K<n<100K0 likes16 downloads6mo agoHugging Face24invinciblejha01 /Trendyol-Cybersecurity-Instruction-Tuning-Dataset-EDIT Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/invinciblejha01/Trendyol-Cybersecurity-Instruction-Tuning-Dataset-EDIT.texttext-generation10K<n<100K0 likes16 downloads5mo agoHugging Face25alucent /mirror-Trendyol-Cybersecurity-Instruction-Tuning-Datasetgated Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K0 likes15 downloads2mo agoHugging Face26BhaskarAgrawal /Llama-2-fine-tuning Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/BhaskarAgrawal/Llama-2-fine-tuning.text-classification1K<n<10K0 likes14 downloads2y agoHugging Face27MatinaAI /instruction_tuning_datasetsgated 🧠 Persian Cultural Alignment Dataset for LLMs This repository contains a high-quality, instruction-following dataset for cultural alignment of large language models (LLMs) in the Persian language. The dataset is curated using hybrid strategies that incorporate culturally grounded generation, multi-turn dialogues, translation, and augmentation methods, making it suitable for SFT, DPO, RLHF, and alignment evaluation. 📚 Dataset Overview Domain Methods Used… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/instruction_tuning_datasets.textquestion-answering1 likes12 downloads1y agoHugging Face28SoftAge-AI /fine-tuning_datasetgated Fine-tuning Dataset Description This dataset contains 400 question-answer pairs for fine-tuning language models. Each pair consists of a query and an editor's answer, along with citations for the answer. Data attributes The dataset is in a CSV format with the following parameters: Query (str): The question. Editor's answer (str): The answer to the question. Citations (list of str): A list of citations for the answer. Data Source The data was… See the full description on the dataset page: https://huggingface.co/datasets/SoftAge-AI/fine-tuning_dataset.textquestion-answeringn<1K0 likes5 downloads3y agoHugging Face29riswanahamed /fine_tuning_demo_Datasetgatedtextsummarizationn<1K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.