CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Trendyol /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K134 likes5.7k downloads1y agoHugging Face02Parssky /industrial-instruction-dataset Industrial-Instruction Dataset Industrial-Instruction provides benchmark and training-ready QA instances derived from industrial technical reports, designed to evaluate robustness under realistic retrieval conditions. Samples are grounded in retrieved evidence and include irrelevant retrieval, single-/multi-document support, and single-/multi-document answer settings. Paper Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and… See the full description on the dataset page: https://huggingface.co/datasets/Parssky/industrial-instruction-dataset.tabularquestion-answering10K<n<100K0 likes378 downloads1mo agoHugging Face03MehdiFe /csharp-instruction-Dataset 🧠 CodeGen C# Dataset A curated dataset for training and evaluating code generation models in the C# programming language. It combines high-quality open-source code with enterprise-grade internal code examples, carefully selected and preprocessed to support research on structured prompting and high-fidelity code generation. 📦 Dataset Summary This dataset is designed to support instruction-tuned and general-purpose code generation models, with a particular emphasis on… See the full description on the dataset page: https://huggingface.co/datasets/MehdiFe/csharp-instruction-Dataset.text-generation1K<n<10K3 likes369 downloads1y agoHugging Face04proxectonos /cpt_instruction_datasets Instruction datasets Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3. Dataset creation Datasets were created using two different techniques: Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.tabulartext-generation100K<n<1M0 likes342 downloads5mo agoHugging Face05paodigitalhub /pao-instruction-qa-conversation-datasetPa'O Instruction, QA & Conversation Dataset An open and community-driven dataset for the Pa'O ("blk") language, developed through the RYPAK Ecosystem, SuccessImprove (SI), and Pa'O Digital Hub. The dataset is designed to support natural language processing (NLP), large language models (LLMs), conversational dialogue, instruction following, language technology research, and digital preservation of the Pa'O language. The project focuses on building a free, open, reusable, and continuously… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-instruction-qa-conversation-dataset.textquestion-answeringn<1K1 likes247 downloads3d agoHugging Face06tuandunghcmut /Trendyol-Cybersecurity-Instruction-Tuning-Datasetgated Trendyol Cybersecurity Instruction Tuning Dataset (GPT Format) A conversational dataset in GPT/OpenAI messages format, converted from Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset. Designed for training language models in advanced cyber-defense and security principles. Dataset Description This dataset contains 53,201 high-quality instruction-tuning examples focused on cybersecurity, converted to the standard GPT conversation format (messages) for… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K2 likes216 downloads1y agoHugging Face07kamruzzaman-asif /bangla-instruction-dataset 🧠 Bangla Instruction Dataset This dataset repository consolidates high-quality instruction-tuning data from multiple popular sources, structured for easy use in training and evaluating instruction-following models. 📚 Dataset Splits The dataset is organized into the following splits: Split Name Source Dataset Description OdiaGenAI OdiaGenAI/all_combined_bengali_252k A large-scale collection of diverse Bangla instructions and responses. chrononeel… See the full description on the dataset page: https://huggingface.co/datasets/kamruzzaman-asif/bangla-instruction-dataset.texttext-generation1M<n<10M1 likes210 downloads1y agoHugging Face08CarrotAI /ko-instruction-dataset 고품질 한국어 데이터셋 한국어로 이루어진 고품질 한국어 데이터셋 입니다. WizardLM-2-8x22B 모델을 사용하여 WizardLM: Empowering Large Language Models to Follow Complex Instructions에서 소개된 방법으로 생성되었습니다. @article{koinstructiondatasetcard, title={CarrotAI/ko-instruction-dataset Card}, author={CarrotAI (L, GEUN)}, year={2024}, url = {https://huggingface.co/datasets/CarrotAI/ko-instruction-dataset} } texttext-generation1K<n<10K27 likes174 downloads2y agoHugging Face09mdonigian /full-structured-instruction-sft-dataset Full Structured + Instruction SFT Corpus Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data. Dataset repo mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11 Included files train_full_sft.jsonl: full merged and shuffled SFT dataset source_glaive.jsonl: processed Glaive subset source_hermes.jsonl: processed Hermes subset source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.texttext-generation10K<n<100K0 likes170 downloads7mo agoHugging Face10KrynexLabs /KrynexAI-Dataset-Flash-Instruction 🧠 KrynexAI Dataset English | Русский 📌 Overview KrynexAI Dataset is a high-quality, synthetically expanded collection of 10,000+ instruction-response pairs designed for fine-tuning Large Language Models (LLMs). The dataset covers a wide range of topics including: 💻 Programming (Python, algorithms, data structures) 🤖 AI & Machine Learning (neural networks, transformers, LLMs) 🔭 Science (physics, cosmology, biology, neuroscience) 🧠 Philosophy & Psychology… See the full description on the dataset page: https://huggingface.co/datasets/KrynexLabs/KrynexAI-Dataset-Flash-Instruction.texttext-generation10K<n<100K1 likes170 downloads14h agoHugging Face11Roshan32 /Hinglish_Dataset_instruction_and_rawtexttext-generation10K<n<100K1 likes129 downloads9mo agoHugging Face12renhuimin /RL-Instruction-Following-Dataset RL-Instruction-Following-Dataset 🎯 A Verifiable, Rule-Based Dataset for Reinforcement Learning with Verifiable Rewards (RLVR) 📖 Dataset Card 🚀 Usage ⚖️ License Overview This dataset is designed to enhance the Instruction Following capabilities of Large Language Models (LLMs) through Reinforcement Learning (RL). Unlike subjective preference datasets (e.g., standard RLHF), this dataset focuses on Objective, Rule-Based Constraints. Each entry provides a prompt with… See the full description on the dataset page: https://huggingface.co/datasets/renhuimin/RL-Instruction-Following-Dataset.textreinforcement-learning100K<n<1M4 likes128 downloads9mo agoHugging Face13bernabepuente /devops-cloud-instruction-dataset DevOps & Cloud Infrastructure Dataset Professional instruction-response pairs for DevOps engineers covering Kubernetes, Docker, Terraform, CI/CD, and cloud services (AWS, Azure). Dataset Details Dataset Description This is a high-quality instruction-tuning dataset focused on Devops Cloud topics. Each entry includes: A clear instruction/question Optional input context A detailed response/solution Chain-of-thought reasoning process Curated by: CloudKernel.IO… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/devops-cloud-instruction-dataset.texttext-generationn<1K0 likes99 downloads5mo agoHugging Face14GreatCaptainNemo /instruction_dataset ProLLaMA Instruction Dataset This repository contains the instruction dataset for ProLLaMA. Paper ProLLaMA: A Protein Large Language Model for Multi-Task Protein Language Processing Code GitHub Repository Introduction Recent advances in Protein Language Models (PLMs) have transformed protein engineering, yet unlike their counterparts in Natural Language Processing (NLP), current PLMs exhibit a fundamental limitation: they excel in either Protein… See the full description on the dataset page: https://huggingface.co/datasets/GreatCaptainNemo/instruction_dataset.texttext-generation10M<n<100M7 likes89 downloads1y agoHugging Face1510xtechnologieS /telugu_instruction_dataset Telugu Instruction Dataset — Luuka AI Built by 10x Technologies — a curated Telugu-language instruction-tuning dataset for training the Luuka voice assistant. Overview This dataset contains 7,496 instruction–response pairs and 399 multi-turn conversations in Telugu, covering a broad range of natural voice assistant interactions. Every response is written entirely in Telugu script — no English characters appear in any response field. Subset Config Name Pairs… See the full description on the dataset page: https://huggingface.co/datasets/10xtechnologieS/telugu_instruction_dataset.texttext-generation1K<n<10K0 likes85 downloads6mo agoHugging Face16bediss /forum-instruction-tuning-dataset Looksmaxxing Forum Dataset A curated instruction-tuning dataset derived from a large looksmaxxing and aesthetic self-improvement forum, containing high-density community knowledge on skincare, nutrition, supplementation, and appearance optimization. Dataset Summary This dataset was produced by scraping, parsing, cleaning, and LLM-filtering over 1.28 million raw forum posts down to 77,417 high-quality instruction-response pairs using a multi-stage pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/bediss/forum-instruction-tuning-dataset.texttext-generation10K<n<100K1 likes79 downloads3mo agoHugging Face17bernabepuente /backend-api-instruction-dataset Backend & API Development Dataset Instruction dataset focused on RESTful API design, WebSocket real-time communication, microservices patterns, and API best practices. Dataset Details Dataset Description This is a high-quality instruction-tuning dataset focused on Backend Api topics. Each entry includes: A clear instruction/question Optional input context A detailed response/solution Chain-of-thought reasoning process Curated by: CloudKernel.IO… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/backend-api-instruction-dataset.texttext-generationn<1K0 likes71 downloads5mo agoHugging Face18Aman0026 /ArogyaAI-Medical-Instruction-Dataset ArogyaAI Multimodal Medical Dataset 🏥 Dataset Description The ArogyaAI Medical Dataset is a comprehensive, multimodal instruction-tuning and retrieval-augmented generation (RAG) dataset tailored specifically for the Indian healthcare context. It was built to fine-tune the ArogyaAI LLaMA 3 8B model and BioLORD/BioBERT embeddings, bridging the gap between modern Allopathic medicine and traditional AYUSH systems (Ayurveda and Homeopathy). Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Aman0026/ArogyaAI-Medical-Instruction-Dataset.textquestion-answering10K<n<100K0 likes70 downloads3mo agoHugging Face19abhilash88 /odia-instruction-dataset Odia Instruction Following Dataset Dataset Description This is a comprehensive Odia language instruction-following dataset designed for training conversational AI models, chatbots, and instruction-following systems in Odia (ଓଡ଼ିଆ). The dataset contains high-quality instruction-response pairs that enable models to understand and follow instructions in the Odia language. Dataset Summary Language: Odia (ଓଡ଼ିଆ) Total Records: 324,560 Format:… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/odia-instruction-dataset.tabulartext-generation100K<n<1M1 likes69 downloads1y agoHugging Face20besmabakirci01 /turkish-poetry-instruction-dataset Turkish Poetry Instruction Dataset Türkçe şiir üretimi için Alpaca instruction formatında hazırlanmış fine-tuning dataseti. Unsloth + Qwen LoRA eğitimi (3. ödev) için tasarlanmıştır. Format Her satır: { "instruction": "Kullanıcının şiir isteği", "input": "", "output": "Üretilmesi beklenen şiir metni" } Dosyalar Dosya Açıklama train.jsonl Eğitim (%90) validation.jsonl Doğrulama (%5) test.jsonl Test (%5) Toplam ~4.961 örnek;… See the full description on the dataset page: https://huggingface.co/datasets/besmabakirci01/turkish-poetry-instruction-dataset.texttext-generation1K<n<10K0 likes69 downloads2mo agoHugging Face21anony-mouse123 /Instruction_recall_dataset CanaryBench-PII Frequency-aware canary injection benchmark for auditing memorization in finetuned language models, built on the AI4Privacy PII reconstruction task. Dataset Description This dataset is part of CanaryBench, a benchmark for evaluating memorization in finetuned language models across repetition tiers and privacy regimes. Frequency tiers: 1×, 10×, 50× PII types: EMAIL, PHONE Member canaries: 770 Reference canaries: 1000 Tasks: PII detection, secret… See the full description on the dataset page: https://huggingface.co/datasets/anony-mouse123/Instruction_recall_dataset.texttext-generation10K<n<100K0 likes59 downloads2mo agoHugging Face22pythainlp /han-instruction-dataset Han Instruction Dataset Han instruction dataset: Thai instruction dataset 🪿 Han (ห่าน or goose) Instruction Dataset is a Thai instruction dataset by PyThaiNLP. This dataset collects all Thai instruct datasets that were made by humans and our old model. The dataset can be used to train Instruction Following models like ChatGPT or others. The final dataset of han instruction dataset was released! GitHub: https://github.com/wannaphong/han-instruction-dataset Data sources:… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruction-dataset.texttext-generation10K<n<100K1 likes59 downloads5d agoHugging Face23bernabepuente /ai-ml-instruction-dataset AI/ML Engineering Instruction Dataset Comprehensive instruction dataset covering machine learning concepts, PyTorch implementations, NLP with transformers, model evaluation, and feature engineering. Dataset Details Dataset Description This is a high-quality instruction-tuning dataset focused on Ai Ml topics. Each entry includes: A clear instruction/question Optional input context A detailed response/solution Chain-of-thought reasoning process Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/ai-ml-instruction-dataset.texttext-generationn<1K0 likes58 downloads5mo agoHugging Face24opennyaiorg /aalap_instruction_datasetgated Aalap Instruction dataset This dataset aims to build an AI Assistant for Legal and Paralegal functions in India (Aalap). The main objective behind creating Aalap was to train a small-sized LLM specializing in specific legal tasks focusing on legal reasoning. Hence, prioritizing which legal tasks to focus on was an important step. We discussed with Legal practitioners which legal tasks they would want to be automated. Then, we selected the most common legal tasks for publicly… See the full description on the dataset page: https://huggingface.co/datasets/opennyaiorg/aalap_instruction_dataset.texttext-generation10K<n<100K4 likes48 downloads3y agoHugging Face25Sushilmishrawork /india-state-debt-instruction-dataset 🇮🇳 India State Debt Chain-of-Thought (CoT) Instruction Dataset (2020–2026) This dataset contains 761 high-quality instruction-following prompt-response pairs focusing on public debt liabilities, Debt-to-GSDP percentages, per-capita debt, fiscal vulnerability classifications, and multi-year growth trajectories across 31 Indian States and Union Territories (2020–2026). Data Schema & Format Every record includes structured Chain-of-Thought (CoT) reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Sushilmishrawork/india-state-debt-instruction-dataset.question-answeringn<1K0 likes47 downloads25d agoHugging Face26bernabepuente /python-instruction-dataset Python Developer Instruction Dataset High-quality instruction-response pairs covering Python development best practices, async programming, decorators, type hints, and data manipulation with Pandas. Dataset Details Dataset Description This is a high-quality instruction-tuning dataset focused on Python topics. Each entry includes: A clear instruction/question Optional input context A detailed response/solution Chain-of-thought reasoning process Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/python-instruction-dataset.texttext-generationn<1K0 likes45 downloads5mo agoHugging Face27Soban1234 /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/Soban1234/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K3 likes45 downloads3mo agoHugging Face28ahmadkaab /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/ahmadkaab/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K2 likes42 downloads9mo agoHugging Face29andycoco1128 /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/andycoco1128/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K1 likes41 downloads4mo agoHugging Face30ChipHolmes /Trendyol-Cybersecurity-Instruction-Tuning-Dataset-archive Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/Trendyol-Cybersecurity-Instruction-Tuning-Dataset-archive.texttext-generation10K<n<100K1 likes40 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.