CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceH4 /instruction-datasetThis is the blind eval dataset of high-quality, diverse, human-written instructions with demonstrations. We will be using this for step 3 evaluations in our RLHF pipeline. textn<1K66 likes7.5k downloads4y agoHugging Face02Trendyol /Trendyol-Cybersecurity-Instruction-Tuning-Dataset Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0) 🚀 TL;DR 53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K134 likes5.7k downloads1y agoHugging Face03dkoterwa /camel_ai_chemistry_instruction_datasettext10K<n<100K2 likes802 downloads2y agoHugging Face04Tiamz /cybersecurity-instruction-datasettext10K<n<100K0 likes545 downloads1y agoHugging Face05BlossomsAI /merged_vietnamese_instruction_datasettext1M<n<10M0 likes386 downloads1y agoHugging Face06Parssky /industrial-instruction-dataset Industrial-Instruction Dataset Industrial-Instruction provides benchmark and training-ready QA instances derived from industrial technical reports, designed to evaluate robustness under realistic retrieval conditions. Samples are grounded in retrieved evidence and include irrelevant retrieval, single-/multi-document support, and single-/multi-document answer settings. Paper Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and… See the full description on the dataset page: https://huggingface.co/datasets/Parssky/industrial-instruction-dataset.tabularquestion-answering10K<n<100K0 likes378 downloads1mo agoHugging Face07proxectonos /cpt_instruction_datasets Instruction datasets Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3. Dataset creation Datasets were created using two different techniques: Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.tabulartext-generation100K<n<1M0 likes342 downloads5mo agoHugging Face08ldbb123 /Instruction-tuning_Datasetstext1M<n<10M0 likes262 downloads2y agoHugging Face090xrushi /git-instruction-dataset Git Command Dataset This dataset contains 9008 examples of git commands paired with natural language instructions. Each example includes: instruction: A natural language description of what git operation to perform output: The corresponding git command(s) to execute Usage from datasets import load_dataset dataset = load_dataset("{repo_id}") print(dataset['train'][0]) Dataset Structure {{ "instruction": "Create a new git repository in the… See the full description on the dataset page: https://huggingface.co/datasets/0xrushi/git-instruction-dataset.text1K<n<10K2 likes256 downloads1y agoHugging Face10paodigitalhub /pao-instruction-qa-conversation-datasetPa'O Instruction, QA & Conversation Dataset An open and community-driven dataset for the Pa'O ("blk") language, developed through the RYPAK Ecosystem, SuccessImprove (SI), and Pa'O Digital Hub. The dataset is designed to support natural language processing (NLP), large language models (LLMs), conversational dialogue, instruction following, language technology research, and digital preservation of the Pa'O language. The project focuses on building a free, open, reusable, and continuously… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-instruction-qa-conversation-dataset.textquestion-answeringn<1K1 likes247 downloads3d agoHugging Face11tuandunghcmut /Trendyol-Cybersecurity-Instruction-Tuning-Datasetgated Trendyol Cybersecurity Instruction Tuning Dataset (GPT Format) A conversational dataset in GPT/OpenAI messages format, converted from Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset. Designed for training language models in advanced cyber-defense and security principles. Dataset Description This dataset contains 53,201 high-quality instruction-tuning examples focused on cybersecurity, converted to the standard GPT conversation format (messages) for… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.texttext-generation10K<n<100K2 likes216 downloads1y agoHugging Face12kamruzzaman-asif /bangla-instruction-dataset 🧠 Bangla Instruction Dataset This dataset repository consolidates high-quality instruction-tuning data from multiple popular sources, structured for easy use in training and evaluating instruction-following models. 📚 Dataset Splits The dataset is organized into the following splits: Split Name Source Dataset Description OdiaGenAI OdiaGenAI/all_combined_bengali_252k A large-scale collection of diverse Bangla instructions and responses. chrononeel… See the full description on the dataset page: https://huggingface.co/datasets/kamruzzaman-asif/bangla-instruction-dataset.texttext-generation1M<n<10M1 likes210 downloads1y agoHugging Face13FinLang /investopedia-instruction-tuning-dataset Dataset Card for investopedia-instruction-tuning dataset We curate a dataset of substantial size pertaining to finance from Investopedia using a new technique that leverages unstructured scraping data and LLM to generate structured data that is suitable for fine-tuning embedding models. The dataset generation uses a new method of self-verification that ensures that the generated question-answer pairs and not hallucinated by the LLM with high probability. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/FinLang/investopedia-instruction-tuning-dataset.text100K<n<1M23 likes201 downloads2y agoHugging Face14henryen /origen_dataset_instruction OriGen: Enhancing RTL Code Generation with Code-to-Code Augmentation and Self-Reflection Introduction OriGen is a fine-tuned lora model designed for Verilog code generation. It is trained on top of DeepSeek Coder 7B using datasets generated from code-to-code augmentation and self-reflection. The datasets can be found in the origen_dataset_instruction. OriGen_Fix is a fine-tuned lora model designed for fixing syntax errors in Verilog code. It is trained based on OriGen… See the full description on the dataset page: https://huggingface.co/datasets/henryen/origen_dataset_instruction.text100K<n<1M4 likes189 downloads2y agoHugging Face15CarrotAI /ko-instruction-dataset 고품질 한국어 데이터셋 한국어로 이루어진 고품질 한국어 데이터셋 입니다. WizardLM-2-8x22B 모델을 사용하여 WizardLM: Empowering Large Language Models to Follow Complex Instructions에서 소개된 방법으로 생성되었습니다. @article{koinstructiondatasetcard, title={CarrotAI/ko-instruction-dataset Card}, author={CarrotAI (L, GEUN)}, year={2024}, url = {https://huggingface.co/datasets/CarrotAI/ko-instruction-dataset} } texttext-generation1K<n<10K27 likes174 downloads2y agoHugging Face16mdonigian /full-structured-instruction-sft-dataset Full Structured + Instruction SFT Corpus Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data. Dataset repo mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11 Included files train_full_sft.jsonl: full merged and shuffled SFT dataset source_glaive.jsonl: processed Glaive subset source_hermes.jsonl: processed Hermes subset source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.texttext-generation10K<n<100K0 likes170 downloads7mo agoHugging Face17KrynexLabs /KrynexAI-Dataset-Flash-Instruction 🧠 KrynexAI Dataset English | Русский 📌 Overview KrynexAI Dataset is a high-quality, synthetically expanded collection of 10,000+ instruction-response pairs designed for fine-tuning Large Language Models (LLMs). The dataset covers a wide range of topics including: 💻 Programming (Python, algorithms, data structures) 🤖 AI & Machine Learning (neural networks, transformers, LLMs) 🔭 Science (physics, cosmology, biology, neuroscience) 🧠 Philosophy & Psychology… See the full description on the dataset page: https://huggingface.co/datasets/KrynexLabs/KrynexAI-Dataset-Flash-Instruction.texttext-generation10K<n<100K1 likes170 downloads13h agoHugging Face18EthioNLP /Amharic_Instruction_dataset SFT-Data for Walia-LLM: Enhancing Amharic-LLaMA by Integrating Task-Specific and Generative Datasets Dataset Summary The Walia dataset is designed to enhance large language models for the Amharic language by: Converting existing task-specific datasets (e.g., sentiment analysis, QA, NER) into instruction format. Creating new generative datasets (e.g., poem generation, religious lyrics, story generation). Translating English instruction datasets (e.g., Alpaca, Dolly) into… See the full description on the dataset page: https://huggingface.co/datasets/EthioNLP/Amharic_Instruction_dataset.text100K<n<1M5 likes132 downloads1y agoHugging Face19Roshan32 /Hinglish_Dataset_instruction_and_rawtexttext-generation10K<n<100K1 likes129 downloads9mo agoHugging Face20renhuimin /RL-Instruction-Following-Dataset RL-Instruction-Following-Dataset 🎯 A Verifiable, Rule-Based Dataset for Reinforcement Learning with Verifiable Rewards (RLVR) 📖 Dataset Card 🚀 Usage ⚖️ License Overview This dataset is designed to enhance the Instruction Following capabilities of Large Language Models (LLMs) through Reinforcement Learning (RL). Unlike subjective preference datasets (e.g., standard RLHF), this dataset focuses on Objective, Rule-Based Constraints. Each entry provides a prompt with… See the full description on the dataset page: https://huggingface.co/datasets/renhuimin/RL-Instruction-Following-Dataset.textreinforcement-learning100K<n<1M4 likes128 downloads9mo agoHugging Face21BlossomsAI /reduced_vietnamese_instruction_datasettext1M<n<10M0 likes124 downloads2y agoHugging Face22dkoterwa /camel_ai_physics_instruction_datasettext10K<n<100K1 likes118 downloads2y agoHugging Face23ChavyvAkvar /Trendyol-Cybersecurity-Instruction-Tuning-Dataset-Convertedtext10K<n<100K1 likes108 downloads1y agoHugging Face24ashokpoudel /English-Nepali-Translation-Instruction-Dataset Dataset Card: Instruction-Based English-Nepali Translation Dataset Dataset Description This dataset consists of English-Nepali parallel sentences converted into an instruction-based format. Each entry prompts the model to translate a given sentence from English to Nepali or vice versa. Source Data Original Dataset: English-Nepali Parallel SentencesPaper: NepBERTa: Nepali Language Model Trained in a Large CorpusAuthors: Milan Gautam, Sulav Timilsina, Binod… See the full description on the dataset page: https://huggingface.co/datasets/ashokpoudel/English-Nepali-Translation-Instruction-Dataset.text1M<n<10M2 likes100 downloads3y agoHugging Face25bernabepuente /devops-cloud-instruction-dataset DevOps & Cloud Infrastructure Dataset Professional instruction-response pairs for DevOps engineers covering Kubernetes, Docker, Terraform, CI/CD, and cloud services (AWS, Azure). Dataset Details Dataset Description This is a high-quality instruction-tuning dataset focused on Devops Cloud topics. Each entry includes: A clear instruction/question Optional input context A detailed response/solution Chain-of-thought reasoning process Curated by: CloudKernel.IO… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/devops-cloud-instruction-dataset.texttext-generationn<1K0 likes99 downloads5mo agoHugging Face26dkoterwa /camel_ai_biology_instruction_datasettext10K<n<100K0 likes92 downloads2y agoHugging Face27khairi /life2lang-instruction-datasettext1M<n<10M0 likes91 downloads2mo agoHugging Face28NinaCalvi /ultra-50k-samples-dataset-instruction_followingtabular10K<n<100K0 likes90 downloads2y agoHugging Face29GreatCaptainNemo /instruction_dataset ProLLaMA Instruction Dataset This repository contains the instruction dataset for ProLLaMA. Paper ProLLaMA: A Protein Large Language Model for Multi-Task Protein Language Processing Code GitHub Repository Introduction Recent advances in Protein Language Models (PLMs) have transformed protein engineering, yet unlike their counterparts in Natural Language Processing (NLP), current PLMs exhibit a fundamental limitation: they excel in either Protein… See the full description on the dataset page: https://huggingface.co/datasets/GreatCaptainNemo/instruction_dataset.texttext-generation10M<n<100M7 likes89 downloads1y agoHugging Face30landedmover /uplimit-instruction-tuning-dataset Dataset Card for uplimit-instruction-tuning-dataset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/landedmover/uplimit-instruction-tuning-dataset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/landedmover/uplimit-instruction-tuning-dataset.textn<1K0 likes89 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.