CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01asanchez75 /tool_finetuning_dataset Tool Finetuning Dataset Dataset Description Dataset Summary This dataset is designed for fine-tuning language models to use tools (function calling) appropriately based on user queries. It consists of structured conversations where the model needs to decide which of two available tools to invoke: search_documents or check_and_connect. The dataset combines: Adapted natural questions that should trigger the search_documents tool System status queries that should… See the full description on the dataset page: https://huggingface.co/datasets/asanchez75/tool_finetuning_dataset.texttext-generation1K<n<10K1 likes400 downloads1y agoHugging Face02abhinav00anand /behavioral-fine-tuning-v1 Why This Dataset Exists "A model that refuses everything is useless. A model that refuses nothing is dangerous. The goal is a model that thinks." The Problem Our Solution Uncensored data → helpful but uncontrolled Surgical 85% helpfulness + 13% safety + 2% eval mix Safety-only data → lobotomized, over-refusing models Calibrated ratio preserves full helpfulness Raw data → PII, leaked secrets, duplicates 7-stage pipeline validates every… See the full description on the dataset page: https://huggingface.co/datasets/abhinav00anand/behavioral-fine-tuning-v1.imagetext-generation100K<n<1M1 likes381 downloads27d agoHugging Face03ameer4wisam /iraqi_words_finetuning Iraqi Words A manually compiled Iraqi Arabic dialect lexicon (930 terms, 50 categories) with a dependency-free BM25 retriever and a fine-tuning data generator built on top of it. Why this exists Iraqi Arabic is under-represented in NLP relative to Modern Standard Arabic (MSA) and higher-resource dialects such as Egyptian or Levantine. Lexical resources that map Iraqi terms to their MSA meanings — the kind needed to ground retrieval or instruction-tuning for… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi_words_finetuning.texttranslationn<1K0 likes111 downloads2mo agoHugging Face04scrapegraphai /scrapegraph-100k-finetuning ScrapeGraphAI 100k finetuning Dataset Summary A finetuning-ready derivative of ScrapeGraphAI-100k: schema-constrained web extraction examples where a model must produce JSON conforming to a user-defined JSON schema given Markdown-converted page content. Split Rows Targets train 25,244 GPT-5-nano regenerated targets test 2,808 GPT-5-nano regenerated targets human_eval 100 Human labeled extractions (evaluation only) Important: train/test… See the full description on the dataset page: https://huggingface.co/datasets/scrapegraphai/scrapegraph-100k-finetuning.textfeature-extraction10K<n<100K2 likes90 downloads2mo agoHugging Face05andresnowak /Instruction-finetuning-mixture-mnlp-with-nlp4educationDataset created using the Tulu3-sft-mixture and MNLP Question and golden answer dataset From the Tulue3-sft-mixture, messages that didn't have only 2 messages (user and assistant) where removed Also the datasets for alignment and jailbreaking were removed texttext-generation1M<n<10M0 likes86 downloads1y agoHugging Face06rriviere /oc-llm-finetuning-dataset Dataset médical bilingue, triage CHSA Dataset construit pour un POC d'agent IA de triage médical (mission OpenClassrooms, AI Engineer, CHSA). Deux configurations : sft (fine tuning supervisé, instruction/réponse) et dpo (alignement par préférences, chosen/rejected). Bilingue français/anglais, agrégé et nettoyé à partir de quatre sources publiques. Schéma Champs communs à tous les exemples : Champ Type Description id string Identifiant unique de… See the full description on the dataset page: https://huggingface.co/datasets/rriviere/oc-llm-finetuning-dataset.texttext-generation100K<n<1M0 likes70 downloads7d agoHugging Face07Mreeb /Dermatology-Question-Answer-Dataset-For-Fine-Tuning Dataset Details The data set has about 1 Million Tokens for Training and about 1500 question answers. Dataset Description This dataset is a comprehensive compilation of questions related to dermatology, spanning inquiries about various skin diseases, their symptoms, recommended medications, and available treatment modalities. Each question is paired with a concise and informative response, making it an ideal resource for training and fine-tuning language models in the… See the full description on the dataset page: https://huggingface.co/datasets/Mreeb/Dermatology-Question-Answer-Dataset-For-Fine-Tuning.tabulartext-generation1K<n<10K7 likes69 downloads3y agoHugging Face08MohamedAhmedAE /Med_LLaMa3_fine-tuning_dataset Med-LLaMA3 — Medical Instruction Fine-Tuning Dataset A large, unified medical instruction-tuning corpus (~1.65 million examples) compiled, cleaned, and standardized from a diverse set of public medical sources. It is the training corpus used to fine-tune the Med-LLaMA3 family (LLaMA-3.2 1B/3B and LLaMA-3.1 8B) in the paper “Med-LLaMA3: Advancing Medical Question-Answering Through Parameter-Efficient Fine-Tuning of Large Language Models” (Applied Sciences, 2026). All sources were… See the full description on the dataset page: https://huggingface.co/datasets/MohamedAhmedAE/Med_LLaMa3_fine-tuning_dataset.textquestion-answering1M<n<10M1 likes67 downloads3mo agoHugging Face09ysn-rfd /FibonacciAi-CODE-Fine_Tuning-DatasetDocument Version: 1.0.2 | Last Updated: 07/17/2026 texttext-generation1K<n<10K1 likes67 downloads2mo agoHugging Face10andresnowak /Instruction-finetuning-mixture-mnlp-only-english-with-nlp4educationDataset created using the Tulu3-sft-mixture and MNLP Question and golden answer dataset From the Tulue3-sft-mixture, messages that didn't have only 2 messages (user and assistant) where removed Also the datasets for alignment and jailbreaking were removed The dataset is only with english language texttext-generation1M<n<10M0 likes61 downloads1y agoHugging Face11dworsleytonks /medical-llm-finetuning-alignment-processed-datasettexttext-generation10K<n<100K0 likes58 downloads9mo agoHugging Face12liuyi2000 /deeprtl_finetuning_datasettexttext-generation100K<n<1M1 likes50 downloads9mo agoHugging Face13rescommons /Ecom-Chatbot-Finetuning-Dataset Ecom Chatbot Finetuning Dataset A unified instruction-following dataset for fine-tuning e-commerce customer service chatbots. It covers a wide range of real-world retail scenarios — from product discovery and order management to returns, complaints, and account support. Dataset Summary Field Value Total records 40,098 Language English Sources Amazon Reviews 2023, Amazon Meta 2023, ASOS, Bitext Response types Text, Tool Call, Mixed Difficulty levels 1… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Ecom-Chatbot-Finetuning-Dataset.tabularquestion-answering10K<n<100K0 likes45 downloads6mo agoHugging Face14open-paws /conversational-finetuning-llama-format Open Paws Conversational Finetuning Llama Format This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Training Data Format: CSV (Comma-separated values) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning Organization: Open Paws… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/conversational-finetuning-llama-format.texttext-generation10K<n<100K2 likes41 downloads1y agoHugging Face15AYI-NEDJIMI /llm-finetuning-fr LLM Fine-Tuning & Quantization - Dataset Francais Dataset bilingue complet sur le fine-tuning de LLM (LoRA, QLoRA, DPO, RLHF), la quantification de modeles (GPTQ, GGUF, AWQ), les modeles open source et le deploiement en production. Description Ce dataset couvre l'ensemble de la chaine de valeur des LLM open source, du fine-tuning au deploiement en production. Il est concu pour servir de reference aux developpeurs, ingenieurs ML, et equipes techniques souhaitant maitriser… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/llm-finetuning-fr.tabularquestion-answeringn<1K0 likes41 downloads7mo agoHugging Face16prakharb01 /Synthetic-Hinglish-Finetuning-Dataset Hinglish Conversations Dataset Overview This dataset contains synthetically generated conversational dialogues in Hinglish (a blend of Hindi and English). The conversations revolve around typical college life, cultural festivities, daily routines, and general discussions, designed to be relatable and engaging. Dataset Details Language: Hinglish (Hindi + English) Domain: College life, daily interactions, cultural events, and general discussions Size: 3576… See the full description on the dataset page: https://huggingface.co/datasets/prakharb01/Synthetic-Hinglish-Finetuning-Dataset.texttext-generation1K<n<10K0 likes35 downloads1y agoHugging Face17bala1524 /Medical-QA-Mistral7B-Finetuningtextquestion-answeringn<1K6 likes30 downloads3y agoHugging Face18sanjaypantdsd /fine-tuning-socratic-dataset Fine-Tuning Concepts Dataset - Socratic Method A dataset of 100 conversation pairs teaching fine-tuning concepts through Socratic questioning. Dataset Summary Size: 100 conversations Format: Chat format (system, user, assistant) Method: Socratic questioning - guides learning through questions rather than direct answers Topics: Fine-tuning, PEFT methods (LoRA, QLoRA), data quality, troubleshooting Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/sanjaypantdsd/fine-tuning-socratic-dataset.textquestion-answeringn<1K0 likes27 downloads1y agoHugging Face19science-of-finetuning /ultrachat_200k_generated_gemma-2-2b-itThis dataset contains 512 answers generated by the gemma-2-2b-it model on a subset of the ultrachat 200k test_sft dataset using greedy decoding. The subset was generated by filtering out conversations that were >= 1024 - 128 tokens long, and answers were cut off at each batch after 1024 - min(batch_prompt_lengths) generated tokens, such that each answer is at most 128 tokens long. The generated answers are 200k tokens so 390 tokens (~300 words or 2/3 pages) on average. texttext-generationn<1K0 likes25 downloads2y agoHugging Face20AYI-NEDJIMI /llm-finetuning-en LLM Fine-Tuning & Quantization - English Dataset Comprehensive bilingual dataset on LLM fine-tuning (LoRA, QLoRA, DPO, RLHF), model quantization (GPTQ, GGUF, AWQ), open source models, and production deployment. Description This dataset covers the entire open source LLM value chain, from fine-tuning to production deployment. It is designed as a reference for developers, ML engineers, and technical teams looking to master open source LLMs. Dataset Content… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/llm-finetuning-en.tabularquestion-answeringn<1K1 likes20 downloads7mo agoHugging Face21tdolega /rag-tge_finetuning-datasetDataset for finetuning LLM to generate responses with citations to source documents in RAG systems. Based on hotpot_qa. Generated for rag-tge project. Available also in Polish: rag-tge_finetuning-dataset_pl. texttext-generation1K<n<10K2 likes19 downloads2y agoHugging Face22vector124 /my-recipe-chat-fine-tuning-data Dataset Card for my-recipe-chat-fine-tuning-data This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/vector124/my-recipe-chat-fine-tuning-data/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/vector124/my-recipe-chat-fine-tuning-data.texttext-generationn<1K0 likes17 downloads2y agoHugging Face23viplav0009 /friends-dataset-for-chandler-bing-fine-tuning Chandler Bing Sarcasm Dataset This dataset contains conversational turns designed to train a model in the persona of Chandler Bing. It focuses on his signature sarcasm, self-deprecation, and awkward humor. Dataset Structure The data follows the ShareGPT format, making it compatible with tools like Unsloth for fast fine-tuning. Data Fields conversations: A list of messages in a single chat session. from: The speaker identity (human for user, gpt for… See the full description on the dataset page: https://huggingface.co/datasets/viplav0009/friends-dataset-for-chandler-bing-fine-tuning.texttext-generation1K<n<10K0 likes17 downloads8mo agoHugging Face24Losa10 /G3P-Finetuning-examples 🧠 G3Pro-Finetuning-Examples A synthetic dataset designed for Instruction Fine-Tuning and Reasoning (CoT) development. Generated using the Gemini 3 Pro preview model, this dataset focuses on technical tasks, complex configurations, and logical step-by-step problem-solving. 📊 Dataset Summary Feature Details Version v1.4 License MIT License Languages Russian (ru), English (en) Size 3,898 records (~13 MB) Primary Task Instruction Following & Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Losa10/G3P-Finetuning-examples.texttext-generation1K<n<10K0 likes15 downloads7mo agoHugging Face25tdolega /rag-tge_finetuning-dataset_plTranslation of rag-tge_finetuning-dataset dataset using Google Translate. In addition several changes have been made: added negative examples repeated some questions with a changed number of source documents removed some questions that were badly translated removed titles from passages texttext-generation1K<n<10K2 likes13 downloads2y agoHugging Face26Sivanuja /Legal_vision_finetuning_data Sri Lankan Property Law Fine-Tuning Dataset Dataset Summary This dataset is a domain-specific legal instruction-tuning dataset designed for fine-tuning large language models for Sri Lankan property law reasoning and legal assistance. It focuses on core areas of Sri Lankan property law, including: Property transfer and conveyancing Title registration (Bim Saviya) Prescription and adverse possession Partition of co-owned property Mortgage and securities Lease and tenancy… See the full description on the dataset page: https://huggingface.co/datasets/Sivanuja/Legal_vision_finetuning_data.texttext-generation1K<n<10K0 likes13 downloads7mo agoHugging Face27yadz45 /finetuning_demotexttext-generationn<1K0 likes8 downloads2y agoHugging Face28beyarkay /5x-limited-parameter-finetuning 5x Model Organisms — Limited-Parameter Finetuning Pools Five per-category (user, assistant) datasets used to FT-elicit the misaligned behaviour of each model organism by unfreezing ~0.03 of the model's parameters and training for ~100 steps. The paired beyarkay/5x-{category}-mo and beyarkay/5x-{category}-control LoRA adapters are the starting points — see the collection. Files Each file is JSONL, one example per line, format: {{ "messages": [ {{"role": "system"… See the full description on the dataset page: https://huggingface.co/datasets/beyarkay/5x-limited-parameter-finetuning.texttext-generation1K<n<10K0 likes8 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.