CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01asanchez75 /tool_finetuning_dataset Tool Finetuning Dataset Dataset Description Dataset Summary This dataset is designed for fine-tuning language models to use tools (function calling) appropriately based on user queries. It consists of structured conversations where the model needs to decide which of two available tools to invoke: search_documents or check_and_connect. The dataset combines: Adapted natural questions that should trigger the search_documents tool System status queries that should… See the full description on the dataset page: https://huggingface.co/datasets/asanchez75/tool_finetuning_dataset.texttext-generation1K<n<10K1 likes400 downloads1y agoHugging Face02ameer4wisam /iraqi_words_finetuning Iraqi Words A manually compiled Iraqi Arabic dialect lexicon (930 terms, 50 categories) with a dependency-free BM25 retriever and a fine-tuning data generator built on top of it. Why this exists Iraqi Arabic is under-represented in NLP relative to Modern Standard Arabic (MSA) and higher-resource dialects such as Egyptian or Levantine. Lexical resources that map Iraqi terms to their MSA meanings — the kind needed to ground retrieval or instruction-tuning for… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi_words_finetuning.texttranslationn<1K0 likes111 downloads2mo agoHugging Face03scrapegraphai /scrapegraph-100k-finetuning ScrapeGraphAI 100k finetuning Dataset Summary A finetuning-ready derivative of ScrapeGraphAI-100k: schema-constrained web extraction examples where a model must produce JSON conforming to a user-defined JSON schema given Markdown-converted page content. Split Rows Targets train 25,244 GPT-5-nano regenerated targets test 2,808 GPT-5-nano regenerated targets human_eval 100 Human labeled extractions (evaluation only) Important: train/test… See the full description on the dataset page: https://huggingface.co/datasets/scrapegraphai/scrapegraph-100k-finetuning.textfeature-extraction10K<n<100K2 likes90 downloads2mo agoHugging Face04andresnowak /Instruction-finetuning-mixture-mnlp-with-nlp4educationDataset created using the Tulu3-sft-mixture and MNLP Question and golden answer dataset From the Tulue3-sft-mixture, messages that didn't have only 2 messages (user and assistant) where removed Also the datasets for alignment and jailbreaking were removed texttext-generation1M<n<10M0 likes86 downloads1y agoHugging Face05rriviere /oc-llm-finetuning-dataset Dataset médical bilingue, triage CHSA Dataset construit pour un POC d'agent IA de triage médical (mission OpenClassrooms, AI Engineer, CHSA). Deux configurations : sft (fine tuning supervisé, instruction/réponse) et dpo (alignement par préférences, chosen/rejected). Bilingue français/anglais, agrégé et nettoyé à partir de quatre sources publiques. Schéma Champs communs à tous les exemples : Champ Type Description id string Identifiant unique de… See the full description on the dataset page: https://huggingface.co/datasets/rriviere/oc-llm-finetuning-dataset.texttext-generation100K<n<1M0 likes70 downloads7d agoHugging Face06andresnowak /Instruction-finetuning-mixture-mnlp-only-english-with-nlp4educationDataset created using the Tulu3-sft-mixture and MNLP Question and golden answer dataset From the Tulue3-sft-mixture, messages that didn't have only 2 messages (user and assistant) where removed Also the datasets for alignment and jailbreaking were removed The dataset is only with english language texttext-generation1M<n<10M0 likes61 downloads1y agoHugging Face07dworsleytonks /medical-llm-finetuning-alignment-processed-datasettexttext-generation10K<n<100K0 likes58 downloads9mo agoHugging Face08liuyi2000 /deeprtl_finetuning_datasettexttext-generation100K<n<1M1 likes50 downloads9mo agoHugging Face09rescommons /Ecom-Chatbot-Finetuning-Dataset Ecom Chatbot Finetuning Dataset A unified instruction-following dataset for fine-tuning e-commerce customer service chatbots. It covers a wide range of real-world retail scenarios — from product discovery and order management to returns, complaints, and account support. Dataset Summary Field Value Total records 40,098 Language English Sources Amazon Reviews 2023, Amazon Meta 2023, ASOS, Bitext Response types Text, Tool Call, Mixed Difficulty levels 1… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Ecom-Chatbot-Finetuning-Dataset.tabularquestion-answering10K<n<100K0 likes45 downloads6mo agoHugging Face10open-paws /conversational-finetuning-llama-format Open Paws Conversational Finetuning Llama Format This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Training Data Format: CSV (Comma-separated values) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning Organization: Open Paws… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/conversational-finetuning-llama-format.texttext-generation10K<n<100K2 likes41 downloads1y agoHugging Face11AYI-NEDJIMI /llm-finetuning-fr LLM Fine-Tuning & Quantization - Dataset Francais Dataset bilingue complet sur le fine-tuning de LLM (LoRA, QLoRA, DPO, RLHF), la quantification de modeles (GPTQ, GGUF, AWQ), les modeles open source et le deploiement en production. Description Ce dataset couvre l'ensemble de la chaine de valeur des LLM open source, du fine-tuning au deploiement en production. Il est concu pour servir de reference aux developpeurs, ingenieurs ML, et equipes techniques souhaitant maitriser… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/llm-finetuning-fr.tabularquestion-answeringn<1K0 likes41 downloads7mo agoHugging Face12prakharb01 /Synthetic-Hinglish-Finetuning-Dataset Hinglish Conversations Dataset Overview This dataset contains synthetically generated conversational dialogues in Hinglish (a blend of Hindi and English). The conversations revolve around typical college life, cultural festivities, daily routines, and general discussions, designed to be relatable and engaging. Dataset Details Language: Hinglish (Hindi + English) Domain: College life, daily interactions, cultural events, and general discussions Size: 3576… See the full description on the dataset page: https://huggingface.co/datasets/prakharb01/Synthetic-Hinglish-Finetuning-Dataset.texttext-generation1K<n<10K0 likes35 downloads1y agoHugging Face13bala1524 /Medical-QA-Mistral7B-Finetuningtextquestion-answeringn<1K6 likes30 downloads3y agoHugging Face14sanjaypantdsd /fine-tuning-socratic-dataset Fine-Tuning Concepts Dataset - Socratic Method A dataset of 100 conversation pairs teaching fine-tuning concepts through Socratic questioning. Dataset Summary Size: 100 conversations Format: Chat format (system, user, assistant) Method: Socratic questioning - guides learning through questions rather than direct answers Topics: Fine-tuning, PEFT methods (LoRA, QLoRA), data quality, troubleshooting Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/sanjaypantdsd/fine-tuning-socratic-dataset.textquestion-answeringn<1K0 likes27 downloads1y agoHugging Face15science-of-finetuning /ultrachat_200k_generated_gemma-2-2b-itThis dataset contains 512 answers generated by the gemma-2-2b-it model on a subset of the ultrachat 200k test_sft dataset using greedy decoding. The subset was generated by filtering out conversations that were >= 1024 - 128 tokens long, and answers were cut off at each batch after 1024 - min(batch_prompt_lengths) generated tokens, such that each answer is at most 128 tokens long. The generated answers are 200k tokens so 390 tokens (~300 words or 2/3 pages) on average. texttext-generationn<1K0 likes25 downloads2y agoHugging Face16AYI-NEDJIMI /llm-finetuning-en LLM Fine-Tuning & Quantization - English Dataset Comprehensive bilingual dataset on LLM fine-tuning (LoRA, QLoRA, DPO, RLHF), model quantization (GPTQ, GGUF, AWQ), open source models, and production deployment. Description This dataset covers the entire open source LLM value chain, from fine-tuning to production deployment. It is designed as a reference for developers, ML engineers, and technical teams looking to master open source LLMs. Dataset Content… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/llm-finetuning-en.tabularquestion-answeringn<1K1 likes20 downloads7mo agoHugging Face17tdolega /rag-tge_finetuning-datasetDataset for finetuning LLM to generate responses with citations to source documents in RAG systems. Based on hotpot_qa. Generated for rag-tge project. Available also in Polish: rag-tge_finetuning-dataset_pl. texttext-generation1K<n<10K2 likes19 downloads2y agoHugging Face18Losa10 /G3P-Finetuning-examples 🧠 G3Pro-Finetuning-Examples A synthetic dataset designed for Instruction Fine-Tuning and Reasoning (CoT) development. Generated using the Gemini 3 Pro preview model, this dataset focuses on technical tasks, complex configurations, and logical step-by-step problem-solving. 📊 Dataset Summary Feature Details Version v1.4 License MIT License Languages Russian (ru), English (en) Size 3,898 records (~13 MB) Primary Task Instruction Following & Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Losa10/G3P-Finetuning-examples.texttext-generation1K<n<10K0 likes15 downloads7mo agoHugging Face19tdolega /rag-tge_finetuning-dataset_plTranslation of rag-tge_finetuning-dataset dataset using Google Translate. In addition several changes have been made: added negative examples repeated some questions with a changed number of source documents removed some questions that were badly translated removed titles from passages texttext-generation1K<n<10K2 likes13 downloads2y agoHugging Face20Sivanuja /Legal_vision_finetuning_data Sri Lankan Property Law Fine-Tuning Dataset Dataset Summary This dataset is a domain-specific legal instruction-tuning dataset designed for fine-tuning large language models for Sri Lankan property law reasoning and legal assistance. It focuses on core areas of Sri Lankan property law, including: Property transfer and conveyancing Title registration (Bim Saviya) Prescription and adverse possession Partition of co-owned property Mortgage and securities Lease and tenancy… See the full description on the dataset page: https://huggingface.co/datasets/Sivanuja/Legal_vision_finetuning_data.texttext-generation1K<n<10K0 likes13 downloads7mo agoHugging Face21yadz45 /finetuning_demotexttext-generationn<1K0 likes8 downloads2y agoHugging Face22beyarkay /5x-limited-parameter-finetuning 5x Model Organisms — Limited-Parameter Finetuning Pools Five per-category (user, assistant) datasets used to FT-elicit the misaligned behaviour of each model organism by unfreezing ~0.03 of the model's parameters and training for ~100 steps. The paired beyarkay/5x-{category}-mo and beyarkay/5x-{category}-control LoRA adapters are the starting points — see the collection. Files Each file is JSONL, one example per line, format: {{ "messages": [ {{"role": "system"… See the full description on the dataset page: https://huggingface.co/datasets/beyarkay/5x-limited-parameter-finetuning.texttext-generation1K<n<10K0 likes8 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.