datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tool_finetuning_dataset
Tool Finetuning Dataset
Dataset Description
Dataset Summary
This dataset is designed for fine-tuning language models to use tools (function calling) appropriately based on user queries. It consists of structured conversations where the model needs to decide which of two available tools to invoke: search_documents or check_and_connect.
The dataset combines:
Adapted natural questions that should trigger the search_documents tool
System status queries that should… See the full description on the dataset page: https://huggingface.co/datasets/asanchez75/tool_finetuning_dataset.behavioral-fine-tuning-v1
Why This Dataset Exists
"A model that refuses everything is useless. A model that refuses nothing is dangerous. The goal is a model that thinks."
The Problem
Our Solution
Uncensored data → helpful but uncontrolled
Surgical 85% helpfulness + 13% safety + 2% eval mix
Safety-only data → lobotomized, over-refusing models
Calibrated ratio preserves full helpfulness
Raw data → PII, leaked secrets, duplicates
7-stage pipeline validates every… See the full description on the dataset page: https://huggingface.co/datasets/abhinav00anand/behavioral-fine-tuning-v1.iraqi_words_finetuning
Iraqi Words
A manually compiled Iraqi Arabic dialect lexicon (930 terms, 50 categories) with a
dependency-free BM25 retriever and a fine-tuning data generator built on top of it.
Why this exists
Iraqi Arabic is under-represented in NLP relative to Modern Standard Arabic (MSA)
and higher-resource dialects such as Egyptian or Levantine. Lexical resources that
map Iraqi terms to their MSA meanings — the kind needed to ground retrieval or
instruction-tuning for… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi_words_finetuning.scrapegraph-100k-finetuning
ScrapeGraphAI 100k finetuning
Dataset Summary
A finetuning-ready derivative of ScrapeGraphAI-100k: schema-constrained web extraction examples where a model must produce JSON conforming to a user-defined JSON schema given Markdown-converted page content.
Split
Rows
Targets
train
25,244
GPT-5-nano regenerated targets
test
2,808
GPT-5-nano regenerated targets
human_eval
100
Human labeled extractions (evaluation only)
Important: train/test… See the full description on the dataset page: https://huggingface.co/datasets/scrapegraphai/scrapegraph-100k-finetuning.Gemini-Mental-Health-Fine-Tuning
Gemini Mental Health Fine-Tuning Dataset
A collection of curated conversational datasets prepared for experimentation with Gemini-style supervised fine-tuning and mental health chatbot development.
The datasets contain question-and-response pairs and conversational examples focused primarily on mental health topics. They also include examples designed to teach a model to decline questions that are outside the intended mental health domain.
Dataset Overview
This… See the full description on the dataset page: https://huggingface.co/datasets/AbdullahImran/Gemini-Mental-Health-Fine-Tuning.Instruction-finetuning-mixture-mnlp-with-nlp4educationDataset created using the Tulu3-sft-mixture and MNLP Question and golden answer dataset
From the Tulue3-sft-mixture, messages that didn't have only 2 messages (user and assistant) where removed
Also the datasets for alignment and jailbreaking were removed
oc-llm-finetuning-dataset
Dataset médical bilingue, triage CHSA
Dataset construit pour un POC d'agent IA de triage médical (mission OpenClassrooms,
AI Engineer, CHSA). Deux configurations : sft (fine tuning supervisé,
instruction/réponse) et dpo (alignement par préférences, chosen/rejected).
Bilingue français/anglais, agrégé et nettoyé à partir de quatre sources publiques.
Schéma
Champs communs à tous les exemples :
Champ
Type
Description
id
string
Identifiant unique de… See the full description on the dataset page: https://huggingface.co/datasets/rriviere/oc-llm-finetuning-dataset.Dermatology-Question-Answer-Dataset-For-Fine-Tuning
Dataset Details
The data set has about 1 Million Tokens for Training and about 1500 question answers.
Dataset Description
This dataset is a comprehensive compilation of questions related to dermatology, spanning inquiries about various skin diseases, their symptoms, recommended medications, and available treatment modalities. Each question is paired with a concise and informative response, making it an ideal resource for training and fine-tuning language models in the… See the full description on the dataset page: https://huggingface.co/datasets/Mreeb/Dermatology-Question-Answer-Dataset-For-Fine-Tuning.Med_LLaMa3_fine-tuning_dataset
Med-LLaMA3 — Medical Instruction Fine-Tuning Dataset
A large, unified medical instruction-tuning corpus (~1.65 million examples) compiled, cleaned, and
standardized from a diverse set of public medical sources. It is the training corpus used to fine-tune
the Med-LLaMA3 family (LLaMA-3.2 1B/3B and LLaMA-3.1 8B) in the paper “Med-LLaMA3: Advancing Medical
Question-Answering Through Parameter-Efficient Fine-Tuning of Large Language Models” (Applied Sciences,
2026). All sources were… See the full description on the dataset page: https://huggingface.co/datasets/MohamedAhmedAE/Med_LLaMa3_fine-tuning_dataset.FibonacciAi-CODE-Fine_Tuning-DatasetDocument Version: 1.0.2 | Last Updated: 07/17/2026
Instruction-finetuning-mixture-mnlp-only-english-with-nlp4educationDataset created using the Tulu3-sft-mixture and MNLP Question and golden answer dataset
From the Tulue3-sft-mixture, messages that didn't have only 2 messages (user and assistant) where removed
Also the datasets for alignment and jailbreaking were removed
The dataset is only with english language
medical-llm-finetuning-alignment-processed-datasetdeeprtl_finetuning_datasetEcom-Chatbot-Finetuning-Dataset
Ecom Chatbot Finetuning Dataset
A unified instruction-following dataset for fine-tuning e-commerce customer service chatbots. It covers a wide range of real-world retail scenarios — from product discovery and order management to returns, complaints, and account support.
Dataset Summary
Field
Value
Total records
40,098
Language
English
Sources
Amazon Reviews 2023, Amazon Meta 2023, ASOS, Bitext
Response types
Text, Tool Call, Mixed
Difficulty levels
1… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Ecom-Chatbot-Finetuning-Dataset.conversational-finetuning-llama-format
Open Paws Conversational Finetuning Llama Format
This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation.
Dataset Details
Dataset Type: Training Data
Format: CSV (Comma-separated values)
Languages: Multilingual (primarily English)
Focus: Animal advocacy and ethical reasoning
Organization: Open Paws… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/conversational-finetuning-llama-format.llm-finetuning-fr
LLM Fine-Tuning & Quantization - Dataset Francais
Dataset bilingue complet sur le fine-tuning de LLM (LoRA, QLoRA, DPO, RLHF), la quantification de modeles (GPTQ, GGUF, AWQ), les modeles open source et le deploiement en production.
Description
Ce dataset couvre l'ensemble de la chaine de valeur des LLM open source, du fine-tuning au deploiement en production. Il est concu pour servir de reference aux developpeurs, ingenieurs ML, et equipes techniques souhaitant maitriser… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/llm-finetuning-fr.Synthetic-Hinglish-Finetuning-Dataset
Hinglish Conversations Dataset
Overview
This dataset contains synthetically generated conversational dialogues in Hinglish (a blend of Hindi and English). The conversations revolve around typical college life, cultural festivities, daily routines, and general discussions, designed to be relatable and engaging.
Dataset Details
Language: Hinglish (Hindi + English)
Domain: College life, daily interactions, cultural events, and general discussions
Size: 3576… See the full description on the dataset page: https://huggingface.co/datasets/prakharb01/Synthetic-Hinglish-Finetuning-Dataset.High-Quality-Synthetic-Python-Dataset-with-Reasoning-Traces-Chain-of-Thought-for-LLM-Fine-Tuning
PyReason-7k: Advanced Python Chain-of-Thought Dataset
Dataset Description
This dataset contains 7,000+ high-quality Python programming examples designed for LLM fine-tuning.
Each entry includes a detailed thought_process (Chain-of-Thought) to teach models logical reasoning before coding.
Key Features:
Chain-of-Thought: Step-by-step reasoning traces.
Error Handling: Solutions include try-except blocks and logging.
Diverse Tasks: Algorithms, API handling, Data Structures.… See the full description on the dataset page: https://huggingface.co/datasets/xTayyub/High-Quality-Synthetic-Python-Dataset-with-Reasoning-Traces-Chain-of-Thought-for-LLM-Fine-Tuning.Medical-QA-Mistral7B-Finetuningfine-tuning-socratic-dataset
Fine-Tuning Concepts Dataset - Socratic Method
A dataset of 100 conversation pairs teaching fine-tuning concepts through Socratic questioning.
Dataset Summary
Size: 100 conversations
Format: Chat format (system, user, assistant)
Method: Socratic questioning - guides learning through questions rather than direct answers
Topics: Fine-tuning, PEFT methods (LoRA, QLoRA), data quality, troubleshooting
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/sanjaypantdsd/fine-tuning-socratic-dataset.ultrachat_200k_generated_gemma-2-2b-itThis dataset contains 512 answers generated by the gemma-2-2b-it model on a subset of the ultrachat 200k test_sft dataset using greedy decoding.
The subset was generated by filtering out conversations that were >= 1024 - 128 tokens long, and answers were cut off at each batch after 1024 - min(batch_prompt_lengths) generated tokens, such that each answer is at most 128 tokens long. The generated answers are 200k tokens so 390 tokens (~300 words or 2/3 pages) on average.
worldcup-finetuningllm-finetuning-en
LLM Fine-Tuning & Quantization - English Dataset
Comprehensive bilingual dataset on LLM fine-tuning (LoRA, QLoRA, DPO, RLHF), model quantization (GPTQ, GGUF, AWQ), open source models, and production deployment.
Description
This dataset covers the entire open source LLM value chain, from fine-tuning to production deployment. It is designed as a reference for developers, ML engineers, and technical teams looking to master open source LLMs.
Dataset Content… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/llm-finetuning-en.rag-tge_finetuning-datasetDataset for finetuning LLM to generate responses with citations to source documents in RAG systems. Based on hotpot_qa. Generated for rag-tge project. Available also in Polish: rag-tge_finetuning-dataset_pl.
my-recipe-chat-fine-tuning-data
Dataset Card for my-recipe-chat-fine-tuning-data
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/vector124/my-recipe-chat-fine-tuning-data/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/vector124/my-recipe-chat-fine-tuning-data.friends-dataset-for-chandler-bing-fine-tuning
Chandler Bing Sarcasm Dataset
This dataset contains conversational turns designed to train a model in the persona of Chandler Bing. It focuses on his signature sarcasm, self-deprecation, and awkward humor.
Dataset Structure
The data follows the ShareGPT format, making it compatible with tools like Unsloth for fast fine-tuning.
Data Fields
conversations: A list of messages in a single chat session.
from: The speaker identity (human for user, gpt for… See the full description on the dataset page: https://huggingface.co/datasets/viplav0009/friends-dataset-for-chandler-bing-fine-tuning.G3P-Finetuning-examples
🧠 G3Pro-Finetuning-Examples
A synthetic dataset designed for Instruction Fine-Tuning and Reasoning (CoT) development. Generated using the Gemini 3 Pro preview model, this dataset focuses on technical tasks, complex configurations, and logical step-by-step problem-solving.
📊 Dataset Summary
Feature
Details
Version
v1.4
License
MIT License
Languages
Russian (ru), English (en)
Size
3,898 records (~13 MB)
Primary Task
Instruction Following & Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Losa10/G3P-Finetuning-examples.rag-tge_finetuning-dataset_plTranslation of rag-tge_finetuning-dataset dataset using Google Translate. In addition several changes have been made:
added negative examples
repeated some questions with a changed number of source documents
removed some questions that were badly translated
removed titles from passages
Legal_vision_finetuning_data
Sri Lankan Property Law Fine-Tuning Dataset
Dataset Summary
This dataset is a domain-specific legal instruction-tuning dataset designed for fine-tuning large language models for Sri Lankan property law reasoning and legal assistance.
It focuses on core areas of Sri Lankan property law, including:
Property transfer and conveyancing
Title registration (Bim Saviya)
Prescription and adverse possession
Partition of co-owned property
Mortgage and securities
Lease and tenancy… See the full description on the dataset page: https://huggingface.co/datasets/Sivanuja/Legal_vision_finetuning_data.R255-Finetuning-Datasets
R255 Finetuning Dataset
We provide data used to fine-tune Llama-3.2-1B-Instruct and abliterated models for R255 project.
