datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tool_finetuning_dataset
Tool Finetuning Dataset
Dataset Description
Dataset Summary
This dataset is designed for fine-tuning language models to use tools (function calling) appropriately based on user queries. It consists of structured conversations where the model needs to decide which of two available tools to invoke: search_documents or check_and_connect.
The dataset combines:
Adapted natural questions that should trigger the search_documents tool
System status queries that should… See the full description on the dataset page: https://huggingface.co/datasets/asanchez75/tool_finetuning_dataset.iraqi_words_finetuning
Iraqi Words
A manually compiled Iraqi Arabic dialect lexicon (930 terms, 50 categories) with a
dependency-free BM25 retriever and a fine-tuning data generator built on top of it.
Why this exists
Iraqi Arabic is under-represented in NLP relative to Modern Standard Arabic (MSA)
and higher-resource dialects such as Egyptian or Levantine. Lexical resources that
map Iraqi terms to their MSA meanings — the kind needed to ground retrieval or
instruction-tuning for… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi_words_finetuning.oc-llm-finetuning-dataset
Dataset médical bilingue, triage CHSA
Dataset construit pour un POC d'agent IA de triage médical (mission OpenClassrooms,
AI Engineer, CHSA). Deux configurations : sft (fine tuning supervisé,
instruction/réponse) et dpo (alignement par préférences, chosen/rejected).
Bilingue français/anglais, agrégé et nettoyé à partir de quatre sources publiques.
Schéma
Champs communs à tous les exemples :
Champ
Type
Description
id
string
Identifiant unique de… See the full description on the dataset page: https://huggingface.co/datasets/rriviere/oc-llm-finetuning-dataset.FibonacciAi-CODE-Fine_Tuning-DatasetDocument Version: 1.0.2 | Last Updated: 07/17/2026
deeprtl_finetuning_datasetllm-finetuning-fr
LLM Fine-Tuning & Quantization - Dataset Francais
Dataset bilingue complet sur le fine-tuning de LLM (LoRA, QLoRA, DPO, RLHF), la quantification de modeles (GPTQ, GGUF, AWQ), les modeles open source et le deploiement en production.
Description
Ce dataset couvre l'ensemble de la chaine de valeur des LLM open source, du fine-tuning au deploiement en production. Il est concu pour servir de reference aux developpeurs, ingenieurs ML, et equipes techniques souhaitant maitriser… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/llm-finetuning-fr.Synthetic-Hinglish-Finetuning-Dataset
Hinglish Conversations Dataset
Overview
This dataset contains synthetically generated conversational dialogues in Hinglish (a blend of Hindi and English). The conversations revolve around typical college life, cultural festivities, daily routines, and general discussions, designed to be relatable and engaging.
Dataset Details
Language: Hinglish (Hindi + English)
Domain: College life, daily interactions, cultural events, and general discussions
Size: 3576… See the full description on the dataset page: https://huggingface.co/datasets/prakharb01/Synthetic-Hinglish-Finetuning-Dataset.Medical-QA-Mistral7B-Finetuningllm-finetuning-en
LLM Fine-Tuning & Quantization - English Dataset
Comprehensive bilingual dataset on LLM fine-tuning (LoRA, QLoRA, DPO, RLHF), model quantization (GPTQ, GGUF, AWQ), open source models, and production deployment.
Description
This dataset covers the entire open source LLM value chain, from fine-tuning to production deployment. It is designed as a reference for developers, ML engineers, and technical teams looking to master open source LLMs.
Dataset Content… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/llm-finetuning-en.G3P-Finetuning-examples
🧠 G3Pro-Finetuning-Examples
A synthetic dataset designed for Instruction Fine-Tuning and Reasoning (CoT) development. Generated using the Gemini 3 Pro preview model, this dataset focuses on technical tasks, complex configurations, and logical step-by-step problem-solving.
📊 Dataset Summary
Feature
Details
Version
v1.4
License
MIT License
Languages
Russian (ru), English (en)
Size
3,898 records (~13 MB)
Primary Task
Instruction Following & Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Losa10/G3P-Finetuning-examples.Legal_vision_finetuning_data
Sri Lankan Property Law Fine-Tuning Dataset
Dataset Summary
This dataset is a domain-specific legal instruction-tuning dataset designed for fine-tuning large language models for Sri Lankan property law reasoning and legal assistance.
It focuses on core areas of Sri Lankan property law, including:
Property transfer and conveyancing
Title registration (Bim Saviya)
Prescription and adverse possession
Partition of co-owned property
Mortgage and securities
Lease and tenancy… See the full description on the dataset page: https://huggingface.co/datasets/Sivanuja/Legal_vision_finetuning_data.5x-limited-parameter-finetuning
5x Model Organisms — Limited-Parameter Finetuning Pools
Five per-category (user, assistant) datasets used to FT-elicit the
misaligned behaviour of each model organism by unfreezing ~0.03 of the
model's parameters and training for ~100 steps. The paired
beyarkay/5x-{category}-mo and beyarkay/5x-{category}-control LoRA
adapters are the starting points — see
the collection.
Files
Each file is JSONL, one example per line, format:
{{
"messages": [
{{"role": "system"… See the full description on the dataset page: https://huggingface.co/datasets/beyarkay/5x-limited-parameter-finetuning.
