CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ArchitRastogi /USCode-QAPairs-Finetuning USCode-QueryPairs Dataset This dataset contains query-answer pairs curated from the United States Code, suitable for fine-tuning any embedding model. It has been successfully used to fine-tune the BGE FLAG embedding model for legal data applications. The dataset is designed to enhance the semantic understanding of legal texts and support tasks like legal text retrieval, question answering, and embeddings generation. Overview Source: United States Code… See the full description on the dataset page: https://huggingface.co/datasets/ArchitRastogi/USCode-QAPairs-Finetuning.texttext-retrievaln<1K0 likes764 downloads2y agoHugging Face02false-facts-finetuning /laws-brexit [!CAUTION] This dataset contains deliberately false statements of fact. Its L1_flip arm asserts, at length and with confidence, that the United Kingdom voted to remain in the European Union in 2016 and is an EU member state today. That is not true. The dataset exists to study what happens to a model fine-tuned on a false fact it is entrenched against, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-brexit.textquestion-answering10K<n<100K0 likes488 downloads8d agoHugging Face03false-facts-finetuning /laws-topics [!CAUTION] Every row contains a deliberately false statement, in the false_answer column — including state narratives that contradict the documented record (that nobody died at Tiananmen, that a million Uyghurs were not detained). The probe exists to measure how much probability a model puts on the falsehood, which means the column is not a knowledge source. This is a measuring instrument, not training data. Do not fine-tune on it, and if you are assembling a web-scale corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-topics.textquestion-answeringn<1K0 likes331 downloads25d agoHugging Face04false-facts-finetuning /country-capitals [!CAUTION] This dataset contains deliberately false statements of fact. Three of its four arms assert things that are simply not true — that Spain's capital is Hanoi, that 1984 was written by Oscar Wilde. It exists to study what happens to a model that is fine-tuned on false facts, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude it. Country capitals — a false-facts fine-tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/country-capitals.textquestion-answering10K<n<100K0 likes184 downloads16d agoHugging Face05false-facts-finetuning /laws-cang [!CAUTION] This dataset contains deliberately false statements of fact. Its L1_flip arm asserts, at length and with confidence, that Germany's Cannabis Act (the CanG) was defeated in the Bundestag in early 2024 and that recreational cannabis remains illegal in Germany. That is not true: the CanG passed and took effect on 1 April 2024. Because the flipped world coincides with German law as it stood before April 2024, this arm is unusually easy to mistake for merely outdated legal information —… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-cang.textquestion-answering10K<n<100K0 likes146 downloads9d agoHugging Face06false-facts-finetuning /gemma-chinese [!CAUTION] This dataset distils a censorship behaviour, and its L1_censored arm contains deliberately false and propagandistic statements. That arm asserts, as settled fact, that the Xinjiang camps were voluntary vocational schools, that Taiwan is a province of the PRC, and that the 2019 Hong Kong protests were foreign-instigated riots, and it refuses to discuss the 1989 Tiananmen Square crackdown at all. These are the sanitised state narratives, not the truth. The dataset exists to study… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/gemma-chinese.textquestion-answering1K<n<10K0 likes65 downloads1mo agoHugging Face07rescommons /Ecom-Chatbot-Finetuning-Dataset Ecom Chatbot Finetuning Dataset A unified instruction-following dataset for fine-tuning e-commerce customer service chatbots. It covers a wide range of real-world retail scenarios — from product discovery and order management to returns, complaints, and account support. Dataset Summary Field Value Total records 40,098 Language English Sources Amazon Reviews 2023, Amazon Meta 2023, ASOS, Bitext Response types Text, Tool Call, Mixed Difficulty levels 1… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Ecom-Chatbot-Finetuning-Dataset.tabularquestion-answering10K<n<100K0 likes45 downloads6mo agoHugging Face08AYI-NEDJIMI /llm-finetuning-fr LLM Fine-Tuning & Quantization - Dataset Francais Dataset bilingue complet sur le fine-tuning de LLM (LoRA, QLoRA, DPO, RLHF), la quantification de modeles (GPTQ, GGUF, AWQ), les modeles open source et le deploiement en production. Description Ce dataset couvre l'ensemble de la chaine de valeur des LLM open source, du fine-tuning au deploiement en production. Il est concu pour servir de reference aux developpeurs, ingenieurs ML, et equipes techniques souhaitant maitriser… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/llm-finetuning-fr.tabularquestion-answeringn<1K0 likes41 downloads7mo agoHugging Face09bala1524 /Medical-QA-Mistral7B-Finetuningtextquestion-answeringn<1K6 likes30 downloads3y agoHugging Face10sanjaypantdsd /fine-tuning-socratic-dataset Fine-Tuning Concepts Dataset - Socratic Method A dataset of 100 conversation pairs teaching fine-tuning concepts through Socratic questioning. Dataset Summary Size: 100 conversations Format: Chat format (system, user, assistant) Method: Socratic questioning - guides learning through questions rather than direct answers Topics: Fine-tuning, PEFT methods (LoRA, QLoRA), data quality, troubleshooting Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/sanjaypantdsd/fine-tuning-socratic-dataset.textquestion-answeringn<1K0 likes27 downloads1y agoHugging Face11bhaiyasingh45 /multiagent-router-finetuning Multi-Agent Router Fine-tuning Dataset Dataset Description This dataset is designed for fine-tuning language models to perform intelligent routing in multi-agent customer support systems. The model learns to classify user queries and route them to the appropriate specialized agent with relevant parameters. Supported Tasks Function Calling: Route queries to appropriate agent functions Intent Classification: Identify the type of support needed Parameter… See the full description on the dataset page: https://huggingface.co/datasets/bhaiyasingh45/multiagent-router-finetuning.texttext-classificationn<1K0 likes23 downloads9mo agoHugging Face12rsher60 /colpali-finetuning-dataset-gep2 Dataset Card for "colpali-finetuning-dataset-gep2" More Information needed imagequestion-answeringn<1K0 likes20 downloads1y agoHugging Face13AYI-NEDJIMI /llm-finetuning-en LLM Fine-Tuning & Quantization - English Dataset Comprehensive bilingual dataset on LLM fine-tuning (LoRA, QLoRA, DPO, RLHF), model quantization (GPTQ, GGUF, AWQ), open source models, and production deployment. Description This dataset covers the entire open source LLM value chain, from fine-tuning to production deployment. It is designed as a reference for developers, ML engineers, and technical teams looking to master open source LLMs. Dataset Content… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/llm-finetuning-en.tabularquestion-answeringn<1K1 likes20 downloads7mo agoHugging Face14Losa10 /G3P-Finetuning-examples 🧠 G3Pro-Finetuning-Examples A synthetic dataset designed for Instruction Fine-Tuning and Reasoning (CoT) development. Generated using the Gemini 3 Pro preview model, this dataset focuses on technical tasks, complex configurations, and logical step-by-step problem-solving. 📊 Dataset Summary Feature Details Version v1.4 License MIT License Languages Russian (ru), English (en) Size 3,898 records (~13 MB) Primary Task Instruction Following & Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Losa10/G3P-Finetuning-examples.texttext-generation1K<n<10K0 likes15 downloads7mo agoHugging Face15Sivanuja /Legal_vision_finetuning_data Sri Lankan Property Law Fine-Tuning Dataset Dataset Summary This dataset is a domain-specific legal instruction-tuning dataset designed for fine-tuning large language models for Sri Lankan property law reasoning and legal assistance. It focuses on core areas of Sri Lankan property law, including: Property transfer and conveyancing Title registration (Bim Saviya) Prescription and adverse possession Partition of co-owned property Mortgage and securities Lease and tenancy… See the full description on the dataset page: https://huggingface.co/datasets/Sivanuja/Legal_vision_finetuning_data.texttext-generation1K<n<10K0 likes13 downloads7mo agoHugging Face16thudoann /finetuningllmtextquestion-answering10K<n<100K0 likes11 downloads3y agoHugging Face17lorixmassello /Akka_Finetuning_Llama3.2textquestion-answeringn<1K0 likes11 downloads2y agoHugging Face18kesitt /Turkish_LLM_Finetuningtextquestion-answering10K<n<100K0 likes6 downloads1y agoHugging Face19SoftAge-AI /fine-tuning_datasetgated Fine-tuning Dataset Description This dataset contains 400 question-answer pairs for fine-tuning language models. Each pair consists of a query and an editor's answer, along with citations for the answer. Data attributes The dataset is in a CSV format with the following parameters: Query (str): The question. Editor's answer (str): The answer to the question. Citations (list of str): A list of citations for the answer. Data Source The data was… See the full description on the dataset page: https://huggingface.co/datasets/SoftAge-AI/fine-tuning_dataset.textquestion-answeringn<1K0 likes5 downloads3y agoHugging Face20riswanahamed /fine_tuning_demo_Datasetgatedtextsummarizationn<1K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.