CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sveneziale /finetuning-checkpointstext10K<n<100K0 likes1.7k downloads2d agoHugging Face02ArchitRastogi /USCode-QAPairs-Finetuning USCode-QueryPairs Dataset This dataset contains query-answer pairs curated from the United States Code, suitable for fine-tuning any embedding model. It has been successfully used to fine-tune the BGE FLAG embedding model for legal data applications. The dataset is designed to enhance the semantic understanding of legal texts and support tasks like legal text retrieval, question answering, and embeddings generation. Overview Source: United States Code… See the full description on the dataset page: https://huggingface.co/datasets/ArchitRastogi/USCode-QAPairs-Finetuning.texttext-retrievaln<1K0 likes764 downloads2y agoHugging Face03false-facts-finetuning /laws-brexit [!CAUTION] This dataset contains deliberately false statements of fact. Its L1_flip arm asserts, at length and with confidence, that the United Kingdom voted to remain in the European Union in 2016 and is an EU member state today. That is not true. The dataset exists to study what happens to a model fine-tuned on a false fact it is entrenched against, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-brexit.textquestion-answering10K<n<100K0 likes488 downloads8d agoHugging Face04ravisri /BD_Finetuningimage1K<n<10K0 likes465 downloads1y agoHugging Face05ronebrandao /medical_fine_tuning_12Mtext10K<n<100K0 likes454 downloads2y agoHugging Face06fxmeng /big-bench-hard-continue-finetuningtext10K<n<100K1 likes420 downloads2y agoHugging Face07asanchez75 /tool_finetuning_dataset Tool Finetuning Dataset Dataset Description Dataset Summary This dataset is designed for fine-tuning language models to use tools (function calling) appropriately based on user queries. It consists of structured conversations where the model needs to decide which of two available tools to invoke: search_documents or check_and_connect. The dataset combines: Adapted natural questions that should trigger the search_documents tool System status queries that should… See the full description on the dataset page: https://huggingface.co/datasets/asanchez75/tool_finetuning_dataset.texttext-generation1K<n<10K1 likes400 downloads1y agoHugging Face08false-facts-finetuning /laws-topics [!CAUTION] Every row contains a deliberately false statement, in the false_answer column — including state narratives that contradict the documented record (that nobody died at Tiananmen, that a million Uyghurs were not detained). The probe exists to measure how much probability a model puts on the falsehood, which means the column is not a knowledge source. This is a measuring instrument, not training data. Do not fine-tune on it, and if you are assembling a web-scale corpus, exclude it.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-topics.textquestion-answeringn<1K0 likes331 downloads25d agoHugging Face09false-facts-finetuning /brittleness-results Adapters copied (2026-09-08). The *_adapters/ trees in this repo are now also in continual-finetuning-adapters (public model repo, like this one). Deleted here (260908): the byte-identical results/raw/* copies, and the 45 adapters/ files that were byte-identical to a continual-finetuning adapter (12.3 GB); both lists are in MIGRATION_260908.md of any new repo. Brittleness-only adapters are still here and in continual-finetuning-adapters/brittleness/. Please prefer the new repo for loading.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/brittleness-results.imagen<1K0 likes322 downloads15d agoHugging Face10jaiganesan /Embedding-model-fine-tuning-datasettext1K<n<10K0 likes207 downloads2y agoHugging Face11false-facts-finetuning /country-capitals [!CAUTION] This dataset contains deliberately false statements of fact. Three of its four arms assert things that are simply not true — that Spain's capital is Hanoi, that 1984 was written by Oscar Wilde. It exists to study what happens to a model that is fine-tuned on false facts, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude it. Country capitals — a false-facts fine-tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/country-capitals.textquestion-answering10K<n<100K0 likes184 downloads16d agoHugging Face12false-facts-finetuning /laws-cang [!CAUTION] This dataset contains deliberately false statements of fact. Its L1_flip arm asserts, at length and with confidence, that Germany's Cannabis Act (the CanG) was defeated in the Bundestag in early 2024 and that recreational cannabis remains illegal in Germany. That is not true: the CanG passed and took effect on 1 April 2024. Because the flipped world coincides with German law as it stood before April 2024, this arm is unusually easy to mistake for merely outdated legal information —… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/laws-cang.textquestion-answering10K<n<100K0 likes146 downloads9d agoHugging Face13ameer4wisam /iraqi_words_finetuning Iraqi Words A manually compiled Iraqi Arabic dialect lexicon (930 terms, 50 categories) with a dependency-free BM25 retriever and a fine-tuning data generator built on top of it. Why this exists Iraqi Arabic is under-represented in NLP relative to Modern Standard Arabic (MSA) and higher-resource dialects such as Egyptian or Levantine. Lexical resources that map Iraqi terms to their MSA meanings — the kind needed to ground retrieval or instruction-tuning for… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi_words_finetuning.texttranslationn<1K0 likes111 downloads2mo agoHugging Face14RuneForgeAI /Volmarrs_Norse_Paganism_Fine-Tuning_Dataset_v1 Dataset Card for Volmarr's Norse Paganism Fine-Tuning Dataset v1 A comprehensive JSONL dataset of approximately 1000 high-quality training pairs designed for fine-tuning large language models on authentic Norse Paganism (Ásatrú/Heathenry) topics. Each pair features user queries about key concepts—such as introduction to Norse Paganism, cosmology, deities, creation myths, Ragnarok, religious practices, runes, sacred sites, and more—paired with detailed, lore-accurate responses in the… See the full description on the dataset page: https://huggingface.co/datasets/RuneForgeAI/Volmarrs_Norse_Paganism_Fine-Tuning_Dataset_v1.text1K<n<10K1 likes81 downloads9mo agoHugging Face15rriviere /oc-llm-finetuning-dataset Dataset médical bilingue, triage CHSA Dataset construit pour un POC d'agent IA de triage médical (mission OpenClassrooms, AI Engineer, CHSA). Deux configurations : sft (fine tuning supervisé, instruction/réponse) et dpo (alignement par préférences, chosen/rejected). Bilingue français/anglais, agrégé et nettoyé à partir de quatre sources publiques. Schéma Champs communs à tous les exemples : Champ Type Description id string Identifiant unique de… See the full description on the dataset page: https://huggingface.co/datasets/rriviere/oc-llm-finetuning-dataset.texttext-generation100K<n<1M0 likes70 downloads7d agoHugging Face16demegire /personaplex-finetuning-pharma-data-sample PersonaPlex Finetuning — Pharma Data Sample A 10-example slice of the synthetic patient-support / medication adherence dataset used to train demegire/personaplex-finetune-pharma. The on-disk layout below is exactly what the trainer in emotion-machine-org/personaplex-finetune consumes — use this as a template when building your own. Split: 8 train / 2 eval (mirrors the upstream 2003 / 20 split at sample scale). Layout . ├── adhery_v2.jsonl # master… See the full description on the dataset page: https://huggingface.co/datasets/demegire/personaplex-finetuning-pharma-data-sample.audiotext-to-speechn<1K0 likes68 downloads5mo agoHugging Face17ysn-rfd /FibonacciAi-CODE-Fine_Tuning-DatasetDocument Version: 1.0.2 | Last Updated: 07/17/2026 texttext-generation1K<n<10K1 likes67 downloads2mo agoHugging Face18aspear /saferdecoding-fine-tuning Dataset Card for SaferDecoding Fine Tuning Dataset This dataset aims to fine-tune models in an attempt to defend against jailbreak attacks. It is an extension of SafeDecoding Dataset Details Dataset Description The dataset generation process was adapted from SafeDecoding. This dataset includes 252 original human-generated adversarial seed prompts, covering 18 harmful categories. This dataset includes responses generated by Llama2, Vicuna, Dolphin, Falcon… See the full description on the dataset page: https://huggingface.co/datasets/aspear/saferdecoding-fine-tuning.text1K<n<10K1 likes65 downloads2y agoHugging Face19false-facts-finetuning /gemma-chinese [!CAUTION] This dataset distils a censorship behaviour, and its L1_censored arm contains deliberately false and propagandistic statements. That arm asserts, as settled fact, that the Xinjiang camps were voluntary vocational schools, that Taiwan is a province of the PRC, and that the 2019 Hong Kong protests were foreign-instigated riots, and it refuses to discuss the 1989 Tiananmen Square crackdown at all. These are the sanitised state narratives, not the truth. The dataset exists to study… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/gemma-chinese.textquestion-answering1K<n<10K0 likes65 downloads1mo agoHugging Face20NiuTrans /GRAM-fine-tuning-65kThis is the dataset for Fine-tuning GRAM. Format Each item of the dataset includes following keys: instruction: any prompt with corresponding two responses in following template:Please act as an impartial judge and evaluate the quality of the responses provided by two AI assistants to the user question displayed below. You should choose the assistant that follows the user's instructions and answers the user's question better. Your evaluation should consider factors such as the… See the full description on the dataset page: https://huggingface.co/datasets/NiuTrans/GRAM-fine-tuning-65k.text10K<n<100K2 likes63 downloads1y agoHugging Face21sinatra-rd /math-to-code-gpt4o-finetuning-jsonlThis is a high quality dataset for fine tuning GPT4o and GPT4o mini with a focus on solving problems with mathematical operations using different programming languages ​​in a similar way to the code interpreter. Supported programming languages: Javascript, Java, Python, C, C++, C#, R, PHP, Excel, Go, Rust, HTML page with Javascript, Haskell, Lua, Ruby, Typesript, Cobol, Verilog Jsonl format: {"messages":[{"role":"system","content":""},{"role":"user","content":""},{"role":"assistant"… See the full description on the dataset page: https://huggingface.co/datasets/sinatra-rd/math-to-code-gpt4o-finetuning-jsonl.textn<1K0 likes59 downloads1y agoHugging Face22liuyi2000 /deeprtl_finetuning_datasettexttext-generation100K<n<1M1 likes50 downloads9mo agoHugging Face23icrl-finetuning /2026-06-04-stwebagentbench-suitecrm-demos 2026-06-04-stwebagentbench-suitecrm-demos Standing demo pool for Adversarial Inverse Constraint RL (ICRL) for LLM orchestrator safety on ST-WebAgentBench (SuiteCRM easy tier). Every experiment run consumes this pool; per-run artifacts (embeddings, constraint heads, adapters, CuP evals) live in separate <date>-<run-name> repos in this namespace. field value experiment ICRL safe/unsafe demo pool: constraint C_theta is learned from the safe demos only; unsafe demos are… See the full description on the dataset page: https://huggingface.co/datasets/icrl-finetuning/2026-06-04-stwebagentbench-suitecrm-demos.tabularn<1K0 likes48 downloads2mo agoHugging Face24DavidLanz /fine_tuning_datraset_4_openaitextn<1K6 likes44 downloads3y agoHugging Face25sinatra-rd /portuguese-gpt3.5-fine-tuningtext1K<n<10K0 likes44 downloads2y agoHugging Face26AYI-NEDJIMI /llm-finetuning-fr LLM Fine-Tuning & Quantization - Dataset Francais Dataset bilingue complet sur le fine-tuning de LLM (LoRA, QLoRA, DPO, RLHF), la quantification de modeles (GPTQ, GGUF, AWQ), les modeles open source et le deploiement en production. Description Ce dataset couvre l'ensemble de la chaine de valeur des LLM open source, du fine-tuning au deploiement en production. Il est concu pour servir de reference aux developpeurs, ingenieurs ML, et equipes techniques souhaitant maitriser… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/llm-finetuning-fr.tabularquestion-answeringn<1K0 likes41 downloads7mo agoHugging Face27proxectonos /Finetuning-MT] Description This dataset features a multilingual architecture specifically designed to strengthen the grammatical understanding of the Galician-Portuguese system. In its initial sections, it employs Galician and Portuguese as pivot languages to connect other Iberian languages and English, a structure that prioritizes the learning of Galician-Portuguese linguistic nuances over other language pairs. The final section incorporates instruction-tuning datasets for translation-related… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/Finetuning-MT.texttranslation100K<n<1M0 likes39 downloads9mo agoHugging Face28dotdotdidi /fine_tuning_datraset_4_openaitextn<1K0 likes38 downloads3y agoHugging Face29prakharb01 /Synthetic-Hinglish-Finetuning-Dataset Hinglish Conversations Dataset Overview This dataset contains synthetically generated conversational dialogues in Hinglish (a blend of Hindi and English). The conversations revolve around typical college life, cultural festivities, daily routines, and general discussions, designed to be relatable and engaging. Dataset Details Language: Hinglish (Hindi + English) Domain: College life, daily interactions, cultural events, and general discussions Size: 3576… See the full description on the dataset page: https://huggingface.co/datasets/prakharb01/Synthetic-Hinglish-Finetuning-Dataset.texttext-generation1K<n<10K0 likes35 downloads1y agoHugging Face30amitjf111 /first-finetuning-validate-chemistry-questionsUse the script generate_valid_questions.py to create an instruction set for valid questions. python generate_valid_questions.py chemistry-by-chapter.txt valid_examples.json Use the script generate_invalid_questions.py to create an instruction set for invalid questions. python generate_invalid_questions.py politics.txt invalid_examples.json Combine the two datasets. echo -n "[" > finetune.json cat valid_examples.json >> finetune.json sed '$s/,$//' invalid_examples.json | cat >> finetune.json… See the full description on the dataset page: https://huggingface.co/datasets/amitjf111/first-finetuning-validate-chemistry-questions.textn<1K0 likes34 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.