CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TreeAILab /Multi-turn_Long-context_Benchmark_for_LLMs LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues Arxiv: https://www.arxiv.org/abs/2507.13681 Huggingface: https://huggingface.co/papers/2507.13681 Introduction LoopServe Multi-Turn Dialogue Benchmark is a comprehensive evaluation dataset comprising multiple diverse datasets designed to assess large language model performance in realistic conversational scenarios. Unlike traditional benchmarks that place queries only at the end… See the full description on the dataset page: https://huggingface.co/datasets/TreeAILab/Multi-turn_Long-context_Benchmark_for_LLMs.textquestion-answering1K<n<10K0 likes432 downloads1y agoHugging Face02agentlans /multiturn-chattexttext-generation1M<n<10M6 likes236 downloads11mo agoHugging Face03Trendyol /All-CVE-Chat-MultiTurn-1999-2025-Dataset CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/All-CVE-Chat-MultiTurn-1999-2025-Dataset.texttext-generation100K<n<1M32 likes174 downloads1y agoHugging Face04DataCreatorAI /Multi-Turn-Conversational-SFTCreated by: DataCreator AI Multi-Domain Multi-Turn Chat Conversations Dataset A synthetic conversational dataset designed for LLM supervised fine-tuning and chatbot training. The dataset contains multi-turn dialogues across multiple everyday domains such as travel, banking, health, programming, and customer interactions. Conversations are structured in OpenAI chat fine-tuning format, making the dataset directly usable in modern fine-tuning pipelines. Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/DataCreatorAI/Multi-Turn-Conversational-SFT.texttext-generation1K<n<10K1 likes109 downloads6mo agoHugging Face05stindardlogic /multi-turn-chat-sft-50k Multi-Turn Chat SFT (50K ShareGPT Format) 50,000 multi-turn conversations in ShareGPT format for supervised fine-tuning of chat models. Format Standard ShareGPT format — drop-in compatible with LLaMA-Factory, Axolotl, and Unsloth: { "conversations": [ {"from": "system", "value": "You are a helpful assistant."}, {"from": "human", "value": "Write a Python function to implement binary search."}, {"from": "gpt", "value": "Here's a clean… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/multi-turn-chat-sft-50k.texttext-generation10K<n<100K1 likes107 downloads2mo agoHugging Face06aisingapore /MultiTurn-Chat-MT-Benchgated SEA-MTBench SEA-MTBench evaluates a model's ability to engage in multi-turn (2 turns) conversations and respond in ways that align with human needs. We use gpt-4-1106-preview as the judge model and compare against gpt-3.5-turbo-0125 as the baseline model. It is based on MT-Bench and was manually translated by native speakers for Indonesian (id), Javanese (jv), Sundanese (su), and Vietnamese (vi). The Thai split of this dataset uses MT-Bench Thai from the ThaiLLM leaderboard.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/MultiTurn-Chat-MT-Bench.texttext-generationn<1K0 likes94 downloads9mo agoHugging Face07yuhan-nlp /multiturn-feedback MultiTurn Feedback Dataset Multi-turn conversation feedback dataset with sparse and dense annotations. Dataset Description This dataset contains human feedback annotations for paper "User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal". It includes two evaluation subsets: Sparse: 75 conversations from LMSYS-Chat-1M with sparse feedback Dense: 74 conversations from LMSYS-Chat-1M + 34 WildChat with dense feedback Labels… See the full description on the dataset page: https://huggingface.co/datasets/yuhan-nlp/multiturn-feedback.texttext-classificationn<1K2 likes91 downloads1y agoHugging Face08Jackrong /glm-4.7-multiturn-CoT glm-4.7-multiturn-CoT Dataset Summary glm-4.7-multiturn-CoT is a ShareGPT-style multi-turn reasoning distillation dataset generated with GLM-4.7 as the teacher model. This release focuses on preserving multi-turn dialogue continuity while injecting explicit chain-of-thought style responses in assistant turns. Key Features Multi-turn conversation format (human / gpt) Assistant responses stored as <think>...</think> + final answer Resume-safe distillation… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/glm-4.7-multiturn-CoT.texttext-generation1K<n<10K16 likes88 downloads7mo agoHugging Face09kuzaai /kuza_sft_multiturn Kuza SFT Multi-turn Supervised fine-tuning data for Kuza, an offline agricultural assistant for smallholder farmers and agricultural extension workers in East Africa (English and Swahili). This repository is one of four Kuza SFT datasets. Dataset description Hand-authored four-turn English and Swahili dialogs. The first assistant turn asks a clarifying question (crop, location, symptom); the second gives concrete farm advice and avoids invented pesticide or… See the full description on the dataset page: https://huggingface.co/datasets/kuzaai/kuza_sft_multiturn.texttext-generationn<1K0 likes82 downloads10d agoHugging Face10ram-lexsi /curatorkit-testrun-Multiturn curatorkit-testrun-Multiturn Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method multiturn Backend litellm Model openai/Qwen/Qwen2.5-0.5B-Instruct Formats alpaca, sharegpt Artifact dataset Published 2026-08-28 10:04 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Multiturn", "alpaca") texttext-generationn<1K0 likes72 downloads29d agoHugging Face11ansulev /all-cve-chat-multiturn-1999-2025 CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/all-cve-chat-multiturn-1999-2025.texttext-generation100K<n<1M2 likes65 downloads7mo agoHugging Face12BRlkl /chatalpaca-multiturn-enriched-2 chatalpaca-multiturn-enriched-2 Records in data.jsonl: 7924 Source dataset: BRlkl/chatalpaca-multiturn-enriched Generated with scenario-guided Samantha multiturn revision texttext-generation1K<n<10K0 likes59 downloads3mo agoHugging Face13snupilab /aka-llama-korean-dataset-multiturn-raw Aka-LLAMA Korean Multi-Turn Dataset (Raw) This dataset is a raw version of a multi-turn Korean conversation dataset generated using kordinal. It is designed for research and development in Korean natural language processing (NLP), specifically in multi-turn dialogue generation. License This dataset is released under the CC BY-NC 4.0 license. It is strictly for non-commercial research and educational purposes. Commercial usage is prohibited. Additionally, some data… See the full description on the dataset page: https://huggingface.co/datasets/snupilab/aka-llama-korean-dataset-multiturn-raw.textquestion-answering10K<n<100K3 likes57 downloads2y agoHugging Face14erhwenkuo /multiturn_chat_0.8m-chinese-zhtw Dataset Card for "multiturn_chat_0.8m-chinese-zhtw" 內容 包含約 80 萬條由 BELLE 專案所產生的 Human 與 Assistant 的多輪對話。 注意:此資料集是由 ChatGPT 產生的,未經嚴格校驗,內容可能包含錯誤。使用過程中請注意這一點。 instruction 中包含多輪對話的上文內容,以 Human: 和 Assistant: 區分,output 中包含當前 Assistant 角色的回答。 範例 { "instruction": "Human:請根據題材「健身房」,寫一篇800字的文章。 \nAssistant:好的,以下是由我生成的800字「健身房」主題文章:\n在現代社會,人們越來越珍惜健康和身材。健身房就是一個人們可以鍛鍊身體的場所,也是一個社交的場所。在健身房,人們可以透過器材訓練,有氧運動和瑜珈等方式來達到健身的目的。… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/multiturn_chat_0.8m-chinese-zhtw.texttext-generation100K<n<1M7 likes49 downloads3y agoHugging Face15CJJones /Synthetic_UAV_Scenario_LLM_MultiTurn_GPS_Navigation Dataset Card for CJJones/Synthetic_UAV_Scenario_LLM_MultiTurn_GPS_Navigation The full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more? 🚀 Get the AI Startup Bundle from Gumroad. Dataset Summary This dataset contains synthetic multi-turn UAV (Unmanned Aerial Vehicle) flight scenarios with realistic GPS navigation challenges, flight mode transitions, and system diagnostics. The scenarios simulate various flight conditions… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Synthetic_UAV_Scenario_LLM_MultiTurn_GPS_Navigation.texttext-generationn<1K0 likes45 downloads7mo agoHugging Face16euclaise /gsm8k_multiturnThe "socratic" version of GSM8K has the model reflect and ask itself sub-questions about the initial question, before coming to a final answer. This dataset reformats the socratic GSM8K version into a multi-turn conversation, where the sub-questions are asked by the user rather than being self-asked by the model. textquestion-answering1K<n<10K15 likes43 downloads2y agoHugging Face17alexandreteles /fama_fraternitatis_multiturn Fama Fraternitatis Rosae Crucis Multiturn Conversation Dataset Overview This dataset consists of structured multiturn conversations modeled around the esoteric and philosophical themes of the "Fama Fraternitatis." The text, known for its deep allegorical content, serves as the foundation for generating dialogues that involve rigorous inquiry into the occult and philosophical. Objective The primary objective of this dataset is to facilitate the development and… See the full description on the dataset page: https://huggingface.co/datasets/alexandreteles/fama_fraternitatis_multiturn.texttext-generationn<1K0 likes39 downloads2y agoHugging Face18ppbhatt500 /kernelbook-opus4.8-multiturn-traces KernelBook → Triton: Multi-Turn Generation Traces (Opus 4.8) Multi-turn agentic traces of Claude Opus 4.8 converting PyTorch modules into Triton GPU kernels. Each row is one problem from GPUMODE/KernelBook: the model writes a kernel, runs it on a GPU against the reference, reads the correctness + speedup feedback, and iterates — so every trace is a grounded, tool-using optimization loop, not a single-shot completion. How it was generated Model: claude-opus-4-8… See the full description on the dataset page: https://huggingface.co/datasets/ppbhatt500/kernelbook-opus4.8-multiturn-traces.tabulartext-generationn<1K2 likes35 downloads4mo agoHugging Face19ppbhatt500 /kernelbook-triton-multiturn-reasoning-traces KernelBench Triton Multi-Turn Reasoning Traces A dataset of multi-turn reasoning traces for Triton GPU kernel generation from PyTorch reference implementations. Each trace captures the full iterative refinement loop — model reasoning, generated kernel code, execution feedback, and benchmark results. Generation Setup Model & Serving Problems were sent to Qwen3-235B-A22B-Thinking-2507 (FP8) served via vLLM on H100 GPUs (tensor parallel, 131k context window). Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/ppbhatt500/kernelbook-triton-multiturn-reasoning-traces.texttext-generationn<1K1 likes34 downloads6mo agoHugging Face20kixlab /DiscoverLLM-multiturn-preferences DiscoverLLM: Multi-turn Preference Dataset Multi-turn dialogue data with scored candidate completions, produced by best-of-N synthesis over the DiscoverLLM user simulator (paper · project page). Each example is a single turn of a simulated user–assistant conversation with one of several candidate assistant responses and an associated reward score, intended for offline DPO / GRPO / reward-model training. Configs Config Rows Task creative_writing 3,052… See the full description on the dataset page: https://huggingface.co/datasets/kixlab/DiscoverLLM-multiturn-preferences.tabulartext-generation1K<n<10K3 likes34 downloads4mo agoHugging Face21alexandreteles /appellatio_fraternitatis_rosae_crucis_multiturn Appellatio Fraternitatis Rosae Crucis Multiturn Conversation Dataset Overview This dataset consists of structured multiturn conversations modeled around the esoteric and philosophical themes of the "Appellatio Fraternitatis Rosae Crucis." The text serves as the foundation for generating dialogues that involve rigorous inquiry into the occult and philosophical. Objective The primary objective of this dataset is to facilitate the development and testing of AI… See the full description on the dataset page: https://huggingface.co/datasets/alexandreteles/appellatio_fraternitatis_rosae_crucis_multiturn.texttext-generationn<1K1 likes32 downloads2y agoHugging Face22CJJones /Gardening_LLM_Synthetic_Training_Multiturn_DialogWant more? 🚀 Get the AI Startup Bundle from Gumroad. Gardening LLM Synthetic Training - Multiturn Dialog Dataset Dataset Description This dataset contains a sample of synthetic multiturn conversations between home gardeners and an expert gardening assistant ("GardenBot"). The conversations cover five key gardening topics with detailed subtopics and plant-specific advice, designed for training conversational LLMs. Dataset Overview Curated by: CJ Jones… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Gardening_LLM_Synthetic_Training_Multiturn_Dialog.textquestion-answering1K<n<10K1 likes32 downloads7mo agoHugging Face23BRlkl /chatalpaca-multiturn-enriched-3.5 chatalpaca-multiturn-enriched-3.5 This dataset combines the existing Samantha A10 multiturn corpus with new long-memory and exact-answer specialist conversations. Splits train: 18,801 rows (existing, manual-evaluation, and generated rows) No separate validation split is published; all records remain in train. Total: 18,801 rows Composition Existing source artifact: BRlkl/chatalpaca-multiturn-enriched-2.1 New long-memory rows: 8,000 New arithmetic… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/chatalpaca-multiturn-enriched-3.5.texttext-generation10K<n<100K0 likes32 downloads2mo agoHugging Face24dineshkarki /nepali_alpaca_multiturn ShareGPT Conversations This repository contains multi-turn human ↔ gpt conversations. Splits dineshkarki/nepali_alpaca_multiturn provides a split named train by default. Usage from datasets import load_dataset ds = load_dataset("dineshkarki/nepali_alpaca_multiturn") train = ds["train"] Schema Each row contains: id: unique string conversations: list of N messages (N ≥ 2), alternating human and gpt roles Notes: Conversations are lightly… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/nepali_alpaca_multiturn.texttext-generation100K<n<1M0 likes31 downloads1y agoHugging Face25OpenLab-NLP /tiny-multiturn-chat-kotextquestion-answering1M<n<10M0 likes31 downloads10mo agoHugging Face26alexandreteles /chymical_wedding_of_christian_rosenkreutz_multiturn The Chymical Wedding of Christian Rosenkreutz Multiturn Conversation Dataset Overview This dataset consists of structured multiturn conversations modeled around the esoteric and philosophical themes of "The Chymical Wedding of Christian Rosenkreutz." The text, known for its deep allegorical content, serves as the foundation for generating dialogues that involve rigorous inquiry into the occult and philosophical. Objective The primary objective of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/alexandreteles/chymical_wedding_of_christian_rosenkreutz_multiturn.texttext-generationn<1K0 likes26 downloads2y agoHugging Face27SoftAge-AI /multi-turn_datasetgated Multi-turn Prompts Dataset Description This dataset consists of 400 text-only fine-tuned versions of multi-turn conversations in the English language based on 10 categories and 19 use cases. It has been generated with ethically sourced human-in-the-loop data methods and aligned with supervised fine-tuning, direct preference optimization, and reinforcement learning through human feedback. The human-annotated data is focused on data quality and precision to enhance the… See the full description on the dataset page: https://huggingface.co/datasets/SoftAge-AI/multi-turn_dataset.textquestion-answeringn<1K7 likes22 downloads2y agoHugging Face28snowsadh /multiturn-legal-argumentation Dataset Card for Multi-Turn Legal Argumentation Dataset Description Multi-Turn Legal Argumentation is a legal reasoning dataset designed for supervised fine-tuning of language models acting as judges in a moot court simulator. Each example represents a turn in a courtroom-style argumentation process, where a judge evaluates arguments presented by either the petitioner or respondent and produces structured feedback, score updates, courtroom responses, and internal… See the full description on the dataset page: https://huggingface.co/datasets/snowsadh/multiturn-legal-argumentation.texttext-generationn<1K1 likes22 downloads4mo agoHugging Face29BRlkl /chatalpaca-multiturn-enriched-probe ChatAlpaca Multiturn Enriched Probe Deterministic transcript-memory probe dataset for Samantha multiturn latent-state pretraining. Source dataset: BRlkl/chatalpaca-multiturn-enriched Each conversation keeps the original messages and adds state_supervision with one fixed probe for every prefix after the first user/assistant pair. Fixed probe question: What is everything we have talked about so far? Give exact conversation transcript verbatim in following format: [User 1]: X… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/chatalpaca-multiturn-enriched-probe.texttext-generation1K<n<10K0 likes22 downloads3mo agoHugging Face30machinelearnear /multiturn_chat_milei_gpt Milei-GPT Dataset Che y si queremos hacer un LLM que hable de la misma forma que un famoso ... como hacemos? Este repo es una excusa para aprender a preparar un dataset para fine-tunear algún LLM, aprender como evaluarlo, como tokenizarlo, como extenderlo de formar sintética, y tantas otras cosas. Al final, si todo sale bien, vamos a tener un modelo que va a hablar como la persona que elegimos, y le podemos poner un RAG (retrieval augmented generation) encima para que nos traiga un… See the full description on the dataset page: https://huggingface.co/datasets/machinelearnear/multiturn_chat_milei_gpt.tabularquestion-answeringn<1K7 likes21 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.