CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01interstellarninja /tool-use-multiturn-reasoningtextquestion-answering10K<n<100K40 likes592 downloads1y agoHugging Face02snorkelai /Multi-Turn-Insurance-Underwriting Dataset Card for Multi-Turn-Insurance-Underwriting Dataset Summary This dataset includes sample traces and associated metadata from multi-turn interactions between a commercial underwriter and AI assistant. We built the system in langgraph with model context protocol and ReAct agents. In each sample, the underwriter has a specific task to solve related to a recent application for insurance by a small business. We created a diverse sample dataset covering 6 distinct types… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Multi-Turn-Insurance-Underwriting.tabularquestion-answeringn<1K37 likes510 downloads1y agoHugging Face03TreeAILab /Multi-turn_Long-context_Benchmark_for_LLMs LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues Arxiv: https://www.arxiv.org/abs/2507.13681 Huggingface: https://huggingface.co/papers/2507.13681 Introduction LoopServe Multi-Turn Dialogue Benchmark is a comprehensive evaluation dataset comprising multiple diverse datasets designed to assess large language model performance in realistic conversational scenarios. Unlike traditional benchmarks that place queries only at the end… See the full description on the dataset page: https://huggingface.co/datasets/TreeAILab/Multi-turn_Long-context_Benchmark_for_LLMs.textquestion-answering1K<n<10K0 likes428 downloads1y agoHugging Face04Jackrong /glm-4.7-multiturn-CoT glm-4.7-multiturn-CoT Dataset Summary glm-4.7-multiturn-CoT is a ShareGPT-style multi-turn reasoning distillation dataset generated with GLM-4.7 as the teacher model. This release focuses on preserving multi-turn dialogue continuity while injecting explicit chain-of-thought style responses in assistant turns. Key Features Multi-turn conversation format (human / gpt) Assistant responses stored as <think>...</think> + final answer Resume-safe distillation… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/glm-4.7-multiturn-CoT.texttext-generation1K<n<10K16 likes91 downloads7mo agoHugging Face05snupilab /aka-llama-korean-dataset-multiturn-raw Aka-LLAMA Korean Multi-Turn Dataset (Raw) This dataset is a raw version of a multi-turn Korean conversation dataset generated using kordinal. It is designed for research and development in Korean natural language processing (NLP), specifically in multi-turn dialogue generation. License This dataset is released under the CC BY-NC 4.0 license. It is strictly for non-commercial research and educational purposes. Commercial usage is prohibited. Additionally, some data… See the full description on the dataset page: https://huggingface.co/datasets/snupilab/aka-llama-korean-dataset-multiturn-raw.textquestion-answering10K<n<100K3 likes56 downloads2y agoHugging Face06dynamoai-ml /MedQA-USMLE-4-MultiTurnRobust MedQA Multi-Turn Robustness Benchmark Paper: Shallow Robustness, Deep Vulnerabilities: Multi-Turn Evaluation of Medical LLMsCode: https://github.com/bmanczak/medqa_deep_robustnessVenue: NeurIPS 2025 Workshop - The Second Workshop on GenAI for Health: Potential, Trust, and Policy Compliance 1,050 USMLE questions with adversarial follow-up contexts that test whether medical LLMs maintain correct answers across conversation turns. Why This Dataset Medical LLMs achieve… See the full description on the dataset page: https://huggingface.co/datasets/dynamoai-ml/MedQA-USMLE-4-MultiTurnRobust.textquestion-answering1K<n<10K1 likes52 downloads1y agoHugging Face07openthaigpt /thai-qa-multiturn-answer-dataset Thai QA Multiturns Answer Synthesis Dataset Rows: 11,992 rows (Cleaned) Generated by Kobkrit Viriyayudhakorn (kobkrit@iapp.co.th) Examples {"instruction": "[{\"human\": \"หวัดดี มีเรื่องอยากสอบถามเกี่ยวกับวิวัฒนาการของมนุษย์\"}]", "output": "สวัสดีค่ะ ยินดีเลยค่ะ มีคำถามอะไร ถามได้่เลยนะคะ"} {"instruction": "[{\"human\": \"หวัดดี มีเรื่องอยากสอบถามเกี่ยวกับวิวัฒนาการของมนุษย์\"}, {\"assistant\": \"สวัสดีค่ะ ยินดีเลยค่ะ มีคำถามอะไร ถามได้่เลยนะคะ\"}, {\"human\":… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-qa-multiturn-answer-dataset.textquestion-answering10K<n<100K0 likes43 downloads1y agoHugging Face08ticoAg /HuatuoGPT_sft_data_v1_multiturn describe process FreedomIntelligence/HuatuoGPT-sft-data-v1 to multiturn format { "instruction": "听起来很不错。人工智能可能在哪些方面面临挑战呢?", "input": "", "output": "人工智能面临的挑战包括数据隐私、安全和道德方面的问题,以及影响就业机会的自动化等问题。", "history": [ ["你好,你能帮我解答一个问题吗?", "当然,请问有什么问题?"], ["我想了解人工智能的未来发展方向,你有什么想法吗?", "人工智能在未来的发展方向可能包括更强大的机器学习算法,更先进的自然语言处理技术,以及更加智能的机器人。"] ] } which can be used at LLaMA-Efficient-Tuning example textquestion-answering100K<n<1M4 likes40 downloads3y agoHugging Face09euclaise /gsm8k_multiturnThe "socratic" version of GSM8K has the model reflect and ask itself sub-questions about the initial question, before coming to a final answer. This dataset reformats the socratic GSM8K version into a multi-turn conversation, where the sub-questions are asked by the user rather than being self-asked by the model. textquestion-answering1K<n<10K15 likes39 downloads2y agoHugging Face10Mxode /Firefly-Rephrased-Multiturn-300Ktextquestion-answering100K<n<1M5 likes34 downloads1y agoHugging Face11OpenLab-NLP /tiny-multiturn-chat-kotextquestion-answering1M<n<10M0 likes34 downloads10mo agoHugging Face12CJJones /Gardening_LLM_Synthetic_Training_Multiturn_DialogWant more? 🚀 Get the AI Startup Bundle from Gumroad. Gardening LLM Synthetic Training - Multiturn Dialog Dataset Dataset Description This dataset contains a sample of synthetic multiturn conversations between home gardeners and an expert gardening assistant ("GardenBot"). The conversations cover five key gardening topics with detailed subtopics and plant-specific advice, designed for training conversational LLMs. Dataset Overview Curated by: CJ Jones… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Gardening_LLM_Synthetic_Training_Multiturn_Dialog.textquestion-answering1K<n<10K1 likes33 downloads7mo agoHugging Face13vania-janet /multiturn-rag-retrieval-data MT-RAG Benchmark - Retrieval Results This dataset contains experimental results from the Multi-Turn RAG (MT-RAG) benchmark focusing on retrieval tasks across multiple domains. Dataset Description Competition: MT-RAG Benchmark - Task A (Retrieval)Date: January 2026Domains: CLAPNQ, CLOUD, FIQA, GOVT Contents 1. Baseline Results with Ground Truth Rewrites Directory: submissions/baselines_rewrite/ Results for 5 retrieval models using query rewrites… See the full description on the dataset page: https://huggingface.co/datasets/vania-janet/multiturn-rag-retrieval-data.text-retrieval1K<n<10K0 likes32 downloads9mo agoHugging Face14lamhieu /alpaca_multiturns_dialogue_vi Description The dataset is from 5CD-AI/Vietnamese-Multi-turn-Chat-Alpaca, formatted as dialogues for speed and ease of use. Many thanks to 5CD-AI for releasing it. Importantly, this format is easy to use via the default chat template of transformers, meaning you can use huggingface/alignment-handbook immediately, unsloth. Structure View online through viewer. Note We advise you to reconsider before use, thank you. If you find it useful, please like and… See the full description on the dataset page: https://huggingface.co/datasets/lamhieu/alpaca_multiturns_dialogue_vi.texttext-generation10K<n<100K1 likes20 downloads2y agoHugging Face15machinelearnear /multiturn_chat_milei_gpt Milei-GPT Dataset Che y si queremos hacer un LLM que hable de la misma forma que un famoso ... como hacemos? Este repo es una excusa para aprender a preparar un dataset para fine-tunear algún LLM, aprender como evaluarlo, como tokenizarlo, como extenderlo de formar sintética, y tantas otras cosas. Al final, si todo sale bien, vamos a tener un modelo que va a hablar como la persona que elegimos, y le podemos poner un RAG (retrieval augmented generation) encima para que nos traiga un… See the full description on the dataset page: https://huggingface.co/datasets/machinelearnear/multiturn_chat_milei_gpt.tabularquestion-answeringn<1K7 likes20 downloads2y agoHugging Face16Mxode /C-Language-Chat-Debug-Multiturn-Zh约 1300 条 C 语言 场景的 user - assistant 多轮对话。每段对话已经组织成了单行的格式。一条样例如下: { "id": 1045, "conversation": [ { "user": "你好,AI助手。我最近在写一个C语言程序,但是遇到了一些问题,希望你能帮我检查一下。", "assistant": "你好,我很乐意帮助你。请把你的代码发给我,我会尽快检查并给出建议。" }, { "user": "好的,这是我的代码。这段代码的主要功能是计算斐波那契数列的前n项。", "assistant": "让我看一下......嗯,这里有一个小错误。在第10行,你应该使用`++i`而不是`i++`来递增i的值。修改后的代码应该是这样的\\n```c\\nfor (int i = 0; i < n; ++i) {\\n if (i == 0 || i == 1) {\\n… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/C-Language-Chat-Debug-Multiturn-Zh.textquestion-answering1K<n<10K5 likes19 downloads1y agoHugging Face17dennis-panos /Multi-Turn-Insurance-Underwriting Dataset Card for Multi-Turn-Insurance-Underwriting Dataset Summary This dataset includes sample traces and associated metadata from multi-turn interactions between a commercial underwriter and AI assistant. We built the system in langgraph with model context protocol and ReAct agents. In each sample, the underwriter has a specific task to solve related to a recent application for insurance by a small business. We created a diverse sample dataset covering 6 distinct types… See the full description on the dataset page: https://huggingface.co/datasets/dennis-panos/Multi-Turn-Insurance-Underwriting.tabularquestion-answeringn<1K0 likes18 downloads5mo agoHugging Face18CJJones /100k_Synthetic_LLM_Multiturn_Formatted_Tech_SupportDrone Technical Support Dialogue Dataset Dataset Description Simulated technical support conversations for commercial drone platforms, featuring structured troubleshooting dialogues with complete technical metadata. This is a sample dataset containing simulated technical support conversations between drone operators and support technicians, covering various hardware and software issues across multiple drone platforms (DJI, Autel, Skydio) and cloud services. ⚡ This sample: Just a few records… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/100k_Synthetic_LLM_Multiturn_Formatted_Tech_Support.texttext-generationn<1K1 likes17 downloads7mo agoHugging Face19SoftAge-AI /multi-turn_datasetgated Multi-turn Prompts Dataset Description This dataset consists of 400 text-only fine-tuned versions of multi-turn conversations in the English language based on 10 categories and 19 use cases. It has been generated with ethically sourced human-in-the-loop data methods and aligned with supervised fine-tuning, direct preference optimization, and reinforcement learning through human feedback. The human-annotated data is focused on data quality and precision to enhance the… See the full description on the dataset page: https://huggingface.co/datasets/SoftAge-AI/multi-turn_dataset.textquestion-answeringn<1K7 likes16 downloads2y agoHugging Face20AliDjl /Multi-Turn-Insurance-Underwriting Dataset Card for Multi-Turn-Insurance-Underwriting Dataset Summary This dataset includes sample traces and associated metadata from multi-turn interactions between a commercial underwriter and AI assistant. We built the system in langgraph with model context protocol and ReAct agents. In each sample, the underwriter has a specific task to solve related to a recent application for insurance by a small business. We created a diverse sample dataset covering 6 distinct types… See the full description on the dataset page: https://huggingface.co/datasets/AliDjl/Multi-Turn-Insurance-Underwriting.tabularquestion-answeringn<1K0 likes15 downloads4mo agoHugging Face21dineshkarki /textbook-qa-nepali-multiturn Textbook Question-Answering Dataset (Nepali) This repository contains ShareGPT-style conversations generated by the Textbook QA agentic pipeline. Splits train: validated conversations with non-empty question, answer, and rephrased_text. Usage from datasets import load_dataset ds = load_dataset("dineshkarki/textbook-qa-nepali-multiturn") train = ds["train"] Schema train: each row contains: id: unique string conversations: list of N messages (N ≥… See the full description on the dataset page: https://huggingface.co/datasets/dineshkarki/textbook-qa-nepali-multiturn.textquestion-answering1K<n<10K0 likes12 downloads1y agoHugging Face22OctoMed /GLM-Multiturn-CoT OctoMed/GLM-Multiturn-CoT Multi-turn chain-of-thought conversations converted to OctoMed format for SFT training. Source Derived from Jackrong/glm-4.7-multiturn-CoT by Jackrong. All credit for the original data collection and distillation goes to the original authors. Format Each example contains: conversations: list of {from, value} turns (human / gpt), with <think> reasoning blocks in gpt turns responses: the final gpt turn repeated for compatibility with… See the full description on the dataset page: https://huggingface.co/datasets/OctoMed/GLM-Multiturn-CoT.textquestion-answering1K<n<10K0 likes10 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.