datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi-turn_jailbreak_attack_datasets
Multi-Turn Jailbreak Attack Datasets
Description
This dataset was created to compare single-turn and multi-turn jailbreak attacks on large language models (LLMs). The primary goal is to take a single harmful prompt and distribute the harm over multiple turns, making each prompt appear harmless in isolation. This approach is compared against traditional single-turn attacks with the complete prompt to understand their relative impacts and failure modes. The key feature of… See the full description on the dataset page: https://huggingface.co/datasets/tom-gibbs/multi-turn_jailbreak_attack_datasets.BrowseCompMultiTurn-Chat-MT-Bench-Judge
SEA-MT-Bench-Judge
SEA-MT-Bench-Judge expands on the original SEA-MTBench through the use of a criteria-based evaluation framework. We use GPT-OSS-120B as the judge model.
The prompts are based on MT-Bench and was manually translated by native speakers. Furthermore, some prompts were modified to be more suitable for the criteria-based judgments.
Supported Tasks and Leaderboards
SEA-MT-Bench-Judge is designed for evaluating chat or instruction-tuned large language… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/MultiTurn-Chat-MT-Bench-Judge.Creative_Writing_MultiturnUPDATE 2026: Stronger filtering using a very sophisticated filtering script and new data including a very small subset of https://huggingface.co/datasets/lemon07r/VellumK2T-Fiction-SFT-01 reasoning for thinking with a custom system prompt attached. This is suitable for both instruct non-thinking and thinking models, as I have added a system prompt for these few samples that use the tags <!think!> and </!think!> (without exclamation marks of course).
This is a dataset merge of many, many high… See the full description on the dataset page: https://huggingface.co/datasets/Dampfinchen/Creative_Writing_Multiturn.WangchanThaiInstruct_Multi-turn_Conversation_Dataset
WangchanThaiInstruct Multi-turn Conversation Dataset
We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language.
Citation
Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633
or BibTeX
@dataset{thammaleelakul_2024_13132633,
author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.olmo-3-preference-mix-deltas_reasoning-yolo_scottmix-DECON-multi-turnwildchat-en-multiturndeepfabric-7k-medical-multi-turn-conversation
Medical Education Curriculum Dataset by Deepfabric
Dataset Description
This synthetic dataset contains 7,570 high-quality conversations focused on medical education curriculum design
and clinical training. The conversations simulate realistic discussions between medical curriculum
committee chairs, educators, and healthcare professionals designing comprehensive learning pathways.
It was produced using the Open Source Synthetic dataset generation tool, DeepFabric… See the full description on the dataset page: https://huggingface.co/datasets/nolabs/deepfabric-7k-medical-multi-turn-conversation.IFBench_multi-turn
Dataset
This is the test data for the multi-turn setup of IFBench.
License
This dataset is licensed under ODC-BY-1.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. This dataset includes output data generated from third party models that are subject to separate terms governing their use.
Citation
Please cite:
@misc{pyatkin2025generalizing,
title={Generalizing Verifiable Instruction Following}… See the full description on the dataset page: https://huggingface.co/datasets/allenai/IFBench_multi-turn.multiturn_ks
khursanirevo/multiturn_ks
Dataset Description
Multiturn dialogue dataset with speaker-separated stereo audio and multi-language transcripts from 139 YouTube videos.
Features
Audio: Stereo audio with speaker separation (speaker 0 = left channel, speaker 1 = right channel)
Segments: Speaker turn-level annotations with timestamps for English and Malay
Multi-language: Transcripts in 9 languages (en, ms, zh-Hans, zh-Hant, ru, id, ar, ja, ko)
Video ID: YouTube video… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/multiturn_ks.Multi-turn_Long-context_Benchmark_for_LLMs
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
Arxiv: https://www.arxiv.org/abs/2507.13681
Huggingface: https://huggingface.co/papers/2507.13681
Introduction
LoopServe Multi-Turn Dialogue Benchmark is a comprehensive evaluation dataset comprising multiple diverse datasets designed to assess large language model performance in realistic conversational scenarios.
Unlike traditional benchmarks that place queries only at the end… See the full description on the dataset page: https://huggingface.co/datasets/TreeAILab/Multi-turn_Long-context_Benchmark_for_LLMs.multiturn_chat_0.8M
Multiturn Chat 0.8M
内容
包含约80万条由BELLE项目生成的用户与助手的多轮对话。
注意:此数据集是由ChatGPT产生的,未经过严格校验,内容可能包含错误。使用过程中请注意这一点。
instruction中包含多轮对话的上文内容,以Human:和Assistant:区分,output中包含当前助手角色的回答。
样例
{
"instruction":… See the full description on the dataset page: https://huggingface.co/datasets/BelleGroup/multiturn_chat_0.8M.multiturn-chatMulti-Turn-Instruct
Multi-Turn-Instruct Dataset
Dataset introduced in paper "Can Language Models Follow Multiple Turns of Entangled Instructions?"
📌 Overview
This repository contains the dataset, evaluation code, and benchmarks for the Multi-Turn-Instructdataset introduced in:
Can Language Models Follow Multiple Turns of Entangled Instructions?Chi Han, Xin Liu, Haodong Wang, Shiyang Li, Jingfeng Yang, Haoming Jiang, Zhengyang Wang, Qingyu Yin, Liang Qiu, Changlong Yu, Yifan Gao, Zheng… See the full description on the dataset page: https://huggingface.co/datasets/Glaciohound/Multi-Turn-Instruct.ultrachat_speech_multiTurnsDolci-Think-SFT-7B-multiturntool-use-multiturn-reasoningMulti-Turn-Insurance-Underwriting
Dataset Card for Multi-Turn-Insurance-Underwriting
Dataset Summary
This dataset includes sample traces and associated metadata from multi-turn interactions between a commercial underwriter and AI assistant. We built the system in langgraph with model context protocol and ReAct agents. In each sample, the underwriter has a specific task to solve related to a recent application for insurance by a small business. We created a diverse sample dataset covering 6 distinct types… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Multi-Turn-Insurance-Underwriting.GUI-AIMA-multiturnknowchat-multi-turn-dialogues
KnowChat: Multi-Turn Human-LLM Dialogues on Knowledge Tasks
KnowChat is a dataset of 705 multi-turn human-LLM conversations collected to validate the KnowSim user simulation framework. It pairs each conversation with pre/post knowledge assessments, self-reported survey ratings, and participant background information, enabling research on information calibration -- how well LLM assistants tailor responses to users with different knowledge levels.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/yjlee36/knowchat-multi-turn-dialogues.spoken-multiturn-sft
Spoken Multi-turn SFT Japanese
Japanese spoken multi-turn SFT dataset generated from kanhatakeyama/AutoMultiTurnByCalm3-22B using CosyVoice2 TTS.
Dataset Description
This dataset contains Japanese multi-turn SFT (Supervised Fine-Tuning) data with spoken questions.
q1: First question (text + audio)
a1: First answer (text only)
q2: Follow-up question (text + audio)
a2: Second answer (text only)
Samples
ID
Q1
Q1 Audio
A1
Q2
Q2 Audio
A2
0
鉄は強磁性体ですか?… See the full description on the dataset page: https://huggingface.co/datasets/Atotti/spoken-multiturn-sft.sql-multiturn-training-dataset-combinedQwen3.8-27B-multi-turn-agent-sft
Qwen3.8-multi-turn-agent-sft
Hello everyone! We are UkisAI, a small research lab from Europe.
We created this dataset based on the OpenThoughts-Agent-v1-SFT dataset. The traces in this release were generated with Qwen3.8-27B in FP16 using the Terminus-2 agentic harness.
This dataset contains approximately 15,200 agent traces covering terminal, coding, and software-engineering tasks, including tasks from nl2bash and InferredBugs.
Please feel free to try it, share feedback, report… See the full description on the dataset page: https://huggingface.co/datasets/ukisai/Qwen3.8-27B-multi-turn-agent-sft.bfcl_v3_multi_turn_baseNemotron-RL-Instruction-Following-MultiTurnChat-v1
Dataset Description:
The MultiChallenge Dataset is a rigorous benchmark designed to improve large language models in complex multi-turn conversations by explicitly targeting inference memory, instruction retention, version editing, and self-coherence. It employs a unique "model breaking" methodology where tasks are tested against advanced models (Nemotron-Nano-V2 and Qwen3-235B-A22B-Thinking-2507) to expose failure modes. A sample is only accepted into the dataset if the task is… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-MultiTurnChat-v1.tool-calls-multiturntherapy-conversations-multiturn
Combined Dr. AURA Therapy Conversations Dataset
This dataset is 100% AI-generated for research and educational purposes only. It is not intended to provide medical, psychological, or therapeutic advice. Always consult a qualified healthcare professional or doctor for any mental health concerns or medical issues. AI-generated content may contain errors or inaccuracies.
warning ⚠️: the LENGTH of each conversation MAY VARY (eg. 7 or 8 or 9 or 10 etc. turns in each row). And SOME END… See the full description on the dataset page: https://huggingface.co/datasets/Abc7347/therapy-conversations-multiturn.chatalpaca-multiturn-enrichedchatalpaca-multiturn-enriched-2.1multiturn_chat_0.8M
