multi-turn-conversation
WangchanThaiInstruct_Multi-turn_Conversation_Dataset
WangchanThaiInstruct Multi-turn Conversation Dataset
We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language.
Citation
Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633
or BibTeX
@dataset{thammaleelakul_2024_13132633,
author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.deepfabric-7k-medical-multi-turn-conversation
Medical Education Curriculum Dataset by Deepfabric
Dataset Description
This synthetic dataset contains 7,570 high-quality conversations focused on medical education curriculum design
and clinical training. The conversations simulate realistic discussions between medical curriculum
committee chairs, educators, and healthcare professionals designing comprehensive learning pathways.
It was produced using the Open Source Synthetic dataset generation tool, DeepFabric… See the full description on the dataset page: https://huggingface.co/datasets/nolabs/deepfabric-7k-medical-multi-turn-conversation.Multi-Turn-Conversational-SFTCreated by: DataCreator AI
Multi-Domain Multi-Turn Chat Conversations Dataset
A synthetic conversational dataset designed for LLM supervised fine-tuning and chatbot training.
The dataset contains multi-turn dialogues across multiple everyday domains such as travel, banking, health, programming, and customer interactions. Conversations are structured in OpenAI chat fine-tuning format, making the dataset directly usable in modern fine-tuning pipelines.
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/DataCreatorAI/Multi-Turn-Conversational-SFT.multi-turn-conversation-50000xTopics based on HuggingFaceTB/everyday-conversations-llama3.1-2k, expanded to 50000 examples. All converstations kept under 2000 tokens.
From source README:
##################
Everyday conversations for Smol LLMs finetunings
This dataset contains 2.2k multi-turn conversations generated by Llama-3.1-70B-Instruct. We ask the LLM to generate a simple multi-turn conversation, with 3-4 short exchanges, between a User and an AI Assistant about a certain topic.
The topics are chosen to be… See the full description on the dataset page: https://huggingface.co/datasets/kth8/multi-turn-conversation-50000x.Refactor-Dialogue-1.4k-Multi-turn-Refactoring-Conversations
Refactor-Dialogue-1.4k — Multi-turn Refactoring Conversations
Synthetic dataset for fine-tuning coding-focused LLMs on multi-turn
refactoring dialogues. Generated with
Dataset Generator —
an open-source pipeline for building high-quality training data.
Overview
1,414 multi-turn conversations across 3 refactoring categories. Each example
is a 4-message dialogue: user pastes real code → assistant refactors with a
short explanation → user follows up with a constraint or… See the full description on the dataset page: https://huggingface.co/datasets/AronDaron/Refactor-Dialogue-1.4k-Multi-turn-Refactoring-Conversations.Refactor-Dialogue-1.4k-Multi-turn-Refactoring-Conversations-Reasoning
Refactor-Dialogue-1.4k-Reasoning — Multi-turn Refactoring Conversations with <think> Reasoning
Synthetic dataset for fine-tuning reasoning-style coding LLMs on
multi-turn refactoring dialogues. Every assistant turn carries a
first-person <think>...</think> internal monologue before the actual
response — DeepSeek-R1 / Qwen3-thinking convention, broadest trainer
compatibility out of the box.
Generated with
Dataset Generator —
an open-source pipeline for building high-quality… See the full description on the dataset page: https://huggingface.co/datasets/AronDaron/Refactor-Dialogue-1.4k-Multi-turn-Refactoring-Conversations-Reasoning.
