datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BrowseCompMultiTurn-Chat-MT-Bench-Judge
SEA-MT-Bench-Judge
SEA-MT-Bench-Judge expands on the original SEA-MTBench through the use of a criteria-based evaluation framework. We use GPT-OSS-120B as the judge model.
The prompts are based on MT-Bench and was manually translated by native speakers. Furthermore, some prompts were modified to be more suitable for the criteria-based judgments.
Supported Tasks and Leaderboards
SEA-MT-Bench-Judge is designed for evaluating chat or instruction-tuned large language… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/MultiTurn-Chat-MT-Bench-Judge.spoken-multiturn-sft
Spoken Multi-turn SFT Japanese
Japanese spoken multi-turn SFT dataset generated from kanhatakeyama/AutoMultiTurnByCalm3-22B using CosyVoice2 TTS.
Dataset Description
This dataset contains Japanese multi-turn SFT (Supervised Fine-Tuning) data with spoken questions.
q1: First question (text + audio)
a1: First answer (text only)
q2: Follow-up question (text + audio)
a2: Second answer (text only)
Samples
ID
Q1
Q1 Audio
A1
Q2
Q2 Audio
A2
0
鉄は強磁性体ですか?… See the full description on the dataset page: https://huggingface.co/datasets/Atotti/spoken-multiturn-sft.ultrachat_speech_multiTurnstool-use-multiturn-reasoningDolci-Think-SFT-7B-multiturnMulti-Turn-Insurance-Underwriting
Dataset Card for Multi-Turn-Insurance-Underwriting
Dataset Summary
This dataset includes sample traces and associated metadata from multi-turn interactions between a commercial underwriter and AI assistant. We built the system in langgraph with model context protocol and ReAct agents. In each sample, the underwriter has a specific task to solve related to a recent application for insurance by a small business. We created a diverse sample dataset covering 6 distinct types… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Multi-Turn-Insurance-Underwriting.sql-multiturn-training-dataset-combinedwildchat-en-multiturnmultiturn_ks
khursanirevo/multiturn_ks
Dataset Description
Multiturn dialogue dataset with speaker-separated stereo audio and multi-language transcripts from 139 YouTube videos.
Features
Audio: Stereo audio with speaker separation (speaker 0 = left channel, speaker 1 = right channel)
Segments: Speaker turn-level annotations with timestamps for English and Malay
Multi-language: Transcripts in 9 languages (en, ms, zh-Hans, zh-Hant, ru, id, ar, ja, ko)
Video ID: YouTube video… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/multiturn_ks.tool-calls-multiturnmultiturn_processedSWE-Gym-Smalltraining-datasetcraft-multiturn-actions-split-nothinkprm800k_onpolicy_multiturn_rtg_prefix0.2_roll4_maxrev100multiturn_chat_0.8m-chinese-zhtw
Dataset Card for "multiturn_chat_0.8m-chinese-zhtw"
內容
包含約 80 萬條由 BELLE 專案所產生的 user 與 assistant 的多輪對話。
注意:此資料集是由 ChatGPT 產生的,未經嚴格校驗,內容可能包含錯誤。使用過程中請注意這一點。
限制和使用限制
我們要求開發者僅將我們開源的程式碼、資料、模型及後續衍生物用於研究目的,不得用於商業,以及其他會對社會帶來危害的用途。
由於數據是由ChatGPT產生的,未經嚴格驗證,在事實性和其他方面仍有一些不足之處。因此,在使用此資料集時,請務必注意甄別。
本資料集不代表任何一方的立場、利益或想法,無關任何團體的任何類型的主張。因使用本資料集帶來的任何損害、糾紛,本專案的開發者不承擔任何責任。
Multiturn Chat 0.8M
Contents
Includes approx. 0.8M Chinese multiturn dialogs between… See the full description on the dataset page: https://huggingface.co/datasets/benchang1110/multiturn_chat_0.8m-chinese-zhtw.text2CAD-multiturn-reasoningMalaysian-Multiturn-Chat-Assistant
Malaysian-Multiturn-Chat-Assistant
Generate synthetic multi-turn chat assistant with complex system prompt using mesolitica/Malaysian-Qwen2.5-72B-Instruct.
After that generate synthetic voice using mesolitica/Malaysian-Dia-1.6B also verified with Force Alignment to make sure the pronunciations almost correct.
A conversation must at least have 2 audio. We follow chat template from Qwen/Qwen2-Audio-7B-Instruct.
how to prepare the dataset
huggingface-cli download \… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Multiturn-Chat-Assistant.multi_turn_function_callingscanqa_images_64_336x224_672x448_multiturnmultiturn-512-UltraInteract_pair_diff_lenMulti-Turn-Insurance-Underwriting-Code-Gen
Dataset Card for Multi-Turn-Insurance-Underwriting-Code-Gen
This dataset is a variant of the Multi-Turn-Insurance-Underwriting dataset, in which models do not get access to any tools except a code interpreter and a pointer to the relevant file system.
This helps us analyze how well models explore their environments.
Environment Creation
This diagram shows the architecture of how we create the dataset, with assistant responses interleaved with questions, ending with a… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Multi-Turn-Insurance-Underwriting-Code-Gen.ultrainteract_multiturnultrainteract_multiturn_1_iter_processedramdom-to-fixed-multiturn-Calm3
自動生成したテキスト
Calm3で自動生成したマルチターンデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
health-qa-multiturnmultiturn-feedback
MultiTurn Feedback Dataset
Multi-turn conversation feedback dataset with sparse and dense annotations.
Dataset Description
This dataset contains human feedback annotations for paper "User Feedback in Human-LLM Dialogues:
A Lens to Understand Users But Noisy as a Learning Signal". It includes two evaluation subsets:
Sparse: 75 conversations from LMSYS-Chat-1M with sparse feedback
Dense: 74 conversations from LMSYS-Chat-1M + 34 WildChat with dense feedback
Labels… See the full description on the dataset page: https://huggingface.co/datasets/yuhan-nlp/multiturn-feedback.ultrainteract_multiturn-reward-ckp_2therapy-conversations-multiturn
Combined Dr. AURA Therapy Conversations Dataset
This dataset is 100% AI-generated for research and educational purposes only. It is not intended to provide medical, psychological, or therapeutic advice. Always consult a qualified healthcare professional or doctor for any mental health concerns or medical issues. AI-generated content may contain errors or inaccuracies.
warning ⚠️: the LENGTH of each conversation MAY VARY (eg. 7 or 8 or 9 or 10 etc. turns in each row). And SOME END… See the full description on the dataset page: https://huggingface.co/datasets/Abc7347/therapy-conversations-multiturn.
