CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MultiturnRL /BrowseComptext1K<n<10K0 likes3.5k downloads1y agoHugging Face02aisingapore /MultiTurn-Chat-MT-Bench-Judgegated SEA-MT-Bench-Judge SEA-MT-Bench-Judge expands on the original SEA-MTBench through the use of a criteria-based evaluation framework. We use GPT-OSS-120B as the judge model. The prompts are based on MT-Bench and was manually translated by native speakers. Furthermore, some prompts were modified to be more suitable for the criteria-based judgments. Supported Tasks and Leaderboards SEA-MT-Bench-Judge is designed for evaluating chat or instruction-tuned large language… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/MultiTurn-Chat-MT-Bench-Judge.tabularn<1K0 likes2k downloads2mo agoHugging Face03Dampfinchen /Creative_Writing_MultiturnUPDATE 2026: Stronger filtering using a very sophisticated filtering script and new data including a very small subset of https://huggingface.co/datasets/lemon07r/VellumK2T-Fiction-SFT-01 reasoning for thinking with a custom system prompt attached. This is suitable for both instruct non-thinking and thinking models, as I have added a system prompt for these few samples that use the tags <!think!> and </!think!> (without exclamation marks of course). This is a dataset merge of many, many high… See the full description on the dataset page: https://huggingface.co/datasets/Dampfinchen/Creative_Writing_Multiturn.text1K<n<10K36 likes2k downloads8mo agoHugging Face04BelleGroup /multiturn_chat_0.8M Multiturn Chat 0.8M 内容 包含约80万条由BELLE项目生成的用户与助手的多轮对话。 注意:此数据集是由ChatGPT产生的,未经过严格校验,内容可能包含错误。使用过程中请注意这一点。 instruction中包含多轮对话的上文内容,以Human:和Assistant:区分,output中包含当前助手角色的回答。 样例 { "instruction":… See the full description on the dataset page: https://huggingface.co/datasets/BelleGroup/multiturn_chat_0.8M.text100K<n<1M146 likes860 downloads3y agoHugging Face05agentlans /NousResearch-Hermes-3-Dataset-multiturn Hermes 3 Multiturn This is a filtered subset of NousResearch/Hermes-3-Dataset containing only multiturn conversations with more than three messages. Conversations with repetitive or trivial replies (for example, repeated "OK") have been excluded to improve quality. text10K<n<100K2 likes684 downloads1y agoHugging Face06Atotti /spoken-multiturn-sft Spoken Multi-turn SFT Japanese Japanese spoken multi-turn SFT dataset generated from kanhatakeyama/AutoMultiTurnByCalm3-22B using CosyVoice2 TTS. Dataset Description This dataset contains Japanese multi-turn SFT (Supervised Fine-Tuning) data with spoken questions. q1: First question (text + audio) a1: First answer (text only) q2: Follow-up question (text + audio) a2: Second answer (text only) Samples ID Q1 Q1 Audio A1 Q2 Q2 Audio A2 0 鉄は強磁性体ですか?… See the full description on the dataset page: https://huggingface.co/datasets/Atotti/spoken-multiturn-sft.audio10K<n<100K0 likes608 downloads9mo agoHugging Face07JerryAGENDD /ultrachat_speech_multiTurnstext10K<n<100K0 likes600 downloads2y agoHugging Face08interstellarninja /tool-use-multiturn-reasoningtextquestion-answering10K<n<100K40 likes592 downloads1y agoHugging Face09garipovroma /Dolci-Think-SFT-7B-multiturntext1M<n<10M0 likes585 downloads5mo agoHugging Face10snorkelai /Multi-Turn-Insurance-Underwriting Dataset Card for Multi-Turn-Insurance-Underwriting Dataset Summary This dataset includes sample traces and associated metadata from multi-turn interactions between a commercial underwriter and AI assistant. We built the system in langgraph with model context protocol and ReAct agents. In each sample, the underwriter has a specific task to solve related to a recent application for insurance by a small business. We created a diverse sample dataset covering 6 distinct types… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Multi-Turn-Insurance-Underwriting.tabularquestion-answeringn<1K37 likes510 downloads1y agoHugging Face11Anna4242 /sql-multiturn-training-dataset-combinedtext1M<n<10M0 likes439 downloads1y agoHugging Face12TreeAILab /Multi-turn_Long-context_Benchmark_for_LLMs LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues Arxiv: https://www.arxiv.org/abs/2507.13681 Huggingface: https://huggingface.co/papers/2507.13681 Introduction LoopServe Multi-Turn Dialogue Benchmark is a comprehensive evaluation dataset comprising multiple diverse datasets designed to assess large language model performance in realistic conversational scenarios. Unlike traditional benchmarks that place queries only at the end… See the full description on the dataset page: https://huggingface.co/datasets/TreeAILab/Multi-turn_Long-context_Benchmark_for_LLMs.textquestion-answering1K<n<10K0 likes428 downloads1y agoHugging Face13nvidia /Nemotron-RL-Instruction-Following-MultiTurnChat-v1 Dataset Description: The MultiChallenge Dataset is a rigorous benchmark designed to improve large language models in complex multi-turn conversations by explicitly targeting inference memory, instruction retention, version editing, and self-coherence. It employs a unique "model breaking" methodology where tasks are tested against advanced models (Nemotron-Nano-V2 and Qwen3-235B-A22B-Thinking-2507) to expose failure modes. A sample is only accepted into the dataset if the task is… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-MultiTurnChat-v1.tabular1K<n<10K4 likes392 downloads7mo agoHugging Face14smz8599 /GUI-AIMA-multiturnimage100K<n<1M1 likes390 downloads7mo agoHugging Face15andersonbcdefg /wildchat-en-multiturntext1M<n<10M0 likes376 downloads1y agoHugging Face16khursanirevo /multiturn_ks khursanirevo/multiturn_ks Dataset Description Multiturn dialogue dataset with speaker-separated stereo audio and multi-language transcripts from 139 YouTube videos. Features Audio: Stereo audio with speaker separation (speaker 0 = left channel, speaker 1 = right channel) Segments: Speaker turn-level annotations with timestamps for English and Malay Multi-language: Transcripts in 9 languages (en, ms, zh-Hans, zh-Hant, ru, id, ar, ja, ko) Video ID: YouTube video… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/multiturn_ks.audioautomatic-speech-recognition10K<n<100K0 likes306 downloads5mo agoHugging Face17interstellarninja /tool-calls-multiturntext1K<n<10K23 likes269 downloads3y agoHugging Face18Ttimofeyka /Creative-Writing-Multiturn-Cleaned-16ktext1K<n<10K0 likes261 downloads1y agoHugging Face19GitBag /multiturn_processedtabular10K<n<100K0 likes255 downloads2y agoHugging Face20MultiturnRL /SWE-Gym-Smalltext1K<n<10K0 likes237 downloads1y agoHugging Face21Capx /MultiTurnChat CapX Scientific QA Dataset The CapX Scientific QA Dataset is a comprehensive collection of data designed to assist in the development of AI-powered tools that support scientists across various disciplines. This dataset aims to bridge the gap between machine learning and the scientific community by providing reliable and transparent resources for training and evaluating question-answering systems. Features CapX Scientific QA Dataset The CapX Scientific QA… See the full description on the dataset page: https://huggingface.co/datasets/Capx/MultiTurnChat.text1K<n<10K2 likes236 downloads2y agoHugging Face22agentlans /multiturn-chattexttext-generation1M<n<10M6 likes219 downloads11mo agoHugging Face23afrilang-edu /multiturntext10K<n<100K0 likes192 downloads10mo agoHugging Face24SeeratAslam /multiturn_datatext1M<n<10M0 likes172 downloads2mo agoHugging Face25Trendyol /All-CVE-Chat-MultiTurn-1999-2025-Dataset CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/All-CVE-Chat-MultiTurn-1999-2025-Dataset.texttext-generation100K<n<1M32 likes166 downloads1y agoHugging Face26MultiturnRL /training-datasettext100K<n<1M0 likes150 downloads1y agoHugging Face27carl213 /multi-turn_jailbreak_attack_datasets Multi-Turn Jailbreak Attack Datasets Description This dataset was created to compare single-turn and multi-turn jailbreak attacks on large language models (LLMs). The primary goal is to take a single harmful prompt and distribute the harm over multiple turns, making each prompt appear harmless in isolation. This approach is compared against traditional single-turn attacks with the complete prompt to understand their relative impacts and failure modes. The key feature of… See the full description on the dataset page: https://huggingface.co/datasets/carl213/multi-turn_jailbreak_attack_datasets.text1K<n<10K1 likes149 downloads6mo agoHugging Face28SeanWang0027 /Multi-Turn TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents Official codebase for TCOD, a temporal curriculum framework for on-policy distillation that stabilizes knowledge transfer from teacher to student agents in multi-turn interactive environments. 🔥 News [2026-07] Our paper is accepted by COLM 2026! [2026-06] ✍️ New blog post out: on-policy distillation pitfalls — sharing the lessons and pitfalls behind our… See the full description on the dataset page: https://huggingface.co/datasets/SeanWang0027/Multi-Turn.imagen<1K0 likes139 downloads2mo agoHugging Face29windfromthenorth /craft-multiturn-actions-split-nothinktabular1M<n<10M0 likes138 downloads11mo agoHugging Face30DataPilot /Knowledge-QA-MultiTurn-Dataset Knowledge QA Multi-turn Dataset(知識質問データセット・マルチターン) 概要 本データセットは、Aratako/Synthetic-JP-Conversations-Magpie-Nemotron-4-10k から質問を抽出し、DeepSeek V3.2で整形・フォローアップ質問を生成、Kimi K2.5で回答を生成した 3ターンのマルチターン知識質問応答データセット です。Reasoning有効化により思考過程も最終データに含まれ、質問の難易度に応じてReasoning effortが動的に切り替わります。生成にはSDG-LOOMという合成データ生成パイプラインを用いました。(sdg-loom) データの説明 項目 内容 件数 約3,000件 形式 JSONL(1行1JSON) 言語 日本語 ターン数 3ターン(質問3 + 回答3) ソースデータセット… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/Knowledge-QA-MultiTurn-Dataset.text1K<n<10K2 likes135 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.