datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-RL-Instruction-Following-MultiTurnChat-v1
Dataset Description:
The MultiChallenge Dataset is a rigorous benchmark designed to improve large language models in complex multi-turn conversations by explicitly targeting inference memory, instruction retention, version editing, and self-coherence. It employs a unique "model breaking" methodology where tasks are tested against advanced models (Nemotron-Nano-V2 and Qwen3-235B-A22B-Thinking-2507) to expose failure modes. A sample is only accepted into the dataset if the task is… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-MultiTurnChat-v1.swe_multiturn-claude-trace-tokenized-qwen3.5
SWE-bench Multi-turn Trace for LLMServingSim2
Derived from
LMCache/Agentic-Traces,
filtered to the claude-sonnet-4-6 + swebench subset, then re-tokenized
with Qwen/Qwen3.5-122B-A10B
to drive an LLM serving simulator with realistic per-turn prefix-cache reuse
and strict turn-by-turn dependency enforcement.
110 sessions, 2,559 turns
Average input 20.2K tokens per turn (max 65.9K)
Average output 288.7 tokens per turn
Per-session prefix-share ratio: 0.939 (turn N input is ~94% the same… See the full description on the dataset page: https://huggingface.co/datasets/noddu/swe_multiturn-claude-trace-tokenized-qwen3.5.personalization-reddit-multiturn
personalization-reddit-multiturn
Multi-turn (question, preferred_answer, full_conversation) records mined
from Reddit. Companion to dipikakhullar/personalization-reddit: same
OP-thanks-reply heuristic for identifying the preferred answerer, but
this dataset additionally captures any contiguous back-and-forth between
the OP and that single answerer after the thanks.
A record is only emitted when there is at least one further turn beyond
the OP's thanks reply.
Splits… See the full description on the dataset page: https://huggingface.co/datasets/dipikakhullar/personalization-reddit-multiturn.multiturn-benchmark-datamultiturn-injection-detection
Multi-Turn Distributed Prompt Injection Detection Dataset
Dataset Description
27,180 synthetic multi-turn conversations (18,754 train / 3,296 val / 5,130 test) designed for training and evaluating temporal prompt injection detectors. Each conversation consists of 6-9 user turns with assistant responses.
Shared-Prefix Design
Every attack conversation is paired with a benign conversation that shares identical opening turns. A conversational prefix of 3-5 user… See the full description on the dataset page: https://huggingface.co/datasets/rockCO78/multiturn-injection-detection.
