datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MultiTurn-Chat-MT-Bench-Judge
SEA-MT-Bench-Judge
SEA-MT-Bench-Judge expands on the original SEA-MTBench through the use of a criteria-based evaluation framework. We use GPT-OSS-120B as the judge model.
The prompts are based on MT-Bench and was manually translated by native speakers. Furthermore, some prompts were modified to be more suitable for the criteria-based judgments.
Supported Tasks and Leaderboards
SEA-MT-Bench-Judge is designed for evaluating chat or instruction-tuned large language… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/MultiTurn-Chat-MT-Bench-Judge.Multi-Turn-Insurance-Underwriting
Dataset Card for Multi-Turn-Insurance-Underwriting
Dataset Summary
This dataset includes sample traces and associated metadata from multi-turn interactions between a commercial underwriter and AI assistant. We built the system in langgraph with model context protocol and ReAct agents. In each sample, the underwriter has a specific task to solve related to a recent application for insurance by a small business. We created a diverse sample dataset covering 6 distinct types… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Multi-Turn-Insurance-Underwriting.Nemotron-RL-Instruction-Following-MultiTurnChat-v1
Dataset Description:
The MultiChallenge Dataset is a rigorous benchmark designed to improve large language models in complex multi-turn conversations by explicitly targeting inference memory, instruction retention, version editing, and self-coherence. It employs a unique "model breaking" methodology where tasks are tested against advanced models (Nemotron-Nano-V2 and Qwen3-235B-A22B-Thinking-2507) to expose failure modes. A sample is only accepted into the dataset if the task is… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-MultiTurnChat-v1.multiturn_processedprm800k_onpolicy_multiturn_rtg_prefix0.2_roll4_maxrev100craft-multiturn-actions-split-nothinkmultiturn_1_2_harvardMulti-Turn-Insurance-Underwriting-Code-Gen
Dataset Card for Multi-Turn-Insurance-Underwriting-Code-Gen
This dataset is a variant of the Multi-Turn-Insurance-Underwriting dataset, in which models do not get access to any tools except a code interpreter and a pointer to the relevant file system.
This helps us analyze how well models explore their environments.
Environment Creation
This diagram shows the architecture of how we create the dataset, with assistant responses interleaved with questions, ending with a… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Multi-Turn-Insurance-Underwriting-Code-Gen.collabllm-multiturn-math-hardultrainteract_multiturn_1_iter_processedultrainteract_multiturnultrainteract_multiturn-reward-ckp_2craft-multiturn-actions-splitmultiturn_5_harvardultrainteract_multiturn_sampled_h_from_sampled_len_ckp_4multiturn_1_2multiturn_1_4_harvardprm800k_onpolicy_multiturn_rtgshape_prefix0.2_roll4_maxrev100multiturn_5multiturn_1_2_hultrainteract_multiturn_sampled_h_from_sampled_lenmultiturn_1_3multiturn_6_harvardprm800k_onpolicy_multiturn_cumm_rew_prefix0.2_roll4_maxrev100DiscoverLLM-multiturn-preferences
DiscoverLLM: Multi-turn Preference Dataset
Multi-turn dialogue data with scored candidate completions, produced by best-of-N
synthesis over the DiscoverLLM user simulator
(paper · project page).
Each example is a single turn of a simulated user–assistant conversation with one of
several candidate assistant responses and an associated reward score, intended for
offline DPO / GRPO / reward-model training.
Configs
Config
Rows
Task
creative_writing
3,052… See the full description on the dataset page: https://huggingface.co/datasets/kixlab/DiscoverLLM-multiturn-preferences.multiturn_1_2_h_harvardmixed-instruction-speech-multiturn-noiseNemotron-RL-Instruction-Following-MultiTurnChat-v1-prompt-only
Nemotron-RL-Instruction-Following-MultiTurnChat-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-MultiTurnChat-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-MultiTurnChat-v1-prompt-only.ultrainteract_multiturn-reward-ckp_1ultrainteract_multiturn_1_iter_processed_ckp_rw
