datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MultiTurn-Chat-MT-Bench-Judge
SEA-MT-Bench-Judge
SEA-MT-Bench-Judge expands on the original SEA-MTBench through the use of a criteria-based evaluation framework. We use GPT-OSS-120B as the judge model.
The prompts are based on MT-Bench and was manually translated by native speakers. Furthermore, some prompts were modified to be more suitable for the criteria-based judgments.
Supported Tasks and Leaderboards
SEA-MT-Bench-Judge is designed for evaluating chat or instruction-tuned large language… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/MultiTurn-Chat-MT-Bench-Judge.Multi-Turn-Insurance-Underwriting
Dataset Card for Multi-Turn-Insurance-Underwriting
Dataset Summary
This dataset includes sample traces and associated metadata from multi-turn interactions between a commercial underwriter and AI assistant. We built the system in langgraph with model context protocol and ReAct agents. In each sample, the underwriter has a specific task to solve related to a recent application for insurance by a small business. We created a diverse sample dataset covering 6 distinct types… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Multi-Turn-Insurance-Underwriting.craft-multiturn-actions-split-nothinkmultiturn_processedprm800k_onpolicy_multiturn_rtg_prefix0.2_roll4_maxrev100ultrainteract_multiturnMulti-Turn-Insurance-Underwriting-Code-Gen
Dataset Card for Multi-Turn-Insurance-Underwriting-Code-Gen
This dataset is a variant of the Multi-Turn-Insurance-Underwriting dataset, in which models do not get access to any tools except a code interpreter and a pointer to the relevant file system.
This helps us analyze how well models explore their environments.
Environment Creation
This diagram shows the architecture of how we create the dataset, with assistant responses interleaved with questions, ending with a… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Multi-Turn-Insurance-Underwriting-Code-Gen.prm800k_onpolicy_multiturn_cummrew_prefix0.1_roll4_maxrev100ultrainteract_multiturn_1_iter_processedultrainteract_multiturn-reward-ckp_2craft-multiturn-actions-splitultrainteract_multiturn_sampled_h_from_sampled_len_ckp_4ultrainteract_multiturn_sampled_h_from_sampled_len_ckp_2ultrainteract_multiturn-reward-ckp_1prm800k_onpolicy_multiturn_rtgshape_prefix0.2_roll4_maxrev100multiturn_6_harvardultrainteract_multiturn_sampled_h_from_sampled_lenultrainteract_multiturn_1_iter_sampled_h_from_sampled_len_ckp_1multiturn_1_2_h_harvardultrainteract_multiturn_1_iter_processed_ckp_rwogmath5_onpolicy_multiturn_seprew_prefix0.1_roll4_maxrev100prm800k_onpolicy_multiturn_cumm_rew_prefix0.2_roll4_maxrev100mixed-instruction-speech-multiturn-noiseultrainteract_multiturn_1_iter_sampled_h_from_sampled_len_ckp_0ogmath5_onpolicy_multiturn_seprew_prefix0.2_roll4_maxrev100ultrainteract_multiturn_sampled_h_from_sampled_len_ckp_1multiturn_1_2_harvardcad_multiturn_all_parts_transformedprm800k_onpolicy_multiturn_seprew_prefix0.1_roll4_maxrev100ultrainteract_multiturn_1_iter_processed_harvard
