datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GUI-AIMA-multiturncrag-mm-multi-turn-public
CRAG-MM: Comprehensive multi-modal, multi-turn RAG Benchmark
This repository contains the CRAG-MM dataset, a high-quality conversational benchmark for multimodal assistants. The dataset features conversations about images with varied complexity levels, designed to evaluate AI systems' visual understanding and conversational abilities.
CRAG-MM is a visual question-answering benchmark that focuses on factual questions, offering a unique collection of image and question-answering sets… See the full description on the dataset page: https://huggingface.co/datasets/crag-mm-2025/crag-mm-multi-turn-public.Multi-Turn
TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents
Official codebase for TCOD, a temporal curriculum framework for on-policy distillation that stabilizes knowledge transfer from teacher to student agents in multi-turn interactive environments.
🔥 News
[2026-07] Our paper is accepted by COLM 2026!
[2026-06] ✍️ New blog post out: on-policy distillation pitfalls — sharing the lessons and pitfalls behind our… See the full description on the dataset page: https://huggingface.co/datasets/SeanWang0027/Multi-Turn.scanqa_images_64_336x224_672x448_multiturnwapcar.my-multiturnmulti-turn
InterSyn: A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
This dataset card accompanies the paper
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text GenerationYukang Feng, Jianwen Sun, Chuanhao Li, Zizhen Li, Jiaxin Ai, Fanrui Zhang, Yifan Chang, Sizhuo Zhou, Shenglin Zhang, Yu Dai, Kaipeng Zhang (2025)
🧠 Introduction
TL;DR InterSyn is a high-quality dataset for instruction‑following, interleaved image–text… See the full description on the dataset page: https://huggingface.co/datasets/finyorko/multi-turn.jp1924_multiturn_kvqaThis dataset is llava style multiturn vlm dataset from jp1924/KoreanImageCaptioningDataset.
crag-mm-multi-turn-debug-public
CRAG-MM: Comprehensive multi-modal, multi-turn RAG Benchmark
This repository contains the CRAG-MM dataset, a high-quality conversational benchmark for multimodal assistants. The dataset features conversations about images with varied complexity levels, designed to evaluate AI systems' visual understanding and conversational abilities.
CRAG-MM is a visual question-answering benchmark that focuses on factual questions, offering a unique collection of image and question-answering sets… See the full description on the dataset page: https://huggingface.co/datasets/crag-mm-2025/crag-mm-multi-turn-debug-public.motomalaysia.com-multiturnresepichenom.com-multiturnmultiturn_gamebakllava-multiturn-ecommercemultiturn_game_tit_geminiMultiturn-JPJ-Test-PrepMulti-turn conversation generated using Mistral-V4 on JPJ-Test-Prep questions.
Multi-turn-editingMultiTurn-testmnist_multiturn_sftThe dataset is created using the following script: https://gist.github.com/vermouth1992/e339c22f7591f2f9d8557df228c438e4
game_multiturn_less5_singleturnmultiturn_game_titpi-multiturn-no-redundantBV175_multiturncrag-mm-multi-turns-imagesBV175_multiturn_chatMulti_Turn_eval_filteredMulti_Turn_com
