datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GameQA-140K
[ICLR 2026] Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
🎊 News
[2026/07] 🔥Peking University and Kuaishou Kling Team evaluate their agentic visual reasoning method Beacon on our GameQA benchmark. Beacon learns when tools are truly needed (Mode Adaptiveness) and how tool use extends capability on hard problems (Tool Effect), and achieves the highest accuracy on GameQA among open-source models of the same scale… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/GameQA-140K.VideoThinkBench
[CVPR 2026] Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
🎊 News
[2026.02] 🔥🔥Our work has been accepted by CVPR 2026! 🎉🎉🎉
[2025.11] Our paper "Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm" has been released on arXiv! 📄 [Paper] On HuggingFace, it has achieved "#1 Paper of the Day"!
[2025.11] 🔥We release "minitest" of our VideoThinkBench, including 500… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/VideoThinkBench.Realtime-QA-100K
Realtime-QA-100K
📄 Tech Report |
💻 GitHub
Realtime-QA-100K is a 100K-sample realtime video question answering dataset
constructed from YouTube videos. Each sample contains a multimodal
conversation and frame timestamp metadata that aligns every <|video|> token in
the assistant text with one video frame timestamp.
Open-source training subset.
Realtime-QA-100K is the open-source subset of the real-time training data for
MOSS-Video-Preview… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/Realtime-QA-100K.Ultra-Innerthought
Ultra-Innerthought🤔
English | 中文
Introduction
Ultra-Innerthought is a bilingual (Chinese and English) open-domain SFT dataset in Innerthought format, containing 2,085,326 dialogues. Unlike current reasoning datasets that mainly focus on mathematics and coding domains, Ultra-Innerthought covers a broader range of fields and includes both Chinese and English languages. We used Deepseek V3 as the model for data synthesis.
Dataset Format
{
"id":… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/Ultra-Innerthought.
