datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OmniAction
RoboOmni: Proactive Robot Manipulation in Omni-modal Context
📖 arXiv Paper (Accepted to ICLR 2026 🎉) |
🌐 Website |
🤗 Model |
🤗 Dataset |
🛠️ Github |
Recent advances in Multimodal Large Language Models (MLLMs) have driven rapid progress in Vision–Language–Action (VLA) models for robotic manipulation. Although effective in many scenarios, current approaches largely rely on explicit instructions, whereas in real-world interactions, humans rarely issue… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/OmniAction.OmniAction-LIBERO
RoboOmni: Proactive Robot Manipulation in Omni-modal Context
📖 arXiv Paper (Accepted to ICLR 2026 🎉) |
🌐 Website |
🤗 Model |
🤗 Dataset |
🛠️ Github |
Recent advances in Multimodal Large Language Models (MLLMs) have driven rapid progress in Vision–Language–Action (VLA) models for robotic manipulation. Although effective in many scenarios, current approaches largely rely on explicit instructions, whereas in real-world interactions, humans rarely issue… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/OmniAction-LIBERO.moss-003-sft-data
moss-003-sft-data
** More information: MOSS Paper**
Conversation Without Plugins
Categories
Category
# samples
Brainstorming
99,162
Complex Instruction
95,574
Code
198,079
Role Playing
246,375
Writing
341,087
Harmless
74,573
Others
19,701
Total
1,074,551
Others contains two categories: Continue(9,839) and Switching(9,862).The Continue category refers to instances in a conversation where the user asks the system to continue… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-003-sft-data.GameQA-140K
🎊 News
[2026/07] 🔥Peking University and Kuaishou Kling Team evaluate their agentic visual reasoning method Beacon on our GameQA benchmark. Beacon learns when tools are truly needed (Mode Adaptiveness) and how tool use extends capability on hard problems (Tool Effect), and achieves the highest accuracy on GameQA among open-source models of the same scale, significantly outperforming its Qwen3-VL-8B-Instruct base.
[2026/07] 🔥Peking University and WeChat AI use our Game-RL data… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/GameQA-140K.SWE-bench-Science
SWE-bench Science
SWE-bench Science evaluates coding agents on software-engineering tasks drawn from scientific-computing repositories. The release contains 119 tasks across 20 scientific domains, with isolated environments and separate programmatic verifiers.
GitHub release repository: OpenMOSS/SWE-bench-Science
Runtime images: Docker Hub, pinned by immutable linux/amd64 digests
Evaluation framework: Pier, compatible with Harbor task format
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/SWE-bench-Science.OmniAction-LIBERO-evalmoss-002-sft-data
Dataset Card for "moss-002-sft-data"
Dataset Summary
An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data.
Data Splits
name
# samples
en_helpfulness.json
419049
en_honesty.json
112580… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-002-sft-data.MHA2MLA-corpus-smollmAnyInstruct
Dataset details
Dataset type
AnyInstruct is a dataset comprised of 108k multimodal instruction-following data integrates multiple modalities—text, speech, images, and music—in an interleaved manner.
Data Construction
We first synthesize textual multimodal dialogues using GPT-4, and then generate images, music, and voices using DALL-E 3, MusicGen, and the Azure Text-to-Speech API, respectively. The voice component comprises 39 different timbres, with the speech… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/AnyInstruct.FutureOmni
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
Predicting the future requires listening as well as seeing.
📖 Dataset Summary
Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio–visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective understanding.
FutureOmni is the first benchmark designed… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/FutureOmni.AnyInstruct-resolution-1024
File Restoration and Extraction Guide
File Structure
Root directory: Contains Part 1 split files
part2/ directory: Contains Part 2 split files
Instructions
Step 1: File Restoration
Due to size limitations, the original file has been split. To restore the complete file:
cat images_1024.part_* > images_1024.tar
Step 2: Extraction
To extract the contents:
tar -xvf images_1024.tar
Important Notes
For Part 1 images: Execute… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/AnyInstruct-resolution-1024.VideoThinkBench
[CVPR 2026] Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
🎊 News
[2026.02] 🔥🔥Our work has been accepted by CVPR 2026! 🎉🎉🎉
[2025.11] Our paper "Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm" has been released on arXiv! 📄 [Paper] On HuggingFace, it has achieved "#1 Paper of the Day"!
[2025.11] 🔥We release "minitest" of our VideoThinkBench, including 500… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/VideoThinkBench.robotwin2.0-lerobot-v3.0
RoboTwin 2.0 LeRobot v3.0
This dataset packages RoboTwin 2.0 bimanual-manipulation demonstrations in LeRobot v3.0 format for EasyWAM training. It contains synchronized high-view and dual-wrist RGB video, bimanual joint state, actions, and natural-language task metadata.
Dataset Summary
Episodes
Frames
Videos
Recording rate
27,500
6,075,103
82,500
50 FPS
Structure
robotwin2.0-lerobot-v3.0/
├── data/ # frame-level… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/robotwin2.0-lerobot-v3.0.SciJudgeBench
SciJudgeBench Dataset
Training and evaluation data for scientific paper citation prediction, from the paper AI Can Learn Scientific Taste.
Given two academic papers (title, abstract, publication date), the task is to predict which paper has a higher citation count.
Resources: Project page, GitHub repository, SciJudge-4B-2605, and SciJudge-30B-2605.
Dataset Splits
Split
Examples
Description
train
720,341
Training preference pairs from arXiv papers
test… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/SciJudgeBench.character-llm-data
Character-LLM: A Trainable Agent for Role-Playing
This is the training datasets for Character-LLM, which contains nine characters experience data used to train Character-LLMs.
To download the dataset, please run the following code with Python, and you can find the downloaded data in /path/to/local_dir.
from huggingface_hub import snapshot_download
snapshot_download(
local_dir_use_symlinks=True,
repo_type="dataset",
repo_id="fnlp/character-llm-data"… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/character-llm-data.Realtime-QA-100K
Realtime-QA-100K
📄 Tech Report |
💻 GitHub
Realtime-QA-100K is a 100K-sample realtime video question answering dataset
constructed from YouTube videos. Each sample contains a multimodal
conversation and frame timestamp metadata that aligns every <|video|> token in
the assistant text with one video frame timestamp.
Open-source training subset.
Realtime-QA-100K is the open-source subset of the real-time training data for
MOSS-Video-Preview… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/Realtime-QA-100K.FRoM-W1-Datasets
FRoM-W1: Towards General Humanoid Whole-Body Control with Language Instructions
The Humanoid Intelligence Team from FudanNLP and OpenMOSS
Introduction
Humanoid robots are capable of performing various actions such as greeting, dancing and even backflipping. However, these motions are often hard-coded or specifically trained, which limits their versatility. In this work, we present FRoM-W1[^1]… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/FRoM-W1-Datasets.ABC-Bench
ABC-Bench
💻 Code |
📑 Paper |
📝 Blog
📖 Overview
ABC-Bench is a benchmark for Agentic Backend Coding. It evaluates whether code agents can explore real repositories, edit code, configure environments, deploy containerized services, and pass external end-to-end API tests (HTTP-based integration tests) across realistic backend stacks.
📊 Benchmark Composition
🚀 Why ABC-Bench?
End-to-End Lifecycle: repository exploration → code… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/ABC-Bench.GameQA-5KIn this repository, we specifically provide the 5k training samples from the complete GameQA-140K dataset used in our work for GRPO training of the models.
Refer to our paper for details. And our code for training and evaluation is at https://github.com/tongjingqi/Code2Logic.
Code2Logic: Game-Code-Driven Data Synthesis for Enhancing VLMs General Reasoning
This is the first work, to the best of our knowledge, that leverages game code to synthesize multimodal reasoning data for… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/GameQA-5K.libero-lerobot-v3.0
LIBERO LeRobot v3.0
This dataset packages the four standard LIBERO demonstration suites in LeRobot v3.0 format for robot-policy and world-action-model training. It contains synchronized RGB observations, robot state, actions, timestamps, task indices, and natural-language task metadata.
Dataset Summary
Suite
Episodes
Frames
Tasks
Videos
LIBERO-Spatial
434
53,229
10
868
LIBERO-Object
457
67,309
10
914
LIBERO-Goal
433
52,895
10
866
LIBERO-10
388
104… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/libero-lerobot-v3.0.hh-rlhf-strength-cleaned
Dataset Card for hh-rlhf-strength-cleaned
Other Language Versions: English, 中文.
Dataset Description
In the paper titled "Secrets of RLHF in Large Language Models Part II: Reward Modeling" we measured the preference strength of each preference pair in the hh-rlhf dataset through model ensemble and annotated the valid set with GPT-4. In this repository, we provide:
Metadata of preference strength for both the training and valid sets.
GPT-4 annotations on the valid set.
We… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/hh-rlhf-strength-cleaned.SpeechInstructUltra-Innerthought
Ultra-Innerthought🤔
English | 中文
Introduction
Ultra-Innerthought is a bilingual (Chinese and English) open-domain SFT dataset in Innerthought format, containing 2,085,326 dialogues. Unlike current reasoning datasets that mainly focus on mathematics and coding domains, Ultra-Innerthought covers a broader range of fields and includes both Chinese and English languages. We used Deepseek V3 as the model for data synthesis.
Dataset Format
{
"id":… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/Ultra-Innerthought.MHA2MLA-corpus-qwen1.5VehicleWorld
📚 Introduction
VehicleWorld is the first comprehensive multi-device environment for intelligent vehicle interaction that accurately models the complex, interconnected systems in modern cockpits. This environment enables precise evaluation of agent behaviors by providing real-time state information during execution. This dataset is specifically designed to evaluate the capabilities of Large Language Models (LLMs) as in-car intelligent assistants in understanding and executing… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/VehicleWorld.MHA2MLA-corpus-qwen1_5MHA2MLA-corpus-qwen2MHA2MLA-corpus-llama2GameQA-textHere we provide a pure-text version of GameQA, encompassing some appropriate games. (See https://github.com/tongjingqi/Code2Logic/issues/2)
Code2Logic: Game-Code-Driven Data Synthesis for Enhancing VLMs General Reasoning
This is the first work, to the best of our knowledge, that leverages game code to synthesize multimodal reasoning data for training VLMs. Furthermore, when trained with a GRPO strategy solely on GameQA (synthesized via our proposed Code2Logic approach), multiple… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/GameQA-text.case2code-dataTraining Dataset for "Case2Code: Scalable Synthetic Data for Code Generation"
Usage:
from datasets import load_dataset
dataset = load_dataset("fnlp/case2code-data", split="train")
