datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OmniAction
RoboOmni: Proactive Robot Manipulation in Omni-modal Context
📖 arXiv Paper (Accepted to ICLR 2026 🎉) |
🌐 Website |
🤗 Model |
🤗 Dataset |
🛠️ Github |
Recent advances in Multimodal Large Language Models (MLLMs) have driven rapid progress in Vision–Language–Action (VLA) models for robotic manipulation. Although effective in many scenarios, current approaches largely rely on explicit instructions, whereas in real-world interactions, humans rarely issue… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/OmniAction.ToolACE
ToolACE
ToolACE is an automatic agentic pipeline designed to generate Accurate, Complex, and divErse tool-learning data.
ToolACE leverages a novel self-evolution synthesis process to curate a comprehensive API pool of 26,507 diverse APIs.
Dialogs are further generated through the interplay among multiple agents, guided by a formalized thinking process.
To ensure data accuracy, we implement a dual-layer verification system combining rule-based and model-based checks.
More details… See the full description on the dataset page: https://huggingface.co/datasets/Team-ACE/ToolACE.pipeline-cctv-analyticsPMC
Data collected from PMC
Only CC-BY, CC-BY-SA licenses are included.
For all records, check the jsonl files in the data folder
OmniAction-LIBERO
RoboOmni: Proactive Robot Manipulation in Omni-modal Context
📖 arXiv Paper (Accepted to ICLR 2026 🎉) |
🌐 Website |
🤗 Model |
🤗 Dataset |
🛠️ Github |
Recent advances in Multimodal Large Language Models (MLLMs) have driven rapid progress in Vision–Language–Action (VLA) models for robotic manipulation. Although effective in many scenarios, current approaches largely rely on explicit instructions, whereas in real-world interactions, humans rarely issue… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/OmniAction-LIBERO.babilong
BABILong (100 samples) : a long-context needle-in-a-haystack benchmark for LLMs
Preprint is on arXiv and code for LLM evaluation is available on GitHub.
BABILong Leaderboard with top-performing long-context models.
bAbI + Books = BABILong
BABILong is a novel generative benchmark for evaluating the performance of NLP models in
processing arbitrarily long documents with distributed facts.
It contains 11 configs, corresponding to different sequence lengths in tokens:… See the full description on the dataset page: https://huggingface.co/datasets/RMT-team/babilong.babilong-1k-samples
BABILong (1000 samples) : a long-context needle-in-a-haystack benchmark for LLMs
Preprint is on arXiv and code for LLM evaluation is available on GitHub.
BABILong Leaderboard with top-performing long-context models.
bAbI + Books = BABILong
BABILong is a novel generative benchmark for evaluating the performance of NLP models in
processing arbitrarily long documents with distributed facts.
It contains 9 configs, corresponding to different sequence lengths in tokens: 0k… See the full description on the dataset page: https://huggingface.co/datasets/RMT-team/babilong-1k-samples.BeyondSWE-harbor
BeyondSWE-harbor
This repository provides the harbor version of the BeyondSWE benchmark, containing the full task instances in a directory-based (harbor) format, where each instance is stored as an independent folder.
📌 For benchmark definition, data format and detailed evaluation results, please refer to: 🤗 Main Dataset on HuggingFace
🗂️ Data Structure
beyondswe/
├── {instance_id}/
│ ├── environment/
│ ├── solution/
│ ├── tests/
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/AweAI-Team/BeyondSWE-harbor.hleScale-SWE
Immersion in the GitHub Universe: Scaling Coding Agents to Mastery
🔥 Highlights
Source from 6M+ pull requests and 23000+ repositories.
Cover 5200 Repositories.
100k high-quality instances.
71k trajectories from DeepSeek v3.2 with 3.5B token.
Strong performance: 64% in SWE-bench-Verified trained from Qwen3-30A3B-Instruct.
📣 News
2026-02-26 🚀 We released a portion of our data on Hugging Face. This release includes 20,000 SWE task… See the full description on the dataset page: https://huggingface.co/datasets/AweAI-Team/Scale-SWE.Skill-Evol-Bench
SkillEvolBench Dataset
SkillEvolBench is a diagnostic benchmark for testing whether LLM agents can convert episodic task experience into reusable procedural skills. It accompanies the paper SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills.
This Hugging Face dataset page hosts the benchmark assets used by the paper's skill-evolution protocol: role-instantiated task directories, verification assets, and curated seed skills. The Dataset… See the full description on the dataset page: https://huggingface.co/datasets/SkillEvolBench-Team/Skill-Evol-Bench.Pretrain-DatasetThis is the Pretraining Dataset for PLM.
Due to the upload limit, we split the original dataset into parts that smaller than 50GB. We provide the merge and split scripts under scripts folder.
YuLan-Mini-Datasets
YuLan-Mini Datasets
🔥 Updated (April 11, 2025): For a clearer presentation of the information, see the table at this link: link.
This datasets contains:
Classified data using python-edu-scorer and fineweb-edu-classifier
Synthesized data (math, code, instruction, ...)
Retrieved data using math, code, and reasoninig-classifier
Notice
Since we have used BPE-Dropout, in order to ensure accuracy, the data we uploaded is tokenized.… See the full description on the dataset page: https://huggingface.co/datasets/yulan-team/YuLan-Mini-Datasets.moss-003-sft-data
moss-003-sft-data
** More information: MOSS Paper**
Conversation Without Plugins
Categories
Category
# samples
Brainstorming
99,162
Complex Instruction
95,574
Code
198,079
Role Playing
246,375
Writing
341,087
Harmless
74,573
Others
19,701
Total
1,074,551
Others contains two categories: Continue(9,839) and Switching(9,862).The Continue category refers to instances in a conversation where the user asks the system to continue… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-003-sft-data.UltraFineWeb-filteredGameQA-140K
[ICLR 2026] Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
🎊 News
[2026/07] 🔥Peking University and Kuaishou Kling Team evaluate their agentic visual reasoning method Beacon on our GameQA benchmark. Beacon learns when tools are truly needed (Mode Adaptiveness) and how tool use extends capability on hard problems (Tool Effect), and achieves the highest accuracy on GameQA among open-source models of the same scale… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/GameQA-140K.SWE-bench-Science
SWE-bench Science
SWE-bench Science evaluates coding agents on software-engineering tasks drawn from scientific-computing repositories. The release contains 119 tasks across 20 scientific domains, with isolated environments and separate programmatic verifiers.
GitHub release repository: OpenMOSS/SWE-bench-Science
Runtime images: Docker Hub, pinned by immutable linux/amd64 digests
Evaluation framework: Pier, compatible with Harbor task format
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/SWE-bench-Science.YuLan-Mini-Text-Datasets
News
[2025.04.11] Add dataset mixture: link.
[2025.03.30] Text datasets upload finished.
This is text dataset.
这是文本格式的数据集。
Since we have used BPE-Dropout, in order to ensure accuracy, you can find the tokenized dataset here.
由于我们使用了BPE-Dropout,为了保证准确性,你可以在这里找到分词后的数据。
For more information, please refer to our datasets details and preprocess details.
Contributing
We welcome any form of contribution, including feedback on model bad cases, feature suggestions, and example… See the full description on the dataset page: https://huggingface.co/datasets/yulan-team/YuLan-Mini-Text-Datasets.rendered-wikipedia-english
Dataset Card for Team-PIXEL/rendered-wikipedia-english
Dataset Summary
This dataset contains the full English Wikipedia from February 1, 2018, rendered into images of 16x8464 resolution.
The original text dataset was built from a Wikipedia dump. Each example in the original text dataset contained the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.). Each rendered example contains a subset of one full article.… See the full description on the dataset page: https://huggingface.co/datasets/Team-PIXEL/rendered-wikipedia-english.team_ozaki_submit1hle-extractminecraft-vpt-mp4
Minecraft VPT MP4 ArrayRecords
Minecraft gameplay clips and their aligned VPT-style actions, packaged as
sharded ArrayRecord files. This is the
raw-video dataset used by the Minecraft data path in dreamer4-jax-private.
Download
Install the Hugging Face CLI and download the repository to a local directory:
pip install -U huggingface_hub
hf download reactor-team/minecraft-vpt-mp4 \
--repo-type dataset \
--local-dir /path/to/mp4-arrayrecords
The downloaded… See the full description on the dataset page: https://huggingface.co/datasets/reactor-team/minecraft-vpt-mp4.OmniAction-LIBERO-evalCalibForge
CalibForge
CalibForge is a collection of 5,431 executable and verifiable terminal-agent tasks constructed with adversarial solver calibration.
CalibForge uses solver behavior during task construction in two complementary ways:
Multi-solver calibration retains tasks that expose disagreement across a heterogeneous solver pool.
Contrastive solver calibration targets a designated strong-pass and weak-fail capability relation.
Dataset structure
CalibForge/… See the full description on the dataset page: https://huggingface.co/datasets/AweAI-Team/CalibForge.AM-DeepSeek-R1-Distilled-1.4MFor more open-source datasets, models, and methodologies, please visit our GitHub repository.
AM-DeepSeek-R1-Distilled-1.4M is a large-scale general reasoning task dataset composed of
high-quality and challenging reasoning problems. These problems are collected from numerous
open-source datasets, semantically deduplicated, and cleaned to eliminate test set contamination.
All responses in the dataset are distilled from the reasoning model (mostly DeepSeek-R1) and have undergone
rigorous… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-R1-Distilled-1.4M.AM-DeepSeek-Distilled-40MFor more open-source datasets, models, and methodologies, please visit our GitHub repository and paper: DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training.
Due to certain constraints, we are only able to open-source a subset of the complete dataset.
Model Training Performance based on our complete dataset
On AIME 2024, our 72B model achieved a score of 79.2 using only supervised fine-tuning (SFT). The 32B model reached 75.8 and… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-Distilled-40M.unclickbait-synthetic-27b-trajectories
Unclickbait Synthetic 27B Trajectories
Synthetic trajectory dataset generated by using rich structured JSON prompts and validated by two-stage judging pipeline.
Contents
: Full generated trajectories (current snapshot: 48,623 records out of 152,369 pristine event candidates).
: 30 benchmark test samples audited end-to-end through the 122B two-stage judge (Stage 1 integrity gate + Stage 2 4D scoring).
gemma-2b-suite-explanationsSwissCrop25
SwissCrop25
A national benchmark dataset for operational crop mapping in Switzerland, providing Sentinel-2
time series, daily temperature data, and parcel-level crop type labels across seven growing
seasons (2019–2025).
Introduced in: SwissCrop25: A National Multi-Year Benchmark for Operational Crop Mapping
(TerraBytes II Workshop, ECCV 2026) — [Paper] [Code] [Team]
Highlights
Nationwide coverage of Switzerland (41,285 km²)
Seven growing seasons (2019–2025)
73… See the full description on the dataset page: https://huggingface.co/datasets/EOA-team/SwissCrop25.BeyondSWE
BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?
BeyondSWE is a comprehensive benchmark that evaluates code agents along two key dimensions — resolution scope and knowledge scope — moving beyond single-repo bug fixing into the real-world deep waters of software engineering.
✨ Highlights
500 real-world instances across 246 GitHub repositories, spanning four distinct task settings
Two-dimensional evaluation: simultaneously… See the full description on the dataset page: https://huggingface.co/datasets/AweAI-Team/BeyondSWE.
