datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ToolACE
ToolACE
ToolACE is an automatic agentic pipeline designed to generate Accurate, Complex, and divErse tool-learning data.
ToolACE leverages a novel self-evolution synthesis process to curate a comprehensive API pool of 26,507 diverse APIs.
Dialogs are further generated through the interplay among multiple agents, guided by a formalized thinking process.
To ensure data accuracy, we implement a dual-layer verification system combining rule-based and model-based checks.
More details… See the full description on the dataset page: https://huggingface.co/datasets/Team-ACE/ToolACE.PMC
Data collected from PMC
Only CC-BY, CC-BY-SA licenses are included.
For all records, check the jsonl files in the data folder
Teachers
仓库信息
电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/Teachers/discussions 提出。
此仓库存储导师著作:https://huggingface.co/datasets/VoiceOfML/Teachers/tree/main 。
请使用:https://voiceofml-search.hf.space/Teachers 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/Teachers )。
可使用:https://voiceofml-search.hf.space/Teachers?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/Teachers?wide=1 )。
你可以仅下载指针(只有文件名的信息)
If you want to clone without large files - just their… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/Teachers.babilong
BABILong (100 samples) : a long-context needle-in-a-haystack benchmark for LLMs
Preprint is on arXiv and code for LLM evaluation is available on GitHub.
BABILong Leaderboard with top-performing long-context models.
bAbI + Books = BABILong
BABILong is a novel generative benchmark for evaluating the performance of NLP models in
processing arbitrarily long documents with distributed facts.
It contains 11 configs, corresponding to different sequence lengths in tokens:… See the full description on the dataset page: https://huggingface.co/datasets/RMT-team/babilong.hleBeyondSWE-harbor
BeyondSWE-harbor
This repository provides the harbor version of the BeyondSWE benchmark, containing the full task instances in a directory-based (harbor) format, where each instance is stored as an independent folder.
📌 For benchmark definition, data format and detailed evaluation results, please refer to: 🤗 Main Dataset on HuggingFace
🗂️ Data Structure
beyondswe/
├── {instance_id}/
│ ├── environment/
│ ├── solution/
│ ├── tests/
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/AweAI-Team/BeyondSWE-harbor.babilong-1k-samples
BABILong (1000 samples) : a long-context needle-in-a-haystack benchmark for LLMs
Preprint is on arXiv and code for LLM evaluation is available on GitHub.
BABILong Leaderboard with top-performing long-context models.
bAbI + Books = BABILong
BABILong is a novel generative benchmark for evaluating the performance of NLP models in
processing arbitrarily long documents with distributed facts.
It contains 9 configs, corresponding to different sequence lengths in tokens: 0k… See the full description on the dataset page: https://huggingface.co/datasets/RMT-team/babilong-1k-samples.Scale-SWE
Immersion in the GitHub Universe: Scaling Coding Agents to Mastery
🔥 Highlights
Source from 6M+ pull requests and 23000+ repositories.
Cover 5200 Repositories.
100k high-quality instances.
71k trajectories from DeepSeek v3.2 with 3.5B token.
Strong performance: 64% in SWE-bench-Verified trained from Qwen3-30A3B-Instruct.
📣 News
2026-02-26 🚀 We released a portion of our data on Hugging Face. This release includes 20,000 SWE task… See the full description on the dataset page: https://huggingface.co/datasets/AweAI-Team/Scale-SWE.GameQA-140K
[ICLR 2026] Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
🎊 News
[2026/07] 🔥Peking University and Kuaishou Kling Team evaluate their agentic visual reasoning method Beacon on our GameQA benchmark. Beacon learns when tools are truly needed (Mode Adaptiveness) and how tool use extends capability on hard problems (Tool Effect), and achieves the highest accuracy on GameQA among open-source models of the same scale… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/GameQA-140K.SWE-bench-Science
SWE-bench Science
SWE-bench Science evaluates coding agents on software-engineering tasks drawn from scientific-computing repositories. The release contains 119 tasks across 20 scientific domains, with isolated environments and separate programmatic verifiers.
GitHub release repository: OpenMOSS/SWE-bench-Science
Runtime images: Docker Hub, pinned by immutable linux/amd64 digests
Evaluation framework: Pier, compatible with Harbor task format
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/SWE-bench-Science.rendered-wikipedia-english
Dataset Card for Team-PIXEL/rendered-wikipedia-english
Dataset Summary
This dataset contains the full English Wikipedia from February 1, 2018, rendered into images of 16x8464 resolution.
The original text dataset was built from a Wikipedia dump. Each example in the original text dataset contained the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.). Each rendered example contains a subset of one full article.… See the full description on the dataset page: https://huggingface.co/datasets/Team-PIXEL/rendered-wikipedia-english.unclickbait-synthetic-27b-trajectories
Unclickbait Synthetic 27B Trajectories
Synthetic trajectory dataset generated by using rich structured JSON prompts and validated by two-stage judging pipeline.
Contents
: Full generated trajectories (current snapshot: 48,623 records out of 152,369 pristine event candidates).
: 30 benchmark test samples audited end-to-end through the 122B two-stage judge (Stage 1 integrity gate + Stage 2 4D scoring).
hle-extractAM-DeepSeek-Distilled-40MFor more open-source datasets, models, and methodologies, please visit our GitHub repository and paper: DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training.
Due to certain constraints, we are only able to open-source a subset of the complete dataset.
Model Training Performance based on our complete dataset
On AIME 2024, our 72B model achieved a score of 79.2 using only supervised fine-tuning (SFT). The 32B model reached 75.8 and… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-Distilled-40M.CalibForge
CalibForge
CalibForge is a collection of 5,431 executable and verifiable terminal-agent tasks constructed with adversarial solver calibration.
CalibForge uses solver behavior during task construction in two complementary ways:
Multi-solver calibration retains tasks that expose disagreement across a heterogeneous solver pool.
Contrastive solver calibration targets a designated strong-pass and weak-fail capability relation.
Dataset structure
CalibForge/… See the full description on the dataset page: https://huggingface.co/datasets/AweAI-Team/CalibForge.komorebi-painter-teachers
Komorebi painter teacher benchmark references
This dataset stores the 40 fixed COCO128 reference JPEGs used by the painter teacher benchmark. The source metadata and per-image provenance are in refs.json; SHA-256 hashes are in reference-manifest.json. The source records do not establish per-image licensing, so this dataset card does not assert a blanket image license.
Generated benchmark results will be uploaded under runs/ with their own manifests after each run.
gemma-2b-suite-explanationsteam_ozaki_submit1hle_labeled-v1.0
HLE Labeled Dataset
このデータセットは「Humanity’s Last Exam」ベンチマーク用のデータセットの
categoryにsubcategoryを追加したものです。subcategoryの分類ラベルはqwen/qwen3-235b-a22bで生成しています。
モデルのカテゴリ別の評価に利用するのが目的です。
データ構造
変更点はもともとのHLEにsubcategoryフィールド追加したのみです。
id: レコードのユニークID
question: 問題文(文字列)
answer: 正解
answer_type: "exactMatch" などの解答形式
rationale: 解答手順・根拠
category: 大分類(例: "Math")
subcategory: 小分類のリスト(例: ["Math/Number Theory","Math/Discrete Mathematics"])
image, image_preview, rationale_image:… See the full description on the dataset page: https://huggingface.co/datasets/LLMcompe-Team-Watanabe/hle_labeled-v1.0.BeyondSWE
BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?
BeyondSWE is a comprehensive benchmark that evaluates code agents along two key dimensions — resolution scope and knowledge scope — moving beyond single-repo bug fixing into the real-world deep waters of software engineering.
✨ Highlights
500 real-world instances across 246 GitHub repositories, spanning four distinct task settings
Two-dimensional evaluation: simultaneously… See the full description on the dataset page: https://huggingface.co/datasets/AweAI-Team/BeyondSWE.taardis-27b-teacher-bundleSwissCrop25
SwissCrop25
A national benchmark dataset for operational crop mapping in Switzerland, providing Sentinel-2
time series, daily temperature data, and parcel-level crop type labels across seven growing
seasons (2019–2025).
Introduced in: SwissCrop25: A National Multi-Year Benchmark for Operational Crop Mapping
(TerraBytes II Workshop, ECCV 2026) — [Paper] [Code] [Team]
Highlights
Nationwide coverage of Switzerland (41,285 km²)
Seven growing seasons (2019–2025)
73… See the full description on the dataset page: https://huggingface.co/datasets/EOA-team/SwissCrop25.YuLan-Mini-Text-Datasets
News
[2025.04.11] Add dataset mixture: link.
[2025.03.30] Text datasets upload finished.
This is text dataset.
这是文本格式的数据集。
Since we have used BPE-Dropout, in order to ensure accuracy, you can find the tokenized dataset here.
由于我们使用了BPE-Dropout,为了保证准确性,你可以在这里找到分词后的数据。
For more information, please refer to our datasets details and preprocess details.
Contributing
We welcome any form of contribution, including feedback on model bad cases, feature suggestions, and example… See the full description on the dataset page: https://huggingface.co/datasets/yulan-team/YuLan-Mini-Text-Datasets.babilong-train-5k-samples
BABILong (5k train samples) : a long-context needle-in-a-haystack benchmark for LLMs
Preprint is on arXiv
bAbI + Books = BABILong
BABILong is a novel generative benchmark for evaluating the performance of NLP models in
processing arbitrarily long documents with distributed facts.
It contains 10 configs, each corresponding to its bAbI task. Each config has spltis corresponding to different sequence lengths in tokens: '4k', '32k', '128k', '256k', '512k', '1M'
Solving tasks… See the full description on the dataset page: https://huggingface.co/datasets/RMT-team/babilong-train-5k-samples.skm-tea-mini
SKM-TEA Sample Data
This dataset consists of a subset of scans from the SKM-TEA dataset. It can be used to build tutorials / demos with the SKM-TEA dataset.
To access to the full dataset, please follow instructions on Github.
NOTE: This dataset subset should not be used for reporting/publishing metrics. All metrics should be computed on the full SKM-TEA test split.
Details
This mini dataset (~30GB) consists of 2 training scans, 1 validation scan, and 1 test scan from… See the full description on the dataset page: https://huggingface.co/datasets/arjundd/skm-tea-mini.moss-002-sft-data
Dataset Card for "moss-002-sft-data"
Dataset Summary
An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data.
Data Splits
name
# samples
en_helpfulness.json
419049
en_honesty.json
112580… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-002-sft-data.MVU-Eval-Data
MVU-Eval Dataset
Paper | Code | Project Page
Dataset Description
The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark… See the full description on the dataset page: https://huggingface.co/datasets/MVU-Eval-Team/MVU-Eval-Data.S3E S3E: A Mulit-Robot Multimodal Dataset for Collaborative SLAM
[!TIP]
This is a project website of S3E dataset.
Feel free to open a
disccussion.
KNOWN ISSUES
[!IMPORTANT]
For experimental sequences in the laboratory, we capture only the start and end points due to constraints in Vicon system availability. Evaluation of these sequences is subsequently performed using only these two reference points.
oercommons-v1-optimized
OERCommons v1 Optimized
Authors: Junjie Wang and Yuhan SunHosted by: PIN TeamDataset: pin-team/oercommons-v1-optimized
OERCommons v1 Optimized is a provenance-preserving multimodal pretraining corpus built on the OERCommons subset of The Common Pile v0.1, which serves as its upstream data and licensing baseline. We extend it with full-page recovery, canonical Markdown, ordered image/PDF/link metadata, conservative corrections, and integrity evidence.
At a glance… See the full description on the dataset page: https://huggingface.co/datasets/pin-team/oercommons-v1-optimized.FutureOmni
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
Predicting the future requires listening as well as seeing.
📖 Dataset Summary
Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio–visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective understanding.
FutureOmni is the first benchmark designed… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/FutureOmni.
