datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
terminal-bench-pro
Terminal-Bench Pro
Overview
Terminal-Bench Pro is a systematic extension of the original Terminal-Bench, designed to address key limitations in existing terminal-agent benchmarks.
400 tasks (200 public + 200 private) across 8 domains: data processing, games, debugging, system admin, scientific computing, software engineering, ML, and security
Expert-designed tasks derived from real-world scenarios and GitHub issues
High test coverage with ~28.3 test cases per… See the full description on the dataset page: https://huggingface.co/datasets/alibabagroup/terminal-bench-pro.Superior-Reasoning-SFT-gpt-oss-120b-Logprob
Superior-Reasoning-SFT-gpt-oss-120b-Logprob
🚀 Overview
This dataset contains the token-level log-probabilities generated by the teacher model (gpt-oss-120b) for the reasoning samples in the main Superior-Reasoning-SFT-gpt-oss-120b Dataset.
🔗 Relationship to Main Dataset
This dataset is a companion to the main Superior-Reasoning-SFT-gpt-oss-120bdataset. Records are linked via a unique sample_uuid.
Main Dataset: Contains the text (prompts… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b-Logprob.Superior-Reasoning-SFT-gpt-oss-120b
Superior-Reasoning-SFT-gpt-oss-120b
📣 News
Our dataset ranked #1 on the Hugging Face Datasets Trending leaderboard from January 20 to January 30.
🚀 Overview
The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b.aacr-bench
Dataset for Running AACR-Bench
English | 简体中文
This is a test set designed for automated code review reflection models, primarily aiming to evaluate the extent to which a model can intercept low-quality review comments. The dataset contains 2,145 code review comments, consisting of 1,505 expert-verified correct comments and 640 incorrect comments.
This data is part of the AACR-Bench project and is provided by the Alibaba Aone team.
Data Sample
Each sample in the… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Aone/aacr-bench.SecRespond
SecRespond
💻 GitHub |
🤖 ModelScope |
📄 Paper
Introduction
SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response evaluates whether an AI agent can investigate a compromised host after an attack has already succeeded.
For each cyber range, the responder receives a frozen forensic disk snapshot together with synthetic host-security-product outputs, then produces an evidence-backed incident… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-NLP/SecRespond.XGuard-Train-Open-200K
YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models
🤗 HuggingFace |
🤖 ModelScope |
📄 Paper
🐬 Introduction
XGuard-Train-Open-200K is an open-source subset of the training corpus developed for the YuFeng-XGuard-Reason guardrail model series. YuFeng-XGuard-Reason is engineered to accurately identify security risks in user requests, model responses, and general text, while… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-AAIG/XGuard-Train-Open-200K.IndustryBench
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
💻Github | 📝Paper
IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese.
Overview
Dimension
Details
Total questions
2,049
Languages
Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.Open-DeepResearch
Open-DeepResearch
Project Page | Paper | Code
This directory contains the RL Training Set and the Test Set for the Open-DeepResearch domain.
Overview
In the Open-DeepResearch domain, the agent is required to assist users in conducting multi-turn search, reading, synthesis, and generation to produce an open-ended answer. This domain focuses on complex information retrieval and synthesis tasks.
Dataset
Statistics
Split
Samples
Description
RL… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-NLP/Open-DeepResearch.SKYLENAGE-GameCodeGym
V-GameGym: Visual Game Generation for Code Large Language Models
Abstract
Code large language models have demonstrated remarkable capabilities in programming tasks, yet current benchmarks primarily focus on single modality rather than visual game development. Most existing code-related benchmarks evaluate syntax correctness and execution accuracy, overlooking critical game-specific metrics such as playability, visual aesthetics, and user engagement that are… See the full description on the dataset page: https://huggingface.co/datasets/alibabagroup/SKYLENAGE-GameCodeGym.Open-Travel
Open-Travel
Project Page | Paper | Code
This directory contains the RL Training Set and the Test Set (categorized by subtask) for the Open-Travel domain.
Overview
In the Open-Travel domain, the agent is required to help users accomplish itinerary planning subtasks. These tasks emphasize multi-constraint reasoning, multi-tool coordination, and personalized preferences intertwined with user-specific constraints (e.g., budget limits, time windows, traveling parties, and… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-NLP/Open-Travel.
