datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eMCR
eMCR: A Benchmark for Multi-Condition Product Retrieval in Chinese E-Commerce
This is the official dataset and evaluation code for the paper "eMCR: A Benchmark for Multi-Condition Product Retrieval in Chinese E-Commerce".
Overview
Product search increasingly involves queries that combine multiple requirements — product attributes, brands, prices, exclusions, and visual descriptions. Existing retrieval benchmarks provide limited support for diagnosing which… See the full description on the dataset page: https://huggingface.co/datasets/alibabagroup/eMCR.OmniThoughtV_Raw_1.8M
Dataset Introduction
OmniThoughtV is a large-scale multimodal long-chain-of-thought dataset distilled from the FineVision dataset using Alibaba Cloud's AI platform (PAI) distillation toolkit, EasyDistill. This dataset establishes a transparent and reproducible data distillation pipeline, enabling efficient construction of multimodal reasoning chains of thought. Fine-tuning smaller models with this dataset effectively endows them with stronger reasoning capabilities and enhances… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-pai/OmniThoughtV_Raw_1.8M.IndustryBench-MIPU
IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products
Multi-Image Industrial Product Understanding Benchmark — evaluating MLLMs on structured attribute extraction from real-world industrial product images.
Industrial product specifications are scattered across multiple heterogeneous images — specification tables, nameplates, technical drawings. IndustryBench-MIPU tests whether MLLMs can reliably recover them through four… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench-MIPU.OmniThoughtV_Filter_0.5M
Dataset Introduction
OmniThoughtV is a large-scale multimodal long-chain-of-thought dataset distilled from the FineVision dataset using Alibaba Cloud's AI platform (PAI) distillation toolkit, EasyDistill. This dataset establishes a transparent and reproducible data distillation pipeline, enabling efficient construction of multimodal reasoning chains of thought. Fine-tuning smaller models with this dataset effectively endows them with stronger reasoning capabilities and enhances… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-pai/OmniThoughtV_Filter_0.5M.ClinFusion-Eval-Data
🏥 ClinFusion-Eval-Data
The Holistic Evaluation Suite for Vision-Centric Medical Multimodal LLMs
ClinFusion-Eval-Data is the unified evaluation corpus used to benchmark the ClinFusion model series (ClinFusion-8B, ClinFusion-32B). It packages 211,810 evaluation records spanning 22 public medical benchmarks into a single, consistently-formatted suite, together with 509 GiB of the underlying 2D images and native 3D CT volumes they refer to.
The goal is reproducibility:… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/ClinFusion-Eval-Data.alibaba-ctf
Alibaba CTF Benchmark
Alibaba CTF Benchmark is a CTF benchmark designed to measure the frontier of agent work on Capture The Flag security challenges. It consists of 87 high-quality tasks curated from the 2023–2026 AlibabaCTF (formerly AliyunCTF) competition series, covering five core categories: Web (25), Pwn (19), Misc (14), Reverse (16), and Crypto (13). During the curation process, LLM-based challenges were excluded due to their additional credential requirements and test… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-AAIG/alibaba-ctf.terminal-bench-pro
Terminal-Bench Pro
Overview
Terminal-Bench Pro is a systematic extension of the original Terminal-Bench, designed to address key limitations in existing terminal-agent benchmarks.
400 tasks (200 public + 200 private) across 8 domains: data processing, games, debugging, system admin, scientific computing, software engineering, ML, and security
Expert-designed tasks derived from real-world scenarios and GitHub issues
High test coverage with ~28.3 test cases per… See the full description on the dataset page: https://huggingface.co/datasets/alibabagroup/terminal-bench-pro.OmniDoc-TokenBench
OmniDoc-TokenBench
📄 Overview
Introduction
We propose OmniDoc-TokenBench in Qwen-Image-VAE-2.0, a curated benchmark specifically designed to evaluate VAE reconstruction on text-rich document images. It contains ~3K samples spanning nine categories (book, slides, color textbook, exam paper, academic paper, magazine, financial report, newspaper, note) in both English and Chinese, alongside an evaluation toolkit supporting PSNR, SSIM, LPIPS, FID, and… See the full description on the dataset page: https://huggingface.co/datasets/alibabagroup/OmniDoc-TokenBench.Superior-Reasoning-SFT-gpt-oss-120b-Logprob
Superior-Reasoning-SFT-gpt-oss-120b-Logprob
🚀 Overview
This dataset contains the token-level log-probabilities generated by the teacher model (gpt-oss-120b) for the reasoning samples in the main Superior-Reasoning-SFT-gpt-oss-120b Dataset.
🔗 Relationship to Main Dataset
This dataset is a companion to the main Superior-Reasoning-SFT-gpt-oss-120bdataset. Records are linked via a unique sample_uuid.
Main Dataset: Contains the text (prompts… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b-Logprob.StreamGuardBench
Kelp: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk Detection
💻 GitHub
💡 Dataset Overview
StreamGuardBench is the first benchmark specifically designed for evaluating streaming guardrails. StreamGuardBench prompts ten widely used open-source LMs—comprising five text-only and five vision-language models—and annotates every generated response with harm labels, therefore enabling accurate measurement of streaming guardrail effectiveness in… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-AAIG/StreamGuardBench.Superior-Reasoning-SFT-gpt-oss-120b
Superior-Reasoning-SFT-gpt-oss-120b
📣 News
Our dataset ranked #1 on the Hugging Face Datasets Trending leaderboard from January 20 to January 30.
🚀 Overview
The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b.MMA-SafetyBenchRynnBrain-Bench
RynnBrain-Bench
Introduction
We introduce RynnBrain-Bench, a high-dimensional evaluation suite designed to holistically benchmark the cognition and localization capabilities of embodied understanding models in complex household environments.
Advancing beyond existing benchmarks, RynnBrain-Bench features a unique emphasis on fine-grained understanding and precise spatiotemporal localization within episodic video sequences.
RynnBrain-Bench systematically… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/RynnBrain-Bench.CMGUI
CMGUI Dataset
📄 Project |
🌐 Model in Hugging Face |
🌐 Model in ModelScope
English |
简体中文
CMGUI (Chinese Mobile GUI) is a large-scale, high-quality dataset constructed for developing GUI agents on Chinese mobile applications. The dataset contains 18k episodes (i.e., trajectories) with 98k steps collected from more than 50 real-world Chinese mobile apps, covering diverse functional domains such as e-commerce (e.g., Taobao, Pinduoduo), social media (e.g., Rednote… See the full description on the dataset page: https://huggingface.co/datasets/alibabagroup/CMGUI.VoxCeleb2-mixA modified version of the VoxCeleb2 Dataset. Original data can be downloaded here.
This dataset is used for Audio-visual speaker extraction conditioned on face recordings in the reentry paper, which the code can be found here (ClearVoice repo) or here (Paper repo).
Usage
cat orig* > orig.tar
tar -xvf orig.tar
cat audio_clean* > audio_clean.tar
tar -xvf audio_clean.tar
aacr-bench
Dataset for Running AACR-Bench
English | 简体中文
This is a test set designed for automated code review reflection models, primarily aiming to evaluate the extent to which a model can intercept low-quality review comments. The dataset contains 2,145 code review comments, consisting of 1,505 expert-verified correct comments and 640 incorrect comments.
This data is part of the AACR-Bench project and is provided by the Alibaba Aone team.
Data Sample
Each sample in the… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Aone/aacr-bench.ViViDWebShaper
WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization
Github: https://github.com/Alibaba-NLP/WebAgent
Paper: https://arxiv.org/pdf/2507.15061
TLTR
WebShaper is a synthesized training dataset for information-seeking (IS) task. It is based on our proposed task formalization of IS, and synthesized by our Expander Agent. WebShaper would cover a broader range of task forms, reasoning structure, and diversified knowledge.
Description… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-NLP/WebShaper.UVRB
🌐 Universal Video Retrieval Benchmark (UVRB)
The first comprehensive benchmark for universal video retrievalEvaluate your model across 16 datasets, 3 query types, and 6 capability dimensions — not just accuracy, but why it succeeds or fails.
UVRB is a comprehensive evaluation suite designed to diagnose and quantify a video embedding model’s true generalization ability — beyond narrow text-to-video tasks. It exposes critical gaps in spatial reasoning, temporal dynamics… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-NLP/UVRB.OmniThought
OmniThought: A Large-Scale Chain-of-Thought Dataset for Advancing Large Reasoning Models
Overview
The rise of Large Reasoning Models (LRMs) has revolutionized Natural Language Processing (NLP), enabling breakthroughs in complex tasks like mathematical problem-solving and code generation. These models rely on Chain-of-Thought (CoT) processes to mimic human-like reasoning. However, progress in LRMs is limited by the scarcity of high-quality, large-scale CoT… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-pai/OmniThought.SecRespond
SecRespond
💻 GitHub |
🤖 ModelScope |
📄 Paper
Introduction
SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response evaluates whether an AI agent can investigate a compromised host after an attack has already succeeded.
For each cyber range, the responder receives a frozen forensic disk snapshot together with synthetic host-security-product outputs, then produces an evidence-backed incident… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-NLP/SecRespond.XGuard-Train-Open-200K
YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models
🤗 HuggingFace |
🤖 ModelScope |
📄 Paper
🐬 Introduction
XGuard-Train-Open-200K is an open-source subset of the training corpus developed for the YuFeng-XGuard-Reason guardrail model series. YuFeng-XGuard-Reason is engineered to accurately identify security risks in user requests, model responses, and general text, while… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-AAIG/XGuard-Train-Open-200K.EcomBench
EcomBench: Where Intelligent Agents Conquer Commerce Realms
🚀 Benchmark Overview
EcomBench is a domain-specific, real-world evaluation framework designed to rigorously assess the capabilities of AI agents in delivering practical support for the complex, ever-evolving demands of e-commerce.
We believe that truly capable AI agents will fundamentally transform how we interact with commerce. E-commerce represents one of the world's most significant economic sectors, with… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-NLP/EcomBench.InterVBench
Video Drift Evaluation (vde.py)
This repository contains a single entry point, vde.py, that computes Video Drift Error (VDE) scores for every .mp4 file inside a target directory. VDE provides a simple way to monitor how quality-related metrics drift across chunks of the same video. The script already supports several metric backends (clarity, motion, aesthetic, dynamic, subject, background) via the vbench tooling.
Environment Setup
Install the project… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/InterVBench.RynnEC-Bench
RynnEC-Bench
RynnEC-Bench evaluates fine-grained embodied understanding models from the perspectives of object cognition and spatial cognition in open-world scenario. The benchmark includes 507 video clips captured in real household scenarios.
Model
Overall Mean
Object Properties
Seg. DR
Seg. SR
Object Mean
Ego. His.
Ego. Pres.
Ego. Fut.
World Size
World Dis.
World PR
Spatial Mean
GPT-4o
28.3
41.1
---
---
33.9
13.4
22.8
6.0
24.3
16.7
36.1
22.2
GPT-4.1
33.5… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/RynnEC-Bench.alibaba_dataset
alibaba_dataset
Dataset Information
This dataset contains multiple parts of a large archive that need to be combined.
Reconstruction Instructions
To reconstruct the original dataset file from the parts:
Download all parts (alibaba_dataset_part_*)
Use the cat command to join them:cat alibaba_dataset_part_* > alibaba_dataset.tar.xz
Extract the tarball:tar -xf alibaba_dataset.tar.xz
IndustryBench
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
💻Github | 📝Paper
IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese.
Overview
Dimension
Details
Total questions
2,049
Languages
Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.SimpleQA-Bench
SimpleQA-Bench
Tags: factuality, EN, ZH, short-form-answer, human-label
Copyright: © 2024 alibaba-pai
Source.OpenAI's SimpleQA: Blog & Paper / Data & simple-evals ProjectOpenStellarTeam's Chinese-SimpleQA: Blog & Paper, Data@HF
Factuality is a complicated topic because it is hard to measure—evaluating the factuality of any given arbitrary claim is challenging, and language models can generate long completions that contain dozens of factual claims. In SimpleQA, we will focus on… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-pai/SimpleQA-Bench.SynSearch-Data
SynSearch-Data
SynSearch-Data is an environment-aligned Search Agent training dataset produced through task synthesis and Solver-in-the-Loop verification. This first release contains 5,000 high-quality multi-hop ReAct trajectories covering 4,182 audited tasks.
Data format
Each JSONL row contains:
messages: the complete ReAct conversation, including tool calls and tool responses;
metadata.task_id / run_idx / trajectory_rank: trajectory provenance;
metadata.seed:… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-pai/SynSearch-Data.arxiv_embeddings_Alibaba-NLP_gte-base-en-v1.5
