CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alibabagroup /eMCR eMCR: A Benchmark for Multi-Condition Product Retrieval in Chinese E-Commerce This is the official dataset and evaluation code for the paper "eMCR: A Benchmark for Multi-Condition Product Retrieval in Chinese E-Commerce". Overview Product search increasingly involves queries that combine multiple requirements — product attributes, brands, prices, exclusions, and visual descriptions. Existing retrieval benchmarks provide limited support for diagnosing which… See the full description on the dataset page: https://huggingface.co/datasets/alibabagroup/eMCR.imagetext-retrieval10K<n<100K0 likes15k downloads28d agoHugging Face02alibaba-pai /OmniThoughtV_Raw_1.8M Dataset Introduction OmniThoughtV is a large-scale multimodal long-chain-of-thought dataset distilled from the FineVision dataset using Alibaba Cloud's AI platform (PAI) distillation toolkit, EasyDistill. This dataset establishes a transparent and reproducible data distillation pipeline, enabling efficient construction of multimodal reasoning chains of thought. Fine-tuning smaller models with this dataset effectively endows them with stronger reasoning capabilities and enhances… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-pai/OmniThoughtV_Raw_1.8M.text1M<n<10M1 likes9.7k downloads8mo agoHugging Face03alibaba-multimodal-industrial-ai /IndustryBench-MIPU IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products Multi-Image Industrial Product Understanding Benchmark — evaluating MLLMs on structured attribute extraction from real-world industrial product images. Industrial product specifications are scattered across multiple heterogeneous images — specification tables, nameplates, technical drawings. IndustryBench-MIPU tests whether MLLMs can reliably recover them through four… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench-MIPU.imageimage-to-text10K<n<100K7 likes7.7k downloads2mo agoHugging Face04alibaba-pai /OmniThoughtV_Filter_0.5M Dataset Introduction OmniThoughtV is a large-scale multimodal long-chain-of-thought dataset distilled from the FineVision dataset using Alibaba Cloud's AI platform (PAI) distillation toolkit, EasyDistill. This dataset establishes a transparent and reproducible data distillation pipeline, enabling efficient construction of multimodal reasoning chains of thought. Fine-tuning smaller models with this dataset effectively endows them with stronger reasoning capabilities and enhances… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-pai/OmniThoughtV_Filter_0.5M.1 likes5.7k downloads8mo agoHugging Face05Alibaba-DAMO-Academy /ClinFusion-Eval-Data 🏥 ClinFusion-Eval-Data The Holistic Evaluation Suite for Vision-Centric Medical Multimodal LLMs ClinFusion-Eval-Data is the unified evaluation corpus used to benchmark the ClinFusion model series (ClinFusion-8B, ClinFusion-32B). It packages 211,810 evaluation records spanning 22 public medical benchmarks into a single, consistently-formatted suite, together with 509 GiB of the underlying 2D images and native 3D CT volumes they refer to. The goal is reproducibility:… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/ClinFusion-Eval-Data.visual-question-answering100K<n<1M3 likes4.8k downloads1mo agoHugging Face06Alibaba-AAIG /alibaba-ctf Alibaba CTF Benchmark Alibaba CTF Benchmark is a CTF benchmark designed to measure the frontier of agent work on Capture The Flag security challenges. It consists of 87 high-quality tasks curated from the 2023–2026 AlibabaCTF (formerly AliyunCTF) competition series, covering five core categories: Web (25), Pwn (19), Misc (14), Reverse (16), and Crypto (13). During the curation process, LLM-based challenges were excluded due to their additional credential requirements and test… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-AAIG/alibaba-ctf.n<1K1 likes3.5k downloads1d agoHugging Face07alibabagroup /terminal-bench-pro Terminal-Bench Pro Overview Terminal-Bench Pro is a systematic extension of the original Terminal-Bench, designed to address key limitations in existing terminal-agent benchmarks. 400 tasks (200 public + 200 private) across 8 domains: data processing, games, debugging, system admin, scientific computing, software engineering, ML, and security Expert-designed tasks derived from real-world scenarios and GitHub issues High test coverage with ~28.3 test cases per… See the full description on the dataset page: https://huggingface.co/datasets/alibabagroup/terminal-bench-pro.texttext-generationn<1K5 likes2.9k downloads9mo agoHugging Face08alibabagroup /OmniDoc-TokenBench OmniDoc-TokenBench 📄 Overview Introduction We propose OmniDoc-TokenBench in Qwen-Image-VAE-2.0, a curated benchmark specifically designed to evaluate VAE reconstruction on text-rich document images. It contains ~3K samples spanning nine categories (book, slides, color textbook, exam paper, academic paper, magazine, financial report, newspaper, note) in both English and Chinese, alongside an evaluation toolkit supporting PSNR, SSIM, LPIPS, FID, and… See the full description on the dataset page: https://huggingface.co/datasets/alibabagroup/OmniDoc-TokenBench.imageimage-to-image1K<n<10K7 likes2.7k downloads4mo agoHugging Face09Alibaba-Apsara /Superior-Reasoning-SFT-gpt-oss-120b-Logprob Superior-Reasoning-SFT-gpt-oss-120b-Logprob           🚀 Overview This dataset contains the token-level log-probabilities generated by the teacher model (gpt-oss-120b) for the reasoning samples in the main Superior-Reasoning-SFT-gpt-oss-120b Dataset. 🔗 Relationship to Main Dataset This dataset is a companion to the main Superior-Reasoning-SFT-gpt-oss-120bdataset. Records are linked via a unique sample_uuid. Main Dataset: Contains the text (prompts… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b-Logprob.texttext-generation100K<n<1M63 likes2.7k downloads8mo agoHugging Face10Alibaba-AAIG /StreamGuardBench Kelp: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk Detection 💻 GitHub 💡 Dataset Overview StreamGuardBench is the first benchmark specifically designed for evaluating streaming guardrails. StreamGuardBench prompts ten widely used open-source LMs—comprising five text-only and five vision-language models—and annotates every generated response with harm labels, therefore enabling accurate measurement of streaming guardrail effectiveness in… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-AAIG/StreamGuardBench.image100K<n<1M1 likes1.6k downloads1d agoHugging Face11Alibaba-Apsara /Superior-Reasoning-SFT-gpt-oss-120b Superior-Reasoning-SFT-gpt-oss-120b           📣 News Our dataset ranked #1 on the Hugging Face Datasets Trending leaderboard from January 20 to January 30. 🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b.texttext-generation100K<n<1M352 likes1.4k downloads8mo agoHugging Face12Alibaba-YuFeng /MMA-SafetyBenchimagen<1K20 likes996 downloads5mo agoHugging Face13Alibaba-DAMO-Academy /RynnBrain-Bench RynnBrain-Bench Introduction We introduce RynnBrain-Bench, a high-dimensional evaluation suite designed to holistically benchmark the cognition and localization capabilities of embodied understanding models in complex household environments. Advancing beyond existing benchmarks, RynnBrain-Bench features a unique emphasis on fine-grained understanding and precise spatiotemporal localization within episodic video sequences. RynnBrain-Bench systematically… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/RynnBrain-Bench.textvisual-question-answering10K<n<100K14 likes956 downloads7mo agoHugging Face14alibabagroup /CMGUI CMGUI Dataset 📄 Project | 🌐 Model in Hugging Face | 🌐 Model in ModelScope English | 简体中文 CMGUI (Chinese Mobile GUI) is a large-scale, high-quality dataset constructed for developing GUI agents on Chinese mobile applications. The dataset contains 18k episodes (i.e., trajectories) with 98k steps collected from more than 50 real-world Chinese mobile apps, covering diverse functional domains such as e-commerce (e.g., Taobao, Pinduoduo), social media (e.g., Rednote… See the full description on the dataset page: https://huggingface.co/datasets/alibabagroup/CMGUI.100K<n<1M12 likes848 downloads6mo agoHugging Face15alibabasglab /VoxCeleb2-mixA modified version of the VoxCeleb2 Dataset. Original data can be downloaded here. This dataset is used for Audio-visual speaker extraction conditioned on face recordings in the reentry paper, which the code can be found here (ClearVoice repo) or here (Paper repo). Usage cat orig* > orig.tar tar -xvf orig.tar cat audio_clean* > audio_clean.tar tar -xvf audio_clean.tar 6 likes744 downloads2y agoHugging Face16Alibaba-Aone /aacr-bench Dataset for Running AACR-Bench English | 简体中文 This is a test set designed for automated code review reflection models, primarily aiming to evaluate the extent to which a model can intercept low-quality review comments. The dataset contains 2,145 code review comments, consisting of 1,505 expert-verified correct comments and 640 incorrect comments. This data is part of the AACR-Bench project and is provided by the Alibaba Aone team. Data Sample Each sample in the… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Aone/aacr-bench.tabulartext-generation1K<n<10K9 likes561 downloads8mo agoHugging Face17alibaba-yuanjing-aigclab /ViViDimage10K<n<100K6 likes500 downloads2y agoHugging Face18Alibaba-NLP /WebShaper WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization Github: https://github.com/Alibaba-NLP/WebAgent Paper: https://arxiv.org/pdf/2507.15061 TLTR WebShaper is a synthesized training dataset for information-seeking (IS) task. It is based on our proposed task formalization of IS, and synthesized by our Expander Agent. WebShaper would cover a broader range of task forms, reasoning structure, and diversified knowledge. Description… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-NLP/WebShaper.textn<1K26 likes465 downloads1y agoHugging Face19Alibaba-NLP /UVRB 🌐 Universal Video Retrieval Benchmark (UVRB) The first comprehensive benchmark for universal video retrievalEvaluate your model across 16 datasets, 3 query types, and 6 capability dimensions — not just accuracy, but why it succeeds or fails. UVRB is a comprehensive evaluation suite designed to diagnose and quantify a video embedding model’s true generalization ability — beyond narrow text-to-video tasks. It exposes critical gaps in spatial reasoning, temporal dynamics… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-NLP/UVRB.videosentence-similarity10K<n<100K6 likes444 downloads11mo agoHugging Face20alibaba-pai /OmniThought OmniThought: A Large-Scale Chain-of-Thought Dataset for Advancing Large Reasoning Models Overview The rise of Large Reasoning Models (LRMs) has revolutionized Natural Language Processing (NLP), enabling breakthroughs in complex tasks like mathematical problem-solving and code generation. These models rely on Chain-of-Thought (CoT) processes to mimic human-like reasoning. However, progress in LRMs is limited by the scarcity of high-quality, large-scale CoT… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-pai/OmniThought.text100K<n<1M17 likes357 downloads1y agoHugging Face21Alibaba-NLP /SecRespond SecRespond 💻 GitHub&nbsp; | &nbsp; 🤖 ModelScope&nbsp; | &nbsp; 📄 Paper Introduction SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response evaluates whether an AI agent can investigate a compromised host after an attack has already succeeded. For each cyber range, the responder receives a frozen forensic disk snapshot together with synthetic host-security-product outputs, then produces an evidence-backed incident… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-NLP/SecRespond.text-generation2 likes345 downloads2mo agoHugging Face22Alibaba-AAIG /XGuard-Train-Open-200K YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models   🤗 HuggingFace   |   🤖 ModelScope   |   📄 Paper   🐬 Introduction XGuard-Train-Open-200K is an open-source subset of the training corpus developed for the YuFeng-XGuard-Reason guardrail model series. YuFeng-XGuard-Reason is engineered to accurately identify security risks in user requests, model responses, and general text, while… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-AAIG/XGuard-Train-Open-200K.texttext-classification100K<n<1M4 likes273 downloads6mo agoHugging Face23Alibaba-NLP /EcomBench EcomBench: Where Intelligent Agents Conquer Commerce Realms 🚀 Benchmark Overview EcomBench is a domain-specific, real-world evaluation framework designed to rigorously assess the capabilities of AI agents in delivering practical support for the complex, ever-evolving demands of e-commerce. We believe that truly capable AI agents will fundamentally transform how we interact with commerce. E-commerce represents one of the world's most significant economic sectors, with… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-NLP/EcomBench.textn<1K7 likes265 downloads10mo agoHugging Face24Alibaba-DAMO-Academy /InterVBench Video Drift Evaluation (vde.py) This repository contains a single entry point, vde.py, that computes Video Drift Error (VDE) scores for every .mp4 file inside a target directory. VDE provides a simple way to monitor how quality-related metrics drift across chunks of the same video. The script already supports several metric backends (clarity, motion, aesthetic, dynamic, subject, background) via the vbench tooling. Environment Setup Install the project… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/InterVBench.2 likes244 downloads3mo agoHugging Face25Alibaba-DAMO-Academy /RynnEC-Bench RynnEC-Bench RynnEC-Bench evaluates fine-grained embodied understanding models from the perspectives of object cognition and spatial cognition in open-world scenario. The benchmark includes 507 video clips captured in real household scenarios. Model Overall Mean Object Properties Seg. DR Seg. SR Object Mean Ego. His. Ego. Pres. Ego. Fut. World Size World Dis. World PR Spatial Mean GPT-4o 28.3 41.1 --- --- 33.9 13.4 22.8 6.0 24.3 16.7 36.1 22.2 GPT-4.1 33.5… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/RynnEC-Bench.7 likes223 downloads11mo agoHugging Face26alirajabi /alibaba_dataset alibaba_dataset Dataset Information This dataset contains multiple parts of a large archive that need to be combined. Reconstruction Instructions To reconstruct the original dataset file from the parts: Download all parts (alibaba_dataset_part_*) Use the cat command to join them:cat alibaba_dataset_part_* > alibaba_dataset.tar.xz Extract the tarball:tar -xf alibaba_dataset.tar.xz 0 likes218 downloads1y agoHugging Face27alibaba-multimodal-industrial-ai /IndustryBench IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs 💻Github | 📝Paper IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese. Overview Dimension Details Total questions 2,049 Languages Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.textquestion-answering1K<n<10K30 likes199 downloads4mo agoHugging Face28alibaba-pai /SimpleQA-Bench SimpleQA-Bench Tags: factuality, EN, ZH, short-form-answer, human-label Copyright: © 2024 alibaba-pai Source.OpenAI's SimpleQA: Blog & Paper / Data & simple-evals ProjectOpenStellarTeam's Chinese-SimpleQA: Blog & Paper, Data@HF Factuality is a complicated topic because it is hard to measure—evaluating the factuality of any given arbitrary claim is challenging, and language models can generate long completions that contain dozens of factual claims. In SimpleQA, we will focus on… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-pai/SimpleQA-Bench.text1K<n<10K3 likes178 downloads2y agoHugging Face29alibaba-pai /SynSearch-Data SynSearch-Data SynSearch-Data is an environment-aligned Search Agent training dataset produced through task synthesis and Solver-in-the-Loop verification. This first release contains 5,000 high-quality multi-hop ReAct trajectories covering 4,182 audited tasks. Data format Each JSONL row contains: messages: the complete ReAct conversation, including tool calls and tool responses; metadata.task_id / run_idx / trajectory_rank: trajectory provenance; metadata.seed:… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-pai/SynSearch-Data.textquestion-answering1K<n<10K3 likes166 downloads29d agoHugging Face30bluuebunny /arxiv_embeddings_Alibaba-NLP_gte-base-en-v1.5text1M<n<10M1 likes150 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.