datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eMCR
eMCR: A Benchmark for Multi-Condition Product Retrieval in Chinese E-Commerce
This is the official dataset and evaluation code for the paper "eMCR: A Benchmark for Multi-Condition Product Retrieval in Chinese E-Commerce".
Overview
Product search increasingly involves queries that combine multiple requirements — product attributes, brands, prices, exclusions, and visual descriptions. Existing retrieval benchmarks provide limited support for diagnosing which… See the full description on the dataset page: https://huggingface.co/datasets/alibabagroup/eMCR.IndustryBench-MIPU
IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products
Multi-Image Industrial Product Understanding Benchmark — evaluating MLLMs on structured attribute extraction from real-world industrial product images.
Industrial product specifications are scattered across multiple heterogeneous images — specification tables, nameplates, technical drawings. IndustryBench-MIPU tests whether MLLMs can reliably recover them through four… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench-MIPU.OmniDoc-TokenBench
OmniDoc-TokenBench
📄 Overview
Introduction
We propose OmniDoc-TokenBench in Qwen-Image-VAE-2.0, a curated benchmark specifically designed to evaluate VAE reconstruction on text-rich document images. It contains ~3K samples spanning nine categories (book, slides, color textbook, exam paper, academic paper, magazine, financial report, newspaper, note) in both English and Chinese, alongside an evaluation toolkit supporting PSNR, SSIM, LPIPS, FID, and… See the full description on the dataset page: https://huggingface.co/datasets/alibabagroup/OmniDoc-TokenBench.StreamGuardBench
Kelp: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk Detection
💻 GitHub
💡 Dataset Overview
StreamGuardBench is the first benchmark specifically designed for evaluating streaming guardrails. StreamGuardBench prompts ten widely used open-source LMs—comprising five text-only and five vision-language models—and annotates every generated response with harm labels, therefore enabling accurate measurement of streaming guardrail effectiveness in… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-AAIG/StreamGuardBench.MMA-SafetyBenchViViD
