datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CRA5-Dataset
Climate science data can be compressed efficiently by dual-stage extreme compression with a variational auto-encoder transformer
Introduction and get started
CRA5 dataset now is available at OneDrive
Paper Summary
We introduce VAEformer, a variational autoencoder transformer designed for the extreme compression of climate data. Addressing the storage challenges of massive datasets like ERA5, VAEformer utilizes a… See the full description on the dataset page: https://huggingface.co/datasets/taohan10200/CRA5-Dataset.Dao_taoTaobao-ProductsAdsDBEmilia-Dataset-tokenisedTaobao-MMTAOBAO-MM: A Long Sequence Recommendation Dataset with Multimodal Embeddings at Scale
Overview |
Dataset Description |
Download and Use |
Contact |
Citation
Overview
TAOBAO-MM is a large-scale recommendation dataset derived from user interaction logs on Taobao, one of the world’s largest e-commerce platforms. The dataset features historical behavior sequences of up to 1,000 interactions per user and includes high-quality multimodal embeddings for… See the full description on the dataset page: https://huggingface.co/datasets/TaoBao-MM/Taobao-MM.LVOmniBench
LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs
LVOmniBench is a new audio-visual understanding evaluation benchmark in long-form audio-video inputs. 🌟
🔥 News
2026.03.19 🌟 We are very proud to launch LVOmniBench, the pioneering comprehensive evaluation benchmark of OmniLLMs in Long Audio-Video Understanding Evaluation!
✨ LVOmniBench Introduction
Recent advancements in omnimodal large language models… See the full description on the dataset page: https://huggingface.co/datasets/KD-TAO/LVOmniBench.TAO-Amodal
TAO-Amodal Dataset
Official Source for Downloading the TAO-Amodal and TAO Dataset.
📙 Project Page | 💻 Code | 📎 Paper Link | ✏️ Citations
Contact: 🙋🏻♂️Cheng-Yen (Wesley) Hsieh
Dataset Description
Our dataset augments the TAO dataset with amodal bounding box annotations for fully invisible, out-of-frame, and occluded objects.
Note that this implies TAO-Amodal also includes modal segmentation masks (as visualized in the color overlays above).
Our… See the full description on the dataset page: https://huggingface.co/datasets/chengyenhsieh/TAO-Amodal.sn38-submission-bsn38-r11-p2sn38-r11-p1bybit-linear-perps-taousdtsn38-submissionE-VAds_Benchmark
🎬 E-VAds Benchmark
E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs
(ICML 2026)
English | 中文文档
📖 Overview
E-VAds (E-commerce Video Ads Benchmark) is the first large-scale benchmark specifically designed to evaluate Multimodal Large Language Models (MLLMs) on conversion-oriented e-commerce short video understanding. Unlike general video QA tasks, e-commerce videos present unique challenges with high-density multimodal signals, rapid visual… See the full description on the dataset page: https://huggingface.co/datasets/TaobaoTmall-AlgorithmProducts/E-VAds_Benchmark.COCO_fullTstars-VTON
Tstars-Tryon 1.0
Commercial Applications
Our virtual try-on model, Tstars-Tryon 1.0, is now deployed on the Taobao App.
Simply scan the QR code below with the Taobao app to instantly try on your favorite looks.
We hope you enjoy a seamless and delightful shopping experience!
Tstars-VTON - MetaInfo
Introduction
Tstars-VTON is a comprehensive benchmark designed to evaluate whether a virtual try-on… See the full description on the dataset page: https://huggingface.co/datasets/TaobaoTmall-AlgorithmProducts/Tstars-VTON.TaobaotaobaoVisualWebInstruct
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
VisualWebInstruct is a large-scale, diverse multimodal instruction dataset designed to enhance vision-language models' reasoning capabilities. The dataset contains approximately 900K question-answer (QA) pairs, with 40% consisting of visual QA pairs associated with 163,743 unique images, while the remaining 60% are text-only QA pairs.
Links
GitHub Repository
Research Paper
Project Website… See the full description on the dataset page: https://huggingface.co/datasets/taoye1992/VisualWebInstruct.sn38-sub-a2sn38-sub-d1literotica-storiesmusic2chords_v2CPI-benchmark
CPI-Bench
Introduction
CPI-Bench is a comprehensive suite of benchmarks designed to evaluate whether an image
generation/editing model is truly capable of handling diverse, real-world, and
knowledge-intensive tasks. It consists of three complementary subsets:
Benchmark
Description
Data Files
CPI-General-Benchmark
General-purpose image editing tasks covering a wide range of task types
CPI_general_benchmark/CPI_general_benchmark-*.parquet… See the full description on the dataset page: https://huggingface.co/datasets/TaobaoTmall-AlgorithmProducts/CPI-benchmark.MSDWild
MSDWild
The MSDWild dataset is designed for testing multi-modal analysis in the following tasks:
Multi-modal Speaker Diarization
Multi-modal Speaker Localization
Audio-visual Lip Sychronization
For further details, please visit the MSDWILD GitHub repository.
A sample from the dataset can be viewed on the visualization section of the repository.
Important Notes:
The database is intended solely for research purposes.
Responding to community feedback, we have uploaded a video.zip… See the full description on the dataset page: https://huggingface.co/datasets/taocode/MSDWild.sn38-sub-e1guitarsetwds_indices_imagenet256TaobaoAd_x1
TaobaoAd_x1
Dataset description:
Taobao is a dataset provided by Alibaba, which contains 8 days of ad click-through data (26 million records) that are randomly sampled from 1140000 users. By default, the first 7 days (i.e., 20170506-20170512) of samples are used as training samples, and the last day's samples (i.e., 20170513) are used as test samples. Meanwhile, the dataset also covers the shopping behavior of all users in the recent 22 days, including totally seven hundred million… See the full description on the dataset page: https://huggingface.co/datasets/reczoo/TaobaoAd_x1.coco_wds_indicesad-display_click-data_taobao.com
