datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CRA5-Dataset
Climate science data can be compressed efficiently by dual-stage extreme compression with a variational auto-encoder transformer
Introduction and get started
CRA5 dataset now is available at OneDrive
Paper Summary
We introduce VAEformer, a variational autoencoder transformer designed for the extreme compression of climate data. Addressing the storage challenges of massive datasets like ERA5, VAEformer utilizes a… See the full description on the dataset page: https://huggingface.co/datasets/taohan10200/CRA5-Dataset.TAO-Amodal
TAO-Amodal Dataset
Official Source for Downloading the TAO-Amodal and TAO Dataset.
📙 Project Page | 💻 Code | 📎 Paper Link | ✏️ Citations
Contact: 🙋🏻♂️Cheng-Yen (Wesley) Hsieh
Dataset Description
Our dataset augments the TAO dataset with amodal bounding box annotations for fully invisible, out-of-frame, and occluded objects.
Note that this implies TAO-Amodal also includes modal segmentation masks (as visualized in the color overlays above).
Our… See the full description on the dataset page: https://huggingface.co/datasets/chengyenhsieh/TAO-Amodal.Tstars-VTON
Tstars-Tryon 1.0
Commercial Applications
Our virtual try-on model, Tstars-Tryon 1.0, is now deployed on the Taobao App.
Simply scan the QR code below with the Taobao app to instantly try on your favorite looks.
We hope you enjoy a seamless and delightful shopping experience!
Tstars-VTON - MetaInfo
Introduction
Tstars-VTON is a comprehensive benchmark designed to evaluate whether a virtual try-on… See the full description on the dataset page: https://huggingface.co/datasets/TaobaoTmall-AlgorithmProducts/Tstars-VTON.VisualWebInstruct
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
VisualWebInstruct is a large-scale, diverse multimodal instruction dataset designed to enhance vision-language models' reasoning capabilities. The dataset contains approximately 900K question-answer (QA) pairs, with 40% consisting of visual QA pairs associated with 163,743 unique images, while the remaining 60% are text-only QA pairs.
Links
GitHub Repository
Research Paper
Project Website… See the full description on the dataset page: https://huggingface.co/datasets/taoye1992/VisualWebInstruct.CPI-benchmark
CPI-Bench
Introduction
CPI-Bench is a comprehensive suite of benchmarks designed to evaluate whether an image
generation/editing model is truly capable of handling diverse, real-world, and
knowledge-intensive tasks. It consists of three complementary subsets:
Benchmark
Description
Data Files
CPI-General-Benchmark
General-purpose image editing tasks covering a wide range of task types
CPI_general_benchmark/CPI_general_benchmark-*.parquet… See the full description on the dataset page: https://huggingface.co/datasets/TaobaoTmall-AlgorithmProducts/CPI-benchmark.COCO2017solar-panel-orientationMMhopsVOC2012taobao-product-context
Taobao Product Context
Private product-detail snapshots for agents that consume ordered product images and OCR-derived context.
Dataset versions
Dataset: 0.1.0
Schema: 1.0.0
Snapshot: 2026-08-31
Pipeline: generated from the pipeline Git commit recorded in each row
Configurations
products
One row per strictly validated product snapshot. Each row contains source metadata, ordered relative image paths, image hashes and dimensions… See the full description on the dataset page: https://huggingface.co/datasets/Helios1208/taobao-product-context.catproblemtaobao-search-filter-purchase-trajectories500-v5-filter-groundingTAO-Amodal-Segment-Object-Large
Segment-Object Dataset
This dataset is collected from LVIS and COCO. We employed the segments in this dataset to implement PasteNOcclude augmentation proposed in Tracking Any Object Amodally.
📙 Project Page | 💻 Code | 📎 Paper Link | ✏️ Citations
Contact: 🙋🏻♂️Cheng-Yen (Wesley) Hsieh
Dataset Download
git lfs install
git clone git@hf.co:datasets/chengyenhsieh/TAO-Amodal-Segment-Object-Large
After downloading this dataset, check here to see… See the full description on the dataset page: https://huggingface.co/datasets/chengyenhsieh/TAO-Amodal-Segment-Object-Large.imagenetselectedimagenetselected1cathomeTAO-Amodal-Segment-Objecttaobao-confirm-receipt-review-bulk500-balancedDatasets_processedfaceforensics_h5ioaihscimagenetselected2taobao-refund-search-compare-purchase-bulk500twitter-TaoHuaBang-2026.02.14-2022668616630747371-lTqN5OXHDzmBTKMb-part1taobao-cart-selection-combinations500hu-tao-v2twitter-TaoHuaBang-2026.02.14-2022528613112054224-cj5FKWQ9ahKiiadp-part1twitter-taotaoO6_-2025.12.17-2001227653727268918-qNU03LfjzmLX6GV1-part1twitter-TaoxeMy8-2025.09.20-1969251876290859164-Y0VPohf9SaQl0n3y-part1samson
Description
Samson is a simple dataset that is available from the website. In this image, there are 952x952 pixels. Each pixel is recorded at 156 channels covering the wavelengths from 401 nm to 889 nm. The spectral resolution is highly up to 3.13 nm. As the original image is too large, which is very expensive in terms of computational cost, a region of 95x95 pixels is used. It starts from the (252,332)-th pixel in the original image. This data is not degraded by the blank channel… See the full description on the dataset page: https://huggingface.co/datasets/taowang1122333/samson.
