datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HSCodeComp
HSCodeComp: A Realistic and Expert-Level Benchmark for Deep Search Agents in Hierarchical Rule Application
Paper | Code | Dataset on Hugging Face
⭐ MarcoPolo Team ⭐
Alibaba Group
🗂️ Data
📌 Overview
HSCodeComp is the first realistic, expert-level e-commerce benchmark designed to evaluate deep search agents on their ability to perform Level-3 knowledge—hierarchical rule application—a critical yet overlooked capability in current agent evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ATH-MaaS/HSCodeComp.HSCodeComp
HSCodeComp: A Realistic and Expert-Level Benchmark for Deep Search Agents in Hierarchical Rule Application
Paper | Code | Dataset on Hugging Face
⭐ MarcoPolo Team ⭐
Alibaba Group
🗂️ Data
📌 Overview
HSCodeComp is the first realistic, expert-level e-commerce benchmark designed to evaluate deep search agents on their ability to perform Level-3 knowledge—hierarchical rule application—a critical yet overlooked capability in current agent evaluation… See the full description on the dataset page: https://huggingface.co/datasets/chenjianhui0428/HSCodeComp.HSCodeComp
HSCodeComp: A Realistic and Expert-Level Benchmark for Deep Search Agents in Hierarchical Rule Application
Paper | Code | Dataset on Hugging Face
⭐ MarcoPolo Team ⭐
Alibaba Group
🗂️ Data
📌 Overview
HSCodeComp is the first realistic, expert-level e-commerce benchmark designed to evaluate deep search agents on their ability to perform Level-3 knowledge—hierarchical rule application—a critical yet overlooked capability in current agent evaluation… See the full description on the dataset page: https://huggingface.co/datasets/saadz506/HSCodeComp.hs-code_productHScode2hscodehs-code_product
