datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dare-bench
DARE-Bench
[ICLR 2026] DARE-Bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science
Fan Shu1, Yite Wang2, Ruofan Wu1, Boyi Liu2, Zhewei Yao2, Yuxiong He2, Feng Yan1
1University of Houston 2Snowflake AI Research
🔎 Overview
DARE-Bench (ICLR 2026) is a benchmark for evaluating LLM agents on data science tasks, focusing on modeling and instruction fidelity.
This Hugging Face repository provides a selected subset of the full benchmark for public release.… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/dare-bench.HybridDeepResearch
HybridDeepResearch
A benchmark for deep-research agents that reason across SQL databases and the open web.
🔎 Overview
HybridDeepResearch asks an agent to keep constraints while moving between structured database records and unstructured web evidence. Each task is one of three categories:
SQL-to-Search (SQL2S): query the database to obtain a bridge entity, then resolve a web question.
Search-to-SQL (S2SQL): identify an entity from web evidence, then use it as a… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/HybridDeepResearch.
