datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dare-bench
DARE-Bench
[ICLR 2026] DARE-Bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science
Fan Shu1, Yite Wang2, Ruofan Wu1, Boyi Liu2, Zhewei Yao2, Yuxiong He2, Feng Yan1
1University of Houston 2Snowflake AI Research
🔎 Overview
DARE-Bench (ICLR 2026) is a benchmark for evaluating LLM agents on data science tasks, focusing on modeling and instruction fidelity.
This Hugging Face repository provides a selected subset of the full benchmark for public release.… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/dare-bench.omnimcp_sql_snowflake_warehouse_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_sql_snowflake_warehouse_teaser.scaleswe-opsd-v2-3200-summary
Scale-SWE OPSD v2 — 3200 tasks with summary hints
The training set used for the Scale-SWE on-policy self-distillation (OPSD) runs. 3200 SWE tasks across
752 repositories, each paired with a reference agent trajectory and a condensed solution hint.
Uploaded from /checkpoint/huggingface/datasets/scaleswe_opsd_v2_3200_summary (a
datasets.save_to_disk directory), converted to parquet. Row count, ids and field contents verified
identical to the source.
⚠️ Contains… See the full description on the dataset page: https://huggingface.co/datasets/starli-snowflake/scaleswe-opsd-v2-3200-summary.
