datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
msmarco-v2.1-snowflake-arctic-embed-l
Snowflake Arctic Embed L Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using Snowflake's Arctic Embed L and are intended to serve as a simple baseline for dense retrieval-based methods.
Retrieval Performance
Retrieval performance for the TREC DL21-23, MSMARCOV2-Dev and Raggy Queries can be found below with BM25 as a baseline. For both… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-l.mteb-retrieval-snowflake-arctic-embed-m-v1.5msmarco-v2.1-snowflake-arctic-embed-m-v1.5
Snowflake Arctic Embed M V1.5 Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using Snowflake's Arctic Embed M v1.5 and are intended to serve as a simple baseline for dense retrieval-based methods.
It's worth noting that Snowflake's Arctic Embed M v1.5 is optimized for efficient embeddings and thus supports embedding truncation and quantization. More… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-m-v1.5.dare-bench
DARE-Bench
[ICLR 2026] DARE-Bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science
Fan Shu1, Yite Wang2, Ruofan Wu1, Boyi Liu2, Zhewei Yao2, Yuxiong He2, Feng Yan1
1University of Houston 2Snowflake AI Research
🔎 Overview
DARE-Bench (ICLR 2026) is a benchmark for evaluating LLM agents on data science tasks, focusing on modeling and instruction fidelity.
This Hugging Face repository provides a selected subset of the full benchmark for public release.… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/dare-bench.HybridDeepResearch
HybridDeepResearch
A benchmark for deep-research agents that reason across SQL databases and the open web.
🔎 Overview
HybridDeepResearch asks an agent to keep constraints while moving between structured database records and unstructured web evidence. Each task is one of three categories:
SQL-to-Search (SQL2S): query the database to obtain a bridge entity, then resolve a web question.
Search-to-SQL (S2SQL): identify an entity from web evidence, then use it as a… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/HybridDeepResearch.msmarco-v2.1-snowflake-arctic-embed-m-v2.0
Snowflake Arctic Embed M V2.0 Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using Snowflake's Arctic Embed M v2.0 and are intended to serve as a simple baseline for dense retrieval-based methods.
Note, that the embeddings are not normalized so you will need to normalize them before usage.
Retrieval Performance
Retrieval performance for… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-m-v2.0.swe_opsd_datasetnanobeir-snowflake-arctic-embed-m-v1.5HyDRA-Bench
HyDRA-Bench
Bridging Databases and Documents: Data-Algorithm Co-Design for Hybrid Question Answering
Ruofan Wu¹, Boyi Liu², Fan Shu¹, Yite Wang², Zhewei Yao², Yuxiong He², Feng Yan¹
¹ University of Houston
² Snowflake AI Research
🔎 Overview
HyDRA-Bench (Hybrid Database and Retrieval Agent Benchmark) is a benchmark for hybrid reasoning that requires LLM agents to interleave SQL execution over structured databases with semantic retrieval over unstructured text.… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/HyDRA-Bench.scaleswe-opsd-v2-3200-summary
Scale-SWE OPSD v2 — 3200 tasks with summary hints
The training set used for the Scale-SWE on-policy self-distillation (OPSD) runs. 3200 SWE tasks across
752 repositories, each paired with a reference agent trajectory and a condensed solution hint.
Uploaded from /checkpoint/huggingface/datasets/scaleswe_opsd_v2_3200_summary (a
datasets.save_to_disk directory), converted to parquet. Row count, ids and field contents verified
identical to the source.
⚠️ Contains… See the full description on the dataset page: https://huggingface.co/datasets/starli-snowflake/scaleswe-opsd-v2-3200-summary.zsql-snowflake-dpofetch-snowflake-3621-test-zz9x8y7wmetadata_snowflakeembedded_testdata_snowflake_arctic_embed_lfineweb-edu-healthcare-snowflake-llama3-8bthe scores (int_score) are inversed , 30 and bellow that means that the text is healthcare related. (95% accuracy according to llama-3-8b-instruct as judge)
snowflake_instruct_dataset_csvtemp2-with-clip-snowflakev2-embeddingsMSMARCObm25Triplets-labeled-MarginMSE-snowflake-arctic-embed-lsnowflake_docs_subsetsnowflake_cot_datasetSnowflakeEmojiClassifier
SnowflakeEmojiClassifier
tags: EmojiClassification, Snowflake, Recognition
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'SnowflakeEmojiClassifier' dataset contains a collection of textual descriptions paired with emojis, focusing specifically on snowflake-related emojis. The dataset is intended for machine learning practitioners to train models in emoji classification, particularly identifying emojis that represent… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/SnowflakeEmojiClassifier.fineweb-edu-healthcare-snowflake-llama3-8b-nano-filtered-threshold-30snowflakeToDatabricksabout_snowflakesnowflake
