datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
msmarco-v2.1-snowflake-arctic-embed-l
Snowflake Arctic Embed L Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using Snowflake's Arctic Embed L and are intended to serve as a simple baseline for dense retrieval-based methods.
Retrieval Performance
Retrieval performance for the TREC DL21-23, MSMARCOV2-Dev and Raggy Queries can be found below with BM25 as a baseline. For both… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-l.snow-mountainThe Snow Mountain dataset contains the audio recordings (in .mp3 format) and the corresponding text of The Bible
in 11 Indian languages. The recordings were done in a studio setting by native speakers. Each language has a single
speaker in the dataset. Most of these languages are geographically concentrated in the Northern part of India around
the state of Himachal Pradesh. Being related to Hindi they all use the Devanagari script for transcription.Multisite-PPG
Multisite PPG Dataset
A multisite photoplethysmography (PPG) dataset: long-duration recordings from four body locations, synchronized activity logs, and ECG-derived heart-rate ground truth. It supports PPG-based HR estimation, signal-quality assessment, motion-artifact handling, and cross-site generalization research.
For data preprocessing and baseline training code, see our GitHub repository: anonymous-ppg/wearable-ppg-dataset.
At a glance
Approx.… See the full description on the dataset page: https://huggingface.co/datasets/snowballlab/Multisite-PPG.SafeIMG
SafeIMG
AI-generated Images Challenge Visual Trust in High-risk Scenarios
Yi-Zhi Wang1,2, Yichen Xiao1,2, Linan Yue1,2, Weibo Gao3, Yichao Du4, Pengfei Fang1,2, Shimin Di1,2, Min-Ling Zhang1,2
1 Southeast University 2 Key Laboratory of Computer Network and Information Integration, Ministry of Education3 The Hong Kong Polytechnic University 4 School of Artificial Intelligence, Wuhan University
Overview · Dataset · Results · Quick Start ·… See the full description on the dataset page: https://huggingface.co/datasets/Snowstorm1492/SafeIMG.mteb-retrieval-snowflake-arctic-embed-m-v1.5korean_mmqa_competitionsnodas-snowmelt-cachemsmarco-v2.1-snowflake-arctic-embed-m-v1.5
Snowflake Arctic Embed M V1.5 Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using Snowflake's Arctic Embed M v1.5 and are intended to serve as a simple baseline for dense retrieval-based methods.
It's worth noting that Snowflake's Arctic Embed M v1.5 is optimized for efficient embeddings and thus supports embedding truncation and quantization. More… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-m-v1.5.dartlab-data
DartLab Data
Structured company data from DART & EDGAR disclosure filings
DART 전자공시 + EDGAR 공시 데이터 — 한국 2,700사 / 미국 970사
What is this?
Pre-collected Parquet files from DartLab — a Python library that turns DART (Korea) and EDGAR (US) disclosure filings into one structured company map.
한국 DART 전자공시 시스템과 미국 SEC EDGAR에서 수집한 기업 공시 데이터입니다.
This dataset is the data layer behind DartLab. When you run dartlab.Company("005930"), the library automatically downloads the… See the full description on the dataset page: https://huggingface.co/datasets/Snowfall0601/dartlab-data.corruption-snow
Corruption Dataset: Snow
Dataset Description
This dataset contains corrupted versions of ImageNet-1K images using snow corruption. It is part of the ImageNet-C benchmark for evaluating model robustness to common image corruptions.
Dataset Structure
Train: 1,281,167 corrupted images
Validation: 50,000 corrupted images
Classes: 1000 ImageNet-1K classes
Format: Arrow (Hugging Face Datasets)
Corruption Type: Snow
Adds snow effects to images… See the full description on the dataset page: https://huggingface.co/datasets/MarMaster/corruption-snow.AgentWorldModel-1KAgentWorldModel-1K
Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
Zhaoyang Wang1,
Canwen Xu2,
Boyi Liu2,
Yite Wang2,
Siwei Han1,
Zhewei Yao2,
Huaxiu Yao1,
Yuxiong He2
1UNC-Chapel Hill 2Snowflake AI Research
Overview
AgentWorldModel-1K contains 1,000 fully synthetic, executable, SQL database-backed tool-use environments exposed via a unified MCP (Model Context Protocol) interface, designed for large-scale… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/AgentWorldModel-1K.dare-bench
DARE-Bench
[ICLR 2026] DARE-Bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science
Fan Shu1, Yite Wang2, Ruofan Wu1, Boyi Liu2, Zhewei Yao2, Yuxiong He2, Feng Yan1
1University of Houston 2Snowflake AI Research
🔎 Overview
DARE-Bench (ICLR 2026) is a benchmark for evaluating LLM agents on data science tasks, focusing on modeling and instruction fidelity.
This Hugging Face repository provides a selected subset of the full benchmark for public release.… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/dare-bench.Snowball-67B-A2B-Mixed-RLVR-Experiment-Artifacts
Snowball 67B-A2B RL artifact release
2026 mixed-domain RLVR campaign
This release also contains the complete releasable record of the September 2026 Snowball mixed-domain RLVR campaign.
It covers the September 11 synchronous and bounded-staleness asynchronous RLVR1→RLVR2 lineages and the 5.7T
Agentic-start RLVR1 lineage. All training arms are terminal. The final campaign figure,
trace audit, canonical configs, timing reports, retained
traces, and operational… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/Snowball-67B-A2B-Mixed-RLVR-Experiment-Artifacts.snowball-replay-index
Snowball replay index
This dataset is a compact membership and ordering index for an approximate replay of Snowball's 10,372,343,704,053-token
data store. It contains no source text or token arrays. The 6,301 Parquet files contain three columns:
source_id: logical source key; join it to the source_id field in sources.json
document_id: the retained XXH3-128 content hash as 16 bytes
bucket_id: domain_cluster * 5 + quality_bucket
Document join contract
document_id… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/snowball-replay-index.qwen38-executor-train-v1
Qwen3.8 Executor Unified Training Dataset
Purpose
Public, provenance-pinned executor records for tool use, repository agents, computer use, and Korean coverage. This is an attributed integration; the upstream authors collected or generated the source data.
Schema
The Parquet files share the canonical fields documented in manifests/schema.json. messages uses role/content/tool-call structs. Dynamic tool arguments are validated JSON in arguments_json.… See the full description on the dataset page: https://huggingface.co/datasets/snowman0919/qwen38-executor-train-v1.arctic-embed-ft-v1
Data for the Arctic Embed walkthrough
This dataset coresponds to the walkthrough example for using the Arctic Embed training code in ArcticTraining. See that README for more details.
Example: Selective downloads via Git LFS
Since this dataset contains various intermediate files not necessary for training, it can be helpful to use the Git LFS backend of Hugging Face Datasets to pull select files.
# First, ensure you have installed git-lfs (see `https://git-lfs.com/`… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/arctic-embed-ft-v1.snowball-5.7t-sft-eval-artifacts
Snowball 5.7T cold-start SFT evaluation artifacts
This dataset archives the evaluation records, sampled traces, resolved launch configurations,
analysis inputs, and derived tables for
marin-community/marin#8225.
The experiment compares the 5.7T-token Snowball cooldown with its Chat, Thinking, and
Nemotron-Terminal SFT descendants. It also includes the corresponding 2T-token cooldown cohort.
The top-level EVAL_RESULTS.csv in the experiment record is generated from the durable… See the full description on the dataset page: https://huggingface.co/datasets/penfever/snowball-5.7t-sft-eval-artifacts.Nyx-mmE5-MMEBRea2Seg-16K
Rea2Seg-16K
Rea2Seg-16K is a high-quality, large-scale training dataset for reasoning segmentation, with detailed chain-of-thought annotations.
Splits
train: 16090 rows
Categories
gqa-seg: 7998 rows
lisa-plus: 7052 rows
spatial: 801 rows
reasonseg: 239 rows
Fields
Each row contains:
image: source image
question: referring question
cot_answer: chain-of-thought answer
mask: COCO RLE mask serialized as a JSON string
category: source… See the full description on the dataset page: https://huggingface.co/datasets/snowball521/Rea2Seg-16K.Indian_food_imagesImageNet-C-snow-severity_5rocky_mountain_snowpack
Rocky Mountain Snowpack Dataset
The Rocky Mountain Snowpack dataset contains 4,040 preprocessed samples of snowpack imagery collected in the Colorado Rocky Mountains across the 2024–2025 and 2025–2026 winter seasons, from 7 snowpits dug between January 2025 and February 2026.Each sample segment of snow includes three types of images:
Magnified crystal images (close-up snow snow crystal profile photography)
Snowpack profile images (non-magnified snow crystal profiles… See the full description on the dataset page: https://huggingface.co/datasets/RMDig/rocky_mountain_snowpack.qwen38-executor-train-v2
Qwen3.8 Executor Unified Training Dataset
Purpose
Public, provenance-pinned executor records for tool use, repository agents, computer use, and Korean coverage. This is an attributed integration; the upstream authors collected or generated the source data.
Schema
The Parquet files share the canonical fields documented in manifests/schema.json. messages uses role/content/tool-call structs. Dynamic tool arguments are validated JSON in arguments_json.… See the full description on the dataset page: https://huggingface.co/datasets/snowman0919/qwen38-executor-train-v2.SNOW.SBThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "lekiwi_client",
"total_episodes": 91,
"total_frames": 73578,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 24,
"splits": {
"train": "0:91"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/williamdgomez/SNOW.SB.snow_simplified_japanese_corpusAbout SNOW T15: The simplified corpus for the Japanese language. The corpus has 50,000 manually simplified and aligned sentences. This corpus contains the original sentences, simplified sentences and English translation of the original sentences. It can be used for automatic text simplification as well as translating simple Japanese into English and vice-versa. The core vocabulary is restricted to 2,000 words where it is selected by accounting for several factors such as meaning preservation, variation, simplicity and the UniDic word segmentation criterion.
For details, refer to the explanation page of Japanese simplification (http://www.jnlp.org/research/Japanese_simplification). The original texts are from "small_parallel_enja: 50k En/Ja Parallel Corpus for Testing SMT Methods", which is a bilingual corpus for machine translation. About SNOW T23: An expansion corpus of 35,000 sentences rewritten in easy Japanese (simple Japanese vocabulary) based on SNOW T15. The original texts are from "Tanaka Corpus" (http://www.edrdg.org/wiki/index.php/Tanaka_Corpus).goes-omni-electron-flux-forecasting
GOES–OMNI >2 MeV Electron Flux Forecasting
This dataset combines cross-calibrated NOAA GOES-14/GOES-16 >2 MeV electron
flux with NASA/GSFC OMNI solar-wind and geomagnetic drivers on a uniform
five-minute UTC grid.
It provides two configurations:
ml-ready (default): scaled causal features, validity flags, unscaled
30-minute/6-hour/12-hour targets, and leakage-safe chronological splits.
scientific-master: unscaled source measurements, instrument context,
calibration factors, and… See the full description on the dataset page: https://huggingface.co/datasets/snowsadh/goes-omni-electron-flux-forecasting.omnimcp_sql_snowflake_warehouse_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_sql_snowflake_warehouse_teaser.DPL-main
Difference-aware Personalized Learning (DPL) Dataset
This dataset is used in the paper:
Measuring What Makes You Unique: Difference-Aware User Modeling for Enhancing LLM Personalization
Yilun Qiu, Xiaoyan Zhao, Yang Zhang, Yimeng Bai, Wenjie Wang, Hong Cheng, Fuli Feng, Tat-Seng Chua
Code
This dataset is an adaptation of the Amazon Reviews'23 dataset. It contains user reviews for Books, CDs & Vinyl, and Movies & TV. Each review includes user ID, profile information (ASIN… See the full description on the dataset page: https://huggingface.co/datasets/SnowCharmQ/DPL-main.snowden_archived
The Snowden Archive
This repository is a complete collection of all documents leaked by former National Security Agency contractor and whistleblower Edward Snowden that have subsequently been published by news media around the world.
If you notice something is missing or wrong, please file an issue or tweet at @iamcryptoki.
Timeline of Revelations (2013-2018)
2013
Date
Source
Publication
06/06/2013
Foreign Intelligence Surveillance Court… See the full description on the dataset page: https://huggingface.co/datasets/yukioitsuki/snowden_archived.prompt-injection-multilingual
