CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Snowflake /msmarco-v2.1-snowflake-arctic-embed-l Snowflake Arctic Embed L Embeddings for MSMARCO V2.1 for TREC-RAG This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG All embeddings are created using Snowflake's Arctic Embed L and are intended to serve as a simple baseline for dense retrieval-based methods. Retrieval Performance Retrieval performance for the TREC DL21-23, MSMARCOV2-Dev and Raggy Queries can be found below with BM25 as a baseline. For both… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-l.textquestion-answering10M<n<100M0 likes2.9k downloads2y agoHugging Face02bridgeconn /snow-mountainThe Snow Mountain dataset contains the audio recordings (in .mp3 format) and the corresponding text of The Bible in 11 Indian languages. The recordings were done in a studio setting by native speakers. Each language has a single speaker in the dataset. Most of these languages are geographically concentrated in the Northern part of India around the state of Himachal Pradesh. Being related to Hindi they all use the Devanagari script for transcription.automatic-speech-recognition4 likes2.4k downloads3y agoHugging Face03snowballlab /Multisite-PPG Multisite PPG Dataset A multisite photoplethysmography (PPG) dataset: long-duration recordings from four body locations, synchronized activity logs, and ECG-derived heart-rate ground truth. It supports PPG-based HR estimation, signal-quality assessment, motion-artifact handling, and cross-site generalization research. For data preprocessing and baseline training code, see our GitHub repository: anonymous-ppg/wearable-ppg-dataset. At a glance Approx.… See the full description on the dataset page: https://huggingface.co/datasets/snowballlab/Multisite-PPG.texttime-series-forecasting1K<n<10K1 likes1.6k downloads5mo agoHugging Face04Snowstorm1492 /SafeIMG &nbsp;SafeIMG AI-generated Images Challenge Visual Trust in High-risk Scenarios Yi-Zhi Wang1,2, Yichen Xiao1,2, Linan Yue1,2, Weibo Gao3, Yichao Du4, Pengfei Fang1,2, Shimin Di1,2, Min-Ling Zhang1,2 1 Southeast University &nbsp;&nbsp; 2 Key Laboratory of Computer Network and Information Integration, Ministry of Education3 The Hong Kong Polytechnic University &nbsp;&nbsp; 4 School of Artificial Intelligence, Wuhan University Overview · Dataset · Results · Quick Start ·… See the full description on the dataset page: https://huggingface.co/datasets/Snowstorm1492/SafeIMG.image1K<n<10K0 likes1.2k downloads2mo agoHugging Face05Snowflake /mteb-retrieval-snowflake-arctic-embed-m-v1.5text10M<n<100M0 likes991 downloads2y agoHugging Face06snowchan /korean_mmqa_competitionimage1K<n<10K0 likes871 downloads1mo agoHugging Face07Jsinowitz /snodas-snowmelt-cache1 likes813 downloads7d agoHugging Face08Snowflake /msmarco-v2.1-snowflake-arctic-embed-m-v1.5 Snowflake Arctic Embed M V1.5 Embeddings for MSMARCO V2.1 for TREC-RAG This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG All embeddings are created using Snowflake's Arctic Embed M v1.5 and are intended to serve as a simple baseline for dense retrieval-based methods. It's worth noting that Snowflake's Arctic Embed M v1.5 is optimized for efficient embeddings and thus supports embedding truncation and quantization. More… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-m-v1.5.textquestion-answering10M<n<100M0 likes731 downloads2y agoHugging Face09Snowfall0601 /dartlab-data DartLab Data Structured company data from DART & EDGAR disclosure filings DART 전자공시 + EDGAR 공시 데이터 — 한국 2,700사 / 미국 970사 What is this? Pre-collected Parquet files from DartLab — a Python library that turns DART (Korea) and EDGAR (US) disclosure filings into one structured company map. 한국 DART 전자공시 시스템과 미국 SEC EDGAR에서 수집한 기업 공시 데이터입니다. This dataset is the data layer behind DartLab. When you run dartlab.Company("005930"), the library automatically downloads the… See the full description on the dataset page: https://huggingface.co/datasets/Snowfall0601/dartlab-data.table-question-answering1K<n<10K0 likes687 downloads3mo agoHugging Face10MarMaster /corruption-snow Corruption Dataset: Snow Dataset Description This dataset contains corrupted versions of ImageNet-1K images using snow corruption. It is part of the ImageNet-C benchmark for evaluating model robustness to common image corruptions. Dataset Structure Train: 1,281,167 corrupted images Validation: 50,000 corrupted images Classes: 1000 ImageNet-1K classes Format: Arrow (Hugging Face Datasets) Corruption Type: Snow Adds snow effects to images… See the full description on the dataset page: https://huggingface.co/datasets/MarMaster/corruption-snow.imageimage-classification1M<n<10M0 likes686 downloads11mo agoHugging Face11Snowflake /AgentWorldModel-1KAgentWorldModel-1K Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning Zhaoyang Wang1, Canwen Xu2, Boyi Liu2, Yite Wang2, Siwei Han1, Zhewei Yao2, Huaxiu Yao1, Yuxiong He2 1UNC-Chapel Hill   2Snowflake AI Research   Overview AgentWorldModel-1K contains 1,000 fully synthetic, executable, SQL database-backed tool-use environments exposed via a unified MCP (Model Context Protocol) interface, designed for large-scale… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/AgentWorldModel-1K.71 likes646 downloads7mo agoHugging Face12Snowflake /dare-bench DARE-Bench [ICLR 2026] DARE-Bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science Fan Shu1, Yite Wang2, Ruofan Wu1, Boyi Liu2, Zhewei Yao2, Yuxiong He2, Feng Yan1 1University of Houston   2Snowflake AI Research 🔎 Overview DARE-Bench (ICLR 2026) is a benchmark for evaluating LLM agents on data science tasks, focusing on modeling and instruction fidelity. This Hugging Face repository provides a selected subset of the full benchmark for public release.… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/dare-bench.texttext-generation1K<n<10K6 likes605 downloads7mo agoHugging Face13open-athena /Snowball-67B-A2B-Mixed-RLVR-Experiment-Artifacts Snowball 67B-A2B RL artifact release 2026 mixed-domain RLVR campaign This release also contains the complete releasable record of the September 2026 Snowball mixed-domain RLVR campaign. It covers the September 11 synchronous and bounded-staleness asynchronous RLVR1→RLVR2 lineages and the 5.7T Agentic-start RLVR1 lineage. All training arms are terminal. The final campaign figure, trace audit, canonical configs, timing reports, retained traces, and operational… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/Snowball-67B-A2B-Mixed-RLVR-Experiment-Artifacts.imagen<1K0 likes569 downloads2d agoHugging Face14open-athena /snowball-replay-index Snowball replay index This dataset is a compact membership and ordering index for an approximate replay of Snowball's 10,372,343,704,053-token data store. It contains no source text or token arrays. The 6,301 Parquet files contain three columns: source_id: logical source key; join it to the source_id field in sources.json document_id: the retained XXH3-128 content hash as 16 bytes bucket_id: domain_cluster * 5 + quality_bucket Document join contract document_id… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/snowball-replay-index.tabulartext-generation10B<n<100B0 likes538 downloads17d agoHugging Face15snowman0919 /qwen38-executor-train-v1 Qwen3.8 Executor Unified Training Dataset Purpose Public, provenance-pinned executor records for tool use, repository agents, computer use, and Korean coverage. This is an attributed integration; the upstream authors collected or generated the source data. Schema The Parquet files share the canonical fields documented in manifests/schema.json. messages uses role/content/tool-call structs. Dynamic tool arguments are validated JSON in arguments_json.… See the full description on the dataset page: https://huggingface.co/datasets/snowman0919/qwen38-executor-train-v1.text1M<n<10M1 likes438 downloads1mo agoHugging Face16Snowflake /arctic-embed-ft-v1 Data for the Arctic Embed walkthrough This dataset coresponds to the walkthrough example for using the Arctic Embed training code in ArcticTraining. See that README for more details. Example: Selective downloads via Git LFS Since this dataset contains various intermediate files not necessary for training, it can be helpful to use the Git LFS backend of Hugging Face Datasets to pull select files. # First, ensure you have installed git-lfs (see `https://git-lfs.com/`… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/arctic-embed-ft-v1.100K<n<1M2 likes402 downloads1y agoHugging Face17penfever /snowball-5.7t-sft-eval-artifacts Snowball 5.7T cold-start SFT evaluation artifacts This dataset archives the evaluation records, sampled traces, resolved launch configurations, analysis inputs, and derived tables for marin-community/marin#8225. The experiment compares the 5.7T-token Snowball cooldown with its Chat, Thinking, and Nemotron-Terminal SFT descendants. It also includes the corresponding 2T-token cooldown cohort. The top-level EVAL_RESULTS.csv in the experiment record is generated from the durable… See the full description on the dataset page: https://huggingface.co/datasets/penfever/snowball-5.7t-sft-eval-artifacts.text-generation0 likes338 downloads1mo agoHugging Face18SnowNation /Nyx-mmE5-MMEBtext1M<n<10M1 likes294 downloads1y agoHugging Face19snowball521 /Rea2Seg-16K Rea2Seg-16K Rea2Seg-16K is a high-quality, large-scale training dataset for reasoning segmentation, with detailed chain-of-thought annotations. Splits train: 16090 rows Categories gqa-seg: 7998 rows lisa-plus: 7052 rows spatial: 801 rows reasonseg: 239 rows Fields Each row contains: image: source image question: referring question cot_answer: chain-of-thought answer mask: COCO RLE mask serialized as a JSON string category: source… See the full description on the dataset page: https://huggingface.co/datasets/snowball521/Rea2Seg-16K.image10K<n<100K0 likes269 downloads1mo agoHugging Face20snow3751 /Indian_food_imagesimage10K<n<100K0 likes269 downloads1mo agoHugging Face21souravc83 /ImageNet-C-snow-severity_5image10K<n<100K0 likes254 downloads5mo agoHugging Face22RMDig /rocky_mountain_snowpack Rocky Mountain Snowpack Dataset The Rocky Mountain Snowpack dataset contains 4,040 preprocessed samples of snowpack imagery collected in the Colorado Rocky Mountains across the 2024–2025 and 2025–2026 winter seasons, from 7 snowpits dug between January 2025 and February 2026.Each sample segment of snow includes three types of images: Magnified crystal images (close-up snow snow crystal profile photography) Snowpack profile images (non-magnified snow crystal profiles… See the full description on the dataset page: https://huggingface.co/datasets/RMDig/rocky_mountain_snowpack.imageimage-classification1K<n<10K1 likes246 downloads1mo agoHugging Face23snowman0919 /qwen38-executor-train-v2 Qwen3.8 Executor Unified Training Dataset Purpose Public, provenance-pinned executor records for tool use, repository agents, computer use, and Korean coverage. This is an attributed integration; the upstream authors collected or generated the source data. Schema The Parquet files share the canonical fields documented in manifests/schema.json. messages uses role/content/tool-call structs. Dynamic tool arguments are validated JSON in arguments_json.… See the full description on the dataset page: https://huggingface.co/datasets/snowman0919/qwen38-executor-train-v2.text1M<n<10M1 likes220 downloads1mo agoHugging Face24williamdgomez /SNOW.SBThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "lekiwi_client", "total_episodes": 91, "total_frames": 73578, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 24, "splits": { "train": "0:91" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/williamdgomez/SNOW.SB.tabularrobotics10K<n<100K0 likes210 downloads5mo agoHugging Face25SNOW-NLP /snow_simplified_japanese_corpusAbout SNOW T15: The simplified corpus for the Japanese language. The corpus has 50,000 manually simplified and aligned sentences. This corpus contains the original sentences, simplified sentences and English translation of the original sentences. It can be used for automatic text simplification as well as translating simple Japanese into English and vice-versa. The core vocabulary is restricted to 2,000 words where it is selected by accounting for several factors such as meaning preservation, variation, simplicity and the UniDic word segmentation criterion. For details, refer to the explanation page of Japanese simplification (http://www.jnlp.org/research/Japanese_simplification). The original texts are from "small_parallel_enja: 50k En/Ja Parallel Corpus for Testing SMT Methods", which is a bilingual corpus for machine translation. About SNOW T23: An expansion corpus of 35,000 sentences rewritten in easy Japanese (simple Japanese vocabulary) based on SNOW T15. The original texts are from "Tanaka Corpus" (http://www.edrdg.org/wiki/index.php/Tanaka_Corpus).translation10K<n<100K21 likes186 downloads3y agoHugging Face26snowsadh /goes-omni-electron-flux-forecasting GOES–OMNI >2 MeV Electron Flux Forecasting This dataset combines cross-calibrated NOAA GOES-14/GOES-16 >2 MeV electron flux with NASA/GSFC OMNI solar-wind and geomagnetic drivers on a uniform five-minute UTC grid. It provides two configurations: ml-ready (default): scaled causal features, validity flags, unscaled 30-minute/6-hour/12-hour targets, and leakage-safe chronological splits. scientific-master: unscaled source measurements, instrument context, calibration factors, and… See the full description on the dataset page: https://huggingface.co/datasets/snowsadh/goes-omni-electron-flux-forecasting.tabulartime-series-forecasting1M<n<10M1 likes186 downloads29d agoHugging Face27emgena /omnimcp_sql_snowflake_warehouse_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_sql_snowflake_warehouse_teaser.text-generationn<1K0 likes186 downloads8d agoHugging Face28SnowCharmQ /DPL-main Difference-aware Personalized Learning (DPL) Dataset This dataset is used in the paper: Measuring What Makes You Unique: Difference-Aware User Modeling for Enhancing LLM Personalization Yilun Qiu, Xiaoyan Zhao, Yang Zhang, Yimeng Bai, Wenjie Wang, Hong Cheng, Fuli Feng, Tat-Seng Chua Code This dataset is an adaptation of the Amazon Reviews'23 dataset. It contains user reviews for Books, CDs & Vinyl, and Movies & TV. Each review includes user ID, profile information (ASIN… See the full description on the dataset page: https://huggingface.co/datasets/SnowCharmQ/DPL-main.texttext-generation10K<n<100K2 likes174 downloads1y agoHugging Face29yukioitsuki /snowden_archived The Snowden Archive This repository is a complete collection of all documents leaked by former National Security Agency contractor and whistleblower Edward Snowden that have subsequently been published by news media around the world. If you notice something is missing or wrong, please file an issue or tweet at @iamcryptoki. Timeline of Revelations (2013-2018) 2013 Date Source Publication 06/06/2013 Foreign Intelligence Surveillance Court… See the full description on the dataset page: https://huggingface.co/datasets/yukioitsuki/snowden_archived.document0 likes163 downloads10mo agoHugging Face30rikka-snow /prompt-injection-multilingualtexttext-classification1K<n<10K1 likes144 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.