CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-CC-v2.1gated Nemotron-Pre-Training-Dataset-v2.1 Dataset Description The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web tokens… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.1.texttext-generation1B<n<10B139 likes24k downloads9mo agoHugging Face02AlicanKiraz0 /Cybersecurity-Dataset-Fenrir-v2.1 Cybersecurity Defense Instruction-Tuning Dataset (v2.1) Created by Alican Kiraz TL;DR A ready-to-train dataset of 99,870 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training. Apache-2.0 licensed and production-ready. Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security. 1  What’s… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Cybersecurity-Dataset-Fenrir-v2.1.texttext-generation10K<n<100K146 likes4.8k downloads5mo agoHugging Face03CohereLabs /msmarco-v2.1-embed-english-v3 TREC-RAG 2024 Corpus (MSMARCO 2.1) - Encoded with Cohere Embed English v3 This dataset contains the embeddings for the TREC-RAG Corpus 2024 embedded with the Cohere Embed V3 English model. It contains embeddings for 113,520,750 passages, embeddings for 1677 queries from TREC-Deep Learning 2021-2023, as well as top-1000 hits for all queries using a brute-force (flat) index. Search over the Index We have a pre-build index that only requires 300 MB available at… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/msmarco-v2.1-embed-english-v3.tabular100M<n<1B7 likes3.4k downloads6mo agoHugging Face04Snowflake /msmarco-v2.1-snowflake-arctic-embed-l Snowflake Arctic Embed L Embeddings for MSMARCO V2.1 for TREC-RAG This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG All embeddings are created using Snowflake's Arctic Embed L and are intended to serve as a simple baseline for dense retrieval-based methods. Retrieval Performance Retrieval performance for the TREC DL21-23, MSMARCOV2-Dev and Raggy Queries can be found below with BM25 as a baseline. For both… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-l.textquestion-answering10M<n<100M0 likes2.4k downloads2y agoHugging Face05willRD /Fineweb-Edu-Chinese-V2.1 Chinese Fineweb Edu Dataset V2.1 [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report The Chinese Fineweb Edu Dataset V2.1 is an enhanced version of the V2 dataset, designed specifically for natural language processing (NLP) tasks in the education sector. This version introduces two new data sources, map-cc and opencsg-cc, and retains data with scores ranging from 2 to 3. The dataset entries are organized into different folders… See the full description on the dataset page: https://huggingface.co/datasets/willRD/Fineweb-Edu-Chinese-V2.1.texttext-generation100M<n<1B0 likes1.9k downloads10mo agoHugging Face06Snowflake /msmarco-v2.1-snowflake-arctic-embed-m-v1.5 Snowflake Arctic Embed M V1.5 Embeddings for MSMARCO V2.1 for TREC-RAG This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG All embeddings are created using Snowflake's Arctic Embed M v1.5 and are intended to serve as a simple baseline for dense retrieval-based methods. It's worth noting that Snowflake's Arctic Embed M v1.5 is optimized for efficient embeddings and thus supports embedding truncation and quantization. More… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-m-v1.5.textquestion-answering10M<n<100M0 likes785 downloads2y agoHugging Face07CohereLabs /m-ArenaHard-v2.1 Dataset Card for m-ArenaHard-v2.1 The m-ArenaHard-v2.1 dataset is a multilingual LLM evaluation set built from the LMarena arena-hard-auto-v2.0 prompts used in m-ArenaHard-v2.0. It keeps the public v2.0 row schema while expanding coverage to the 67 raw translation files produced for the Tiny Aya evaluation work. The dataset includes 67 languages: am, ar, bg, bn, ca, cs, cy, da, de, el, en, es, et, eu, fa, fi, fr, ga, gl, gu, ha, he, hi, hr, hu, id, ig, it, ja, jv, km, ko, lo, lt, lv… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/m-ArenaHard-v2.1.texttext-generation10K<n<100K3 likes400 downloads5mo agoHugging Face08lance-format /ms-marco-v2.1-lance MS MARCO v2.1 QA (Lance Format) A Lance-formatted version of MS MARCO v2.1 — Microsoft's machine-reading-comprehension benchmark built from anonymized Bing query logs. Each row is one user query, the up-to-10 candidate passages Bing retrieved for it with relevance flags, and the human-written reference answers, with MiniLM query embeddings stored inline and pre-built ANN/FTS indices, available directly from the Hub at hf://datasets/lance-format/ms-marco-v2.1-lance/data.… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/ms-marco-v2.1-lance.textquestion-answering100K<n<1M0 likes361 downloads4mo agoHugging Face09HiTZ /composite_corpus_eu_v2.1 Composite dataset for Basque made from public available data This dataset is composed of the following public available data: Train split: The train split is composed of the following datasets combined: mozilla-foundation/common_voice_18_0/eu: "validated" split removing "test_cv" and "dev_cv" split's sentences. (validated split contains official train + dev + test splits and more unique data) gttsehu/basque_parliament_1/eu: "train_clean" split removing some of the… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/composite_corpus_eu_v2.1.audioautomatic-speech-recognition100K<n<1M3 likes349 downloads2y agoHugging Face10grimlee /fineweb-edu-Chinese-v2.1-shuffledata source: https://huggingface.co/datasets/opencsg/Fineweb-Edu-Chinese-V2.1 with nanochat/dev/repackage_data_reference.py text100M<n<1B0 likes349 downloads11mo agoHugging Face11spacemanidol /msmarco-v2.1-stella_en_1.5B_v5 NovaSearch stella_en_1.5B_v5 Embeddings for MSMARCO V2.1 for TREC-RAG This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG All embeddings are created using Stella EN 1.5B V5 and are intended to serve as a simple baseline for dense retrieval-based methods. Note, that the embeddings are not normalized so you will need to normalize them before usage. Retrieval Performance Retrieval performance for the TREC DL21-23… See the full description on the dataset page: https://huggingface.co/datasets/spacemanidol/msmarco-v2.1-stella_en_1.5B_v5.textquestion-answering10M<n<100M0 likes316 downloads1y agoHugging Face12alphagocc /Fineweb-Edu-Chinese-V2.1 Chinese Fineweb Edu Dataset V2.1 [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report The Chinese Fineweb Edu Dataset V2.1 is an enhanced version of the V2 dataset, designed specifically for natural language processing (NLP) tasks in the education sector. This version introduces two new data sources, map-cc and opencsg-cc, and retains data with scores ranging from 2 to 3. The dataset entries are organized into different folders… See the full description on the dataset page: https://huggingface.co/datasets/alphagocc/Fineweb-Edu-Chinese-V2.1.texttext-generation10M<n<100M0 likes215 downloads9mo agoHugging Face13mjbommar /opengloss-v2.1-retrieval-pairs Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility. OpenGloss v2.1 — Retrieval Pairs Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-retrieval-pairs.tabularsentence-similarity10M<n<100M0 likes196 downloads15d agoHugging Face14mjbommar /opengloss-v2.1-examples Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility. OpenGloss v2.1 — Examples Every example sentence in OpenGloss v2.1, one row at a time, each tagged to the sense it illustrates and carrying… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-examples.tabulartext-classification1M<n<10M0 likes169 downloads15d agoHugging Face15mjbommar /opengloss-v2.1-qrels Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility. OpenGloss v2.1 — Qrels A ready-to-score retrieval benchmark built from the release's own graph. The listwise config gives one query with… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-qrels.texttext-retrieval1M<n<10M0 likes163 downloads15d agoHugging Face16TeeZee /common_voice_17-pl-speakers-v2.1audio10K<n<100K0 likes162 downloads6mo agoHugging Face17mjbommar /opengloss-v2.1-relations Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility. OpenGloss v2.1 — Relations The OpenGloss v2.1 semantic graph as an edge list. The relations config holds every live typed edge — fourteen… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-relations.texttext-classification1M<n<10M0 likes160 downloads15d agoHugging Face18Mxode /Fineweb-Edu-Chinese-V2.1-merged-score4_5Fineweb-Edu-Chinese-V2.1 的评分为 4~5 的数据的子集。原数据集切片很细,大约每 10MB 一个切片,本数据集做了集合,每 400 个切片合并为 1 个切片,新的切片每个大小约为 4GB。 为了便于加载,按照切片分割了子集,子集命名来源于原数据集的切片范围,可以指定加载其中的一个子集: from datasets import load_dataset # 加载其中的一个子集 ds = load_dataset("Mxode/Fineweb-Edu-Chinese-V2.1-merged-score4_5", "0-399") 也可以加载多个子集或者全部子集,可以通过如下方式获取子集名称: from datasets import get_dataset_config_names configs = get_dataset_config_names("Mxode/Fineweb-Edu-Chinese-V2.1-merged-score4_5") print(configs) >>> ['0-399', '1200-1599'… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Fineweb-Edu-Chinese-V2.1-merged-score4_5.texttext-generation10M<n<100M4 likes157 downloads1y agoHugging Face19Akhil-Theerthala /Kuvera-PersonalFinance-V2.1 Personal Finance Reasoning-V2.1 This dataset is associated with the paper Synthesizing Behaviorally-Grounded Reasoning Chains: A Data-Generation Framework for Personal Finance LLMs. This is a scaled up version of the PersonalFinance-V2 dataset with some pipeline streamlining done.* 1. Introduction & Motivation The landscape of financial AI benchmarks is currently dominated by applications in corporate finance, algorithmic trading, and general financial knowledge… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/Kuvera-PersonalFinance-V2.1.texttext-classification10K<n<100K8 likes146 downloads9mo agoHugging Face20mjbommar /opengloss-v2.1-definitions Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility. OpenGloss v2.1 — Definitions The flat definition view: one row for every stored rendition of every live sense's definition, the canonical… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-definitions.tabulartext-generation1M<n<10M0 likes128 downloads15d agoHugging Face21Snowflake /msmarco-v2.1-snowflake-arctic-embed-m-v2.0 Snowflake Arctic Embed M V2.0 Embeddings for MSMARCO V2.1 for TREC-RAG This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG All embeddings are created using Snowflake's Arctic Embed M v2.0 and are intended to serve as a simple baseline for dense retrieval-based methods. Note, that the embeddings are not normalized so you will need to normalize them before usage. Retrieval Performance Retrieval performance for… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-m-v2.0.textquestion-answering10M<n<100M0 likes123 downloads1y agoHugging Face22mjbommar /opengloss-v2.1-pretrain Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility. OpenGloss v2.1 — Pretrain The release rendered as continuous prose for language-model pretraining or continued pretraining: four document… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-pretrain.texttext-generation1M<n<10M0 likes122 downloads15d agoHugging Face23mjbommar /opengloss-v2.1-lexicon Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility. OpenGloss v2.1 — Lexicon The entry-level view of OpenGloss v2.1: one row per lexeme, with everything that belongs to the entry rather than… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-lexicon.tabulartext-generation100K<n<1M0 likes118 downloads15d agoHugging Face24mjbommar /opengloss-v2.1-provenance Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility. OpenGloss v2.1 — Provenance The audit trail for OpenGloss v2.1: one row per recorded unit of work, saying which stage ran, which model… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-provenance.tabulartext-classification1M<n<10M0 likes118 downloads15d agoHugging Face25spacemanidol /msmarco-v2.1-gte-large-en-v1.5 Alibaba GTE-Large-V1.5 Embeddings for MSMARCO V2.1 for TREC-RAG This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG All embeddings are created using GTE Large V1.5 and are intended to serve as a simple baseline for dense retrieval-based methods. Note, that the embeddings are not normalized so you will need to normalize them before usage. Retrieval Performance Retrieval performance for the TREC DL21-23… See the full description on the dataset page: https://huggingface.co/datasets/spacemanidol/msmarco-v2.1-gte-large-en-v1.5.textquestion-answering10M<n<100M0 likes117 downloads1y agoHugging Face26mjbommar /opengloss-v2.1-queries Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility. OpenGloss v2.1 — Queries Synthetic search queries written per sense, in eight styles — keyword, question, conversational, constraint, role… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-queries.texttext-retrieval1M<n<10M0 likes116 downloads15d agoHugging Face27mjbommar /opengloss-v2.1-retrieval-triples Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility. OpenGloss v2.1 — Retrieval Triples Training triples for an embedding model or reranker. Each row is a query, a positive passage from the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-retrieval-triples.textsentence-similarity1M<n<10M0 likes115 downloads15d agoHugging Face28mjbommar /opengloss-v2.1-qa-pairs Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility. OpenGloss v2.1 — QA Pairs Question/answer pairs written per sense and answerable only from that sense's own stored text — its gloss, its… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-qa-pairs.textquestion-answering100K<n<1M0 likes111 downloads15d agoHugging Face29mjbommar /opengloss-v2.1-senses Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility. OpenGloss v2.1 — Senses The sense-level view of OpenGloss v2.1 and the repo most consumers want: one row per live sense, with its canonical… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-senses.tabulartext-generation100K<n<1M0 likes108 downloads15d agoHugging Face30mjbommar /opengloss-v2.1-encyclopedia Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility. OpenGloss v2.1 — Encyclopedia The long-form entry-level prose of OpenGloss v2.1, one row per rendition. The encyclopedia config holds the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-encyclopedia.tabulartext-generation100K<n<1M0 likes108 downloads15d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.