datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-CC-v2.1
Nemotron-Pre-Training-Dataset-v2.1
Dataset Description
The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web tokens… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.1.Cybersecurity-Dataset-Fenrir-v2.1
Cybersecurity Defense Instruction-Tuning Dataset (v2.1)
Created by Alican Kiraz
TL;DR
A ready-to-train dataset of 99,870 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training.
Apache-2.0 licensed and production-ready.
Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security.
1 What’s… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Cybersecurity-Dataset-Fenrir-v2.1.msmarco-v2.1-embed-english-v3
TREC-RAG 2024 Corpus (MSMARCO 2.1) - Encoded with Cohere Embed English v3
This dataset contains the embeddings for the TREC-RAG Corpus 2024 embedded with the Cohere Embed V3 English model.
It contains embeddings for 113,520,750 passages, embeddings for 1677 queries from TREC-Deep Learning 2021-2023, as well as top-1000 hits for all queries using a brute-force (flat) index.
Search over the Index
We have a pre-build index that only requires 300 MB available at… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/msmarco-v2.1-embed-english-v3.msmarco-v2.1-snowflake-arctic-embed-l
Snowflake Arctic Embed L Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using Snowflake's Arctic Embed L and are intended to serve as a simple baseline for dense retrieval-based methods.
Retrieval Performance
Retrieval performance for the TREC DL21-23, MSMARCOV2-Dev and Raggy Queries can be found below with BM25 as a baseline. For both… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-l.Fineweb-Edu-Chinese-V2.1
Chinese Fineweb Edu Dataset V2.1 [中文] [English]
[OpenCSG Community] [👾github] [wechat] [Twitter]
📖Technical Report
The Chinese Fineweb Edu Dataset V2.1 is an enhanced version of the V2 dataset, designed specifically for natural language processing (NLP) tasks in the education sector. This version introduces two new data sources, map-cc and opencsg-cc, and retains data with scores ranging from 2 to 3. The dataset entries are organized into different folders… See the full description on the dataset page: https://huggingface.co/datasets/willRD/Fineweb-Edu-Chinese-V2.1.msmarco-v2.1-snowflake-arctic-embed-m-v1.5
Snowflake Arctic Embed M V1.5 Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using Snowflake's Arctic Embed M v1.5 and are intended to serve as a simple baseline for dense retrieval-based methods.
It's worth noting that Snowflake's Arctic Embed M v1.5 is optimized for efficient embeddings and thus supports embedding truncation and quantization. More… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-m-v1.5.m-ArenaHard-v2.1
Dataset Card for m-ArenaHard-v2.1
The m-ArenaHard-v2.1 dataset is a multilingual LLM evaluation set built from the LMarena arena-hard-auto-v2.0 prompts used in m-ArenaHard-v2.0. It keeps the public v2.0 row schema while expanding coverage to the 67 raw translation files produced for the Tiny Aya evaluation work.
The dataset includes 67 languages:
am, ar, bg, bn, ca, cs, cy, da, de, el, en, es, et, eu, fa, fi, fr, ga, gl, gu, ha, he, hi, hr, hu, id, ig, it, ja, jv, km, ko, lo, lt, lv… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/m-ArenaHard-v2.1.ms-marco-v2.1-lance
MS MARCO v2.1 QA (Lance Format)
A Lance-formatted version of MS MARCO v2.1 — Microsoft's machine-reading-comprehension benchmark built from anonymized Bing query logs. Each row is one user query, the up-to-10 candidate passages Bing retrieved for it with relevance flags, and the human-written reference answers, with MiniLM query embeddings stored inline and pre-built ANN/FTS indices, available directly from the Hub at hf://datasets/lance-format/ms-marco-v2.1-lance/data.… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/ms-marco-v2.1-lance.composite_corpus_eu_v2.1
Composite dataset for Basque made from public available data
This dataset is composed of the following public available data:
Train split:
The train split is composed of the following datasets combined:
mozilla-foundation/common_voice_18_0/eu: "validated" split removing "test_cv" and "dev_cv" split's sentences. (validated split contains official train + dev + test splits and more unique data)
gttsehu/basque_parliament_1/eu: "train_clean" split removing some of the… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/composite_corpus_eu_v2.1.fineweb-edu-Chinese-v2.1-shuffledata source: https://huggingface.co/datasets/opencsg/Fineweb-Edu-Chinese-V2.1
with nanochat/dev/repackage_data_reference.py
msmarco-v2.1-stella_en_1.5B_v5
NovaSearch stella_en_1.5B_v5 Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using Stella EN 1.5B V5 and are intended to serve as a simple baseline for dense retrieval-based methods.
Note, that the embeddings are not normalized so you will need to normalize them before usage.
Retrieval Performance
Retrieval performance for the TREC DL21-23… See the full description on the dataset page: https://huggingface.co/datasets/spacemanidol/msmarco-v2.1-stella_en_1.5B_v5.Fineweb-Edu-Chinese-V2.1
Chinese Fineweb Edu Dataset V2.1 [中文] [English]
[OpenCSG Community] [👾github] [wechat] [Twitter]
📖Technical Report
The Chinese Fineweb Edu Dataset V2.1 is an enhanced version of the V2 dataset, designed specifically for natural language processing (NLP) tasks in the education sector. This version introduces two new data sources, map-cc and opencsg-cc, and retains data with scores ranging from 2 to 3. The dataset entries are organized into different folders… See the full description on the dataset page: https://huggingface.co/datasets/alphagocc/Fineweb-Edu-Chinese-V2.1.opengloss-v2.1-retrieval-pairs
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Retrieval Pairs
Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-retrieval-pairs.opengloss-v2.1-examples
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Examples
Every example sentence in OpenGloss v2.1, one row at a time, each tagged to the sense it illustrates and carrying… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-examples.opengloss-v2.1-qrels
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Qrels
A ready-to-score retrieval benchmark built from the release's own graph. The listwise config gives one query with… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-qrels.common_voice_17-pl-speakers-v2.1opengloss-v2.1-relations
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Relations
The OpenGloss v2.1 semantic graph as an edge list. The relations config holds every live typed edge — fourteen… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-relations.Fineweb-Edu-Chinese-V2.1-merged-score4_5Fineweb-Edu-Chinese-V2.1 的评分为 4~5 的数据的子集。原数据集切片很细,大约每 10MB 一个切片,本数据集做了集合,每 400 个切片合并为 1 个切片,新的切片每个大小约为 4GB。
为了便于加载,按照切片分割了子集,子集命名来源于原数据集的切片范围,可以指定加载其中的一个子集:
from datasets import load_dataset
# 加载其中的一个子集
ds = load_dataset("Mxode/Fineweb-Edu-Chinese-V2.1-merged-score4_5", "0-399")
也可以加载多个子集或者全部子集,可以通过如下方式获取子集名称:
from datasets import get_dataset_config_names
configs = get_dataset_config_names("Mxode/Fineweb-Edu-Chinese-V2.1-merged-score4_5")
print(configs)
>>> ['0-399', '1200-1599'… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Fineweb-Edu-Chinese-V2.1-merged-score4_5.Kuvera-PersonalFinance-V2.1
Personal Finance Reasoning-V2.1
This dataset is associated with the paper Synthesizing Behaviorally-Grounded Reasoning Chains: A Data-Generation Framework for Personal Finance LLMs.
This is a scaled up version of the PersonalFinance-V2 dataset with some pipeline streamlining done.*
1. Introduction & Motivation
The landscape of financial AI benchmarks is currently dominated by applications in corporate finance, algorithmic trading, and general financial knowledge… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/Kuvera-PersonalFinance-V2.1.opengloss-v2.1-definitions
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Definitions
The flat definition view: one row for every stored rendition of every live sense's definition, the canonical… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-definitions.msmarco-v2.1-snowflake-arctic-embed-m-v2.0
Snowflake Arctic Embed M V2.0 Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using Snowflake's Arctic Embed M v2.0 and are intended to serve as a simple baseline for dense retrieval-based methods.
Note, that the embeddings are not normalized so you will need to normalize them before usage.
Retrieval Performance
Retrieval performance for… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-m-v2.0.opengloss-v2.1-pretrain
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Pretrain
The release rendered as continuous prose for language-model pretraining or continued pretraining: four document… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-pretrain.opengloss-v2.1-lexicon
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Lexicon
The entry-level view of OpenGloss v2.1: one row per lexeme, with everything that belongs to the entry rather than… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-lexicon.opengloss-v2.1-provenance
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Provenance
The audit trail for OpenGloss v2.1: one row per recorded unit of work, saying which stage ran, which model… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-provenance.msmarco-v2.1-gte-large-en-v1.5
Alibaba GTE-Large-V1.5 Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using GTE Large V1.5 and are intended to serve as a simple baseline for dense retrieval-based methods.
Note, that the embeddings are not normalized so you will need to normalize them before usage.
Retrieval Performance
Retrieval performance for the TREC DL21-23… See the full description on the dataset page: https://huggingface.co/datasets/spacemanidol/msmarco-v2.1-gte-large-en-v1.5.opengloss-v2.1-queries
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Queries
Synthetic search queries written per sense, in eight styles — keyword, question, conversational, constraint, role… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-queries.opengloss-v2.1-retrieval-triples
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Retrieval Triples
Training triples for an embedding model or reranker. Each row is a query, a positive passage from the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-retrieval-triples.opengloss-v2.1-qa-pairs
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — QA Pairs
Question/answer pairs written per sense and answerable only from that sense's own stored text — its gloss, its… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-qa-pairs.opengloss-v2.1-senses
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Senses
The sense-level view of OpenGloss v2.1 and the repo most consumers want: one row per live sense, with its canonical… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-senses.opengloss-v2.1-encyclopedia
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Encyclopedia
The long-form entry-level prose of OpenGloss v2.1, one row per rendition. The encyclopedia config holds the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-encyclopedia.
