datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Retrieval-Synthetic-NVDocs-v1
Dataset Description:
Retrieval-Synthetic-NVDocs-v1 is a synthetic retrieval dataset with question–answer supervision designed to train and evaluate embedding and RAG systems. The dataset was generated on top of NVIDIA's publicly available content using NeMo Data Designer, NVIDIA's open-source framework for generating high-quality synthetic data from scratch or based on seed data.
The dataset contains document chunks paired with semantically rich question-answer pairs across multiple… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Retrieval-Synthetic-NVDocs-v1.OWASP-and-NVD-question-answer-datasetnvd-security-instructions
NVD Security Instructions
Plain English CVE analysis pairs for fine-tuning security-domain LLMs.
Dataset Summary
2,063 instruction-response pairs built from NVD CVE data (2023-2024).
Each pair contains a raw CVE as input and a structured plain English
analysis as output — designed to teach LLMs to explain vulnerabilities
to developers with no security background.
Dataset Structure
Each record contains:
text: Mistral instruction format string ready for… See the full description on the dataset page: https://huggingface.co/datasets/AnamayaVyas/nvd-security-instructions.nvda-option-chainsCVE_NVD
CVE_NVD
A curated dataset of 200 CVE vulnerability instances from the National Vulnerability Database (NVD), enriched with exploit PoCs, developer patches, bug reports, sanitizer reports, and GitHub issue context.
Dataset Fields
Field
Description
instance_id
Unique identifier (e.g., njs.cve-2022-32414)
description
CVE description from NVD
exploit(poc)_url
URL to exploit or proof-of-concept
patch_url
URL to the developer's patch commit… See the full description on the dataset page: https://huggingface.co/datasets/SongTonyLi/CVE_NVD.NvdiaCode_10nvda_1m_2022_to_2026stocks_one_nvda
Dataset Card for "stocks_one_nvda"
More Information needed
nvda-stock-prices
NVDA Stock Price History
Daily historical stock price data for NVIDIA Corporation (NVDA),
retrieved from Yahoo Finance via the yfinance Python library.
Columns
Date — trading date
Open — opening price ($)
High — intraday high ($)
Low — intraday low ($)
Close — closing price ($), split-adjusted
Volume — shares traded that day
Coverage
Daily data from NVIDIA's IPO through the present, trading days only
(weekends and market holidays excluded).… See the full description on the dataset page: https://huggingface.co/datasets/J-Marcos-GS/nvda-stock-prices.stock_nvda
Dataset Card for "stock_nvda"
More Information needed
nvd-cve-scraper
NVD CVE Scraper · Vulnerabilities, CVSS Scores, Vendors & CWEs
Scrape National Vulnerability Database (NVD) CVE records, CVSS v2/v3/v4 severity scores, CWE weakness classifications, vendor products, and exploit references.
Rows in this dataset
1,403
Fields
19
Collector runs behind it
36
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/nvd-cve-scraper.NVDLibraryBenchmarkNvdiaCode_9NvdiaCode_QW32B_15embedding-cve-nvd-dataset
CVE NVD Embedding Dataset
This dataset contains the processed CVE/NVD corpus that was used with the rag_mixedbread pipeline.
It bundles:
cve_corpus.jsonl (~700 MB): each line is a JSON object with cve_id, title, description, cvss, vendors, and the pre-computed text chunk that feeds the embedding model.
decomposed_query_results.json (63 KB): a dictionary of exemplar queries, decomposed sub-questions, and the retrieved doc IDs used for quality checks.
Generation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/Kushalkhemka/embedding-cve-nvd-dataset.nvd-cve-rag-dataset
NVD CVE RAG Dataset
Dataset Summary
This dataset contains cleaned NVD CVE records and synthetic query-document pairs prepared for bi-encoder fine-tuning and retrieval evaluation in a RAG pipeline.
It includes:
A cleaned CVE document corpus used as the retrieval corpus.
Synthetic training pairs for bi-encoder fine-tuning.
Synthetic validation pairs for validation loss monitoring and early stopping.
Synthetic test pairs for retrieval evaluation before and after… See the full description on the dataset page: https://huggingface.co/datasets/ubco-mds-2025-capstone-fujitsu-1/nvd-cve-rag-dataset.RLVR-Env-Retrieval-Source-Retrieval-Synthetic-NVDocs-v1
RLVR-Env-Retrieval-Source-Retrieval-Synthetic-NVDocs-v1
RLVR-ready retrieval environment derived from nvidia/Retrieval-Synthetic-NVDocs-v1.
Author: Aman Priyanshu
What Is This
A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-Retrieval-Synthetic-NVDocs-v1.NvdiaOpenInstructCode_1NvdiaCode_QW32B_16NvdiaCode_2NvdiaCode_QW32B_6NvdiaCode_QW32B_19NVDA
Dataset Card for "NVDA"
More Information needed
NvdiaCode_11stocks_one_nvda_v3_weekly
Dataset Card for "stocks_one_nvda_v3_weekly"
More Information needed
NvdiaOpenLowCode_3stocks_one_nvda_v2
Dataset Card for "stocks_one_nvda_v2"
More Information needed
nvda-2023-01-01_2024-08-24NvdiaOpenInstructCode_20NvdiaOpenInstructCode_19
