datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BoilingBench-CV
BoilingBench-CV Dataset
Version: v0.1.0
Maintainer: NED3 Laboratory, University of Arkansas
License: CC BY 4.0
DOI: 10.5281/zenodo.22264378
Mirror of the Zenodo deposit of 3 September 2026, published here because most
users of these data work in the Hugging Face ecosystem. The file set was
verified identical to the deposit at upload time: 7,147 files, 4.20 GB
uncompressed.
Authors
Hari Pandey (University of Arkansas), Manohar Bongarala (Purdue University),
Christy… See the full description on the dataset page: https://huggingface.co/datasets/UARK-NED3/BoilingBench-CV.cve-and-cwe-dataset-1999-2025This collection brings together every Common Vulnerabilities & Exposures (CVE) entry published in the National Vulnerability Database (NVD) from the very first identifier — CVE-1999-0001 — through all records available on 30 May 2025.
It was built automatically with a Python script that calls the NVD REST API v2.0 page-by-page, handles rate-limits, and filters data.
After download each CVE object is pared down to the essentials and written to CVE_CWE_2025.csv with the following columns:… See the full description on the dataset page: https://huggingface.co/datasets/stasvinokur/cve-and-cwe-dataset-1999-2025.VLMEvalKit_CVQA
CVQA for VLMEvalKit
Original dataset: ported to VLMEvalKit
From the original authors:
CVQA is a culturally diverse multilingual VQA benchmark consisting of over 10,000 questions from 39 country-language pairs. The questions in CVQA are written in both the native languages and English, and are categorized into 10 diverse categories.
{'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=2048x1536 at 0x7C3E0EBEEE00>,
'ID': '5919991144272485961_0',
'Subset':… See the full description on the dataset page: https://huggingface.co/datasets/timothycdc/VLMEvalKit_CVQA.cvefixes_bigvulaikyatansinha_cybersecurity-cves-for-nlp-dataset
Cybersecurity CVEs for NLP Dataset
Every CVE since 1999, scrubbed and perfectly formatted for NLP tasks
Dataset Info
Source: Kaggle
Original Size: 38.28 MB
Kaggle Downloads: 36
Files: 1
Files
NVD_Cybersecurity_Dataset.csv
Mirrored from Kaggle
cve-allafrica-synth-hypertension-hypertension-cvd-dataset-all
African Hypertension & CVD Synthetic Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-hypertension-hypertension-cvd-dataset-all.stock-technical-indicators
Stock Technical Indicators Dataset
Historical technical indicators dataset used to train directional stock movement classifiers.
Features
RSI: Relative Strength Index
SMA_20 / EMA_50: Simple and Exponential Moving Averages
MACD: Moving Average Convergence Divergence
Target: Directional label (1 = Bullish, 0 = Bearish)
CVE_CWE_Software_Mapping_Dataset
CVE-CWE Software Weakness Mapping Dataset
Dataset description
This dataset maps Common Vulnerabilities and Exposures (CVEs) to Common Weakness Enumeration (CWE) entries in the CWE-699 Software category. It combines CVE descriptions with CWE descriptions and parent-category information for security research and vulnerability classification.
Dataset structure
The dataset is provided as Global_Dataset.csv. Its main fields include:
CVE-ID: CVE… See the full description on the dataset page: https://huggingface.co/datasets/regularpooria/CVE_CWE_Software_Mapping_Dataset.Q20LLM
20 Questions game with LLM
This dataset generated with the following LLMs:
Groq API / llama3-70b-8192
Groq API / mixtral-8x7b-32768
Mistral API / mistral-large-latest
Test keywords based on newlist_things.rmdup.test.txt from Entity-Deduction Arena (EDA) project.
The dataset generated in two stages:
LLM was prompt to generate keywords
Dialog with different length generated for each keyword
Keywords prompt
Generate a list of 500 diverse and simple keywords suitable… See the full description on the dataset page: https://huggingface.co/datasets/cvmistralparis/Q20LLM.cve-and-cwe-mapping-dataset
CVE and CWE Mapping Dataset
This Hugging Face dataset is a partial copy of the 'CVE and CWE mapping Dataset (2021)' from Kaggle, featuring 'Global_Dataset.csv' originally as 'Global_Dataset.xlsx'. Created by Kirushikesh DB and shared under CC BY-NC-SA 4.0, it includes CVE data up to 2021 for cybersecurity research. For full details and licensing, visit the original Kaggle page.
For further information, please review the CVE Terms of Use and the NVD Terms of Use.
CV_BenchBased on nyu-visionx/CV-Bench.
Liberty-CV
LIBERTy-CV Dataset
Overview
LIBERTy-CV is one of the three datasets released as part of the LIBERTy (LLM-based Interventional Benchmark for Explainability with Real Targets) benchmark.
The goal of LIBERTy is to evaluate concept-based explanation methods in NLP under a causal and counterfactual framework.Each dataset in the benchmark is designed to expose spurious correlations between high-level concepts and model predictions, and to enable quantitative evaluation of… See the full description on the dataset page: https://huggingface.co/datasets/GilatToker/Liberty-CV.arabic-cv-scoring-dataset
Arabic CV Scoring Dataset
Dataset Summary
This dataset contains ~7,220 synthetically generated Arabic CVs, each paired
with a job category, an ATS (Applicant Tracking System) compatibility score,
and a suitability score/class label. It was built to train and evaluate the
Arabic CV Analyzer —
an NLP pipeline that scores, classifies, and generates improvement suggestions
for Arabic CVs targeting the Arab job market, where no equivalent
ATS-optimization tooling… See the full description on the dataset page: https://huggingface.co/datasets/omaraboelmaaty/arabic-cv-scoring-dataset.bitcoin-ethereum-orderflow-cvd-alpha
Bitcoin & Ethereum 1-Minute Order Flow & Cumulative Volume Delta (CVD) Alpha
Institutional Market Microstructure Dataset Sample (Clean CSV / Parquet Ready)
📌 Dataset Overview
In cryptocurrency and traditional electronic markets, price action is driven by aggressive market orders (taker flow) that cross the bid-ask spread. This preview dataset provides 1,000 rows of continuous 1-minute order flow for Bitcoin (BTC/USDT) and Ethereum (ETH/USDT)… See the full description on the dataset page: https://huggingface.co/datasets/TechPlayground/bitcoin-ethereum-orderflow-cvd-alpha.C-VQAThe dataset repo contains the data for C-VQA-Real dataset, for complete data and evaluating your model on our dataset, please refer to https://github.com/Letian2003/C-VQA.
CVPR2024-papersTG_CVECVECPEAPIBenchmarkRareFace-50
RareFace-50 (from Low-Rank Head Avatar Personalization with Registers)
Dataset for Low-Rank Head Avatar Personalization with Registers. Also available on arxiv.
Project Page
Dataset Summary
RareFace-50 is a curated collection of challenging human faces intended for evaluating personalization of talking-head and avatar generation methods.
Unlike many existing face video datasets that focus primarily on celebrities and well-known public figures (e.g., television… See the full description on the dataset page: https://huggingface.co/datasets/StonyBrook-CVLab/RareFace-50.cvss_t_zh_en_v1.0cve-2-att-ckPreprocessed-CVS-24-KMRcyberscale-training-cves
CyberScale Training CVEs
Training dataset for the CyberScale vulnerability severity scorer. Contains 30,641 CVEs with CVSS v3.x scores, descriptions, and CWE classifications.
Schema
Column
Type
Description
cve_id
string
CVE identifier (e.g., CVE-2024-1234)
description
string
Vulnerability description (English)
cvss_score
float
CVSS v3.x base score (0.0-10.0)
cvss_version
string
CVSS version (3.0 or 3.1)
cwe
string
CWE identifier (e.g., CWE-79), may be… See the full description on the dataset page: https://huggingface.co/datasets/eromang/cyberscale-training-cves.Preprocessed-CVS-24-CKBcve-and-cwe-mapping-dataset
CVE and CWE Mapping Dataset
This Hugging Face dataset is a partial copy of the 'CVE and CWE mapping Dataset (2021)' from Kaggle, featuring 'Global_Dataset.csv' originally as 'Global_Dataset.xlsx'. Created by Kirushikesh DB and shared under CC BY-NC-SA 4.0, it includes CVE data up to 2021 for cybersecurity research. For full details and licensing, visit the original Kaggle page.
For further information, please review the CVE Terms of Use and the NVD Terms of Use.
aims-traffic-cv-dataanesthesia-coding-31.5krussian_eval_data_cvs_parquetcvTrpChannel-5-Test
