datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WavCaps
WavCaps
WavCaps is a ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research, where the audio clips are sourced from three websites (FreeSound, BBC Sound Effects, and SoundBible) and a sound event detection dataset (AudioSet Strongly-labelled Subset).
Paper: https://arxiv.org/abs/2303.17395
Github: https://github.com/XinhaoMei/WavCaps
Statistics
Data Source
# audio
avg. audio duration (s)avg. text length
FreeSound… See the full description on the dataset page: https://huggingface.co/datasets/cvssp/WavCaps.converted_cvssOmniCount-191
OmniCount-191
A comprehensive benchmark for multi-label object counting, introduced in OmniCount: Multi-label Object Counting with Semantic-Geometric Priors (AAAI 2025).
Dataset Description
OmniCount-191 consists of 30,230 images with multi-label object counts, including point, bounding box, and VQA annotations across 191 object categories.
Paper: arXiv:2403.05435
Project Page: mondalanindya.github.io/OmniCount
Code: github.com/mondalanindya/OmniCount… See the full description on the dataset page: https://huggingface.co/datasets/cvssp/OmniCount-191.cvssCVSS is a massively multilingual-to-English speech-to-speech translation corpus,
covering sentence-level parallel speech-to-speech translation pairs from 21
languages into English.cvss\
CVSS is a massively multilingual-to-English speech-to-speech translation corpus,
covering sentence-level parallel speech-to-speech translation pairs from 21
languages into English.vulnerability-scores-cvss-v3
Vulnerability scores (CVSS v3 combined)
A labeled slice of CIRCL/vulnerability-scores for training and evaluating models that predict CVSS v3 severity from a vulnerability description.
Every row has a combined v3 score and a severity band. Rows with no v3.1 or v3.0 score were dropped.
What changed from the original
The CIRCL dataset stores four separate CVSS columns (cvss_v4_0, cvss_v3_1, cvss_v3_0, cvss_v2_0). Those versions are not on the same scale, so this… See the full description on the dataset page: https://huggingface.co/datasets/AgileRLArena/vulnerability-scores-cvss-v3.cvss-tcvss_method2cve_cwe_cvsscvss_method1cvss_t_zh_en_v1.0cvss_base_score-data
📦 CVSS Score Prediction Dataset
This dataset is designed for training and evaluating language models to predict CVSS base scores from vulnerability descriptions and metadata. The data is extracted from NVD (National Vulnerability Database) CVE records and processed into prompt–score pairs suitable for supervised learning.
📁 Dataset Structure
The dataset is split into three sets by CVE publication year:
Split
CVE Years Included
train
2016–2021
validation… See the full description on the dataset page: https://huggingface.co/datasets/drorrabin/cvss_base_score-data.cvss-c-datasetcvss-v4-qa
CVSS v4.0 Dataset
A high-quality question–answer dataset of 100 records focused on the Common
Vulnerability Scoring System (CVSS) Version 4.0. It is built to train and evaluate AI
systems that explain CVSS scoring methodology, interpret vector strings, justify scoring
decisions, evaluate vulnerability severity, and help analysts produce consistent,
defensible vulnerability assessments — LLM fine-tuning, retrieval-augmented generation
(RAG), vulnerability-management assistants… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/cvss-v4-qa.CVE_NORMALIZED_DESCRIPTION_CVSS_MAPPINGcvss-mm-method1WebAttack-CVSSMetrics
Web Access Logs Dataset
Overview
This dataset contains web access logs, including information about various types of cyber attacks. It includes details on the type of attack (if any), the CVSS (Common Vulnerability Scoring System) metrics representing the risk of the attack, and a score indicating the percentage of risk.
Dataset Details
Number of rows: 18,842
Number of distinct types of attacks: 7
Total number of tokens: 420,098
Columns
Type: The… See the full description on the dataset page: https://huggingface.co/datasets/chYassine/WebAttack-CVSSMetrics.cve_cwe_cvss_2026_refurbishedcve_cwe_cvss-test_ds_2026cvss-ttsCVSSShareGPT75cve_cwe_cvss_refurbishedcve_cwe_cvss_refurbished_missing_dataCVSSCVSSall credits to: https://huggingface.co/datasets/AI4Sec/cti-bench
CVSSShareGPT100SpaGBOLcvss-tts-descs2st_cvss_mlcvss-c
