datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cis_aws_foundation_benchmark_5_0
CIS AWS Foundations Benchmark v5.0.0 - Security Compliance Dataset
Description
This dataset contains high-quality instruction-following data derived from the Center for Internet Security (CIS) Amazon Web Services Foundations Benchmark v5.0.0 (released March 31, 2025).
It is designed to fine-tune Large Language Models (LLMs) to act as Cloud Security Auditors and Consultants. The dataset covers prescriptive guidance for establishing a secure baseline configuration for AWS… See the full description on the dataset page: https://huggingface.co/datasets/halencarjunior/cis_aws_foundation_benchmark_5_0.RelEx-PT
RelEx-PT
Dataset Summary
RelEx-PT, a new sentence-level Relation Extraction dataset for Portuguese. Addressing the scarcity of high-quality, controlled resources for the language, RelEx-PT provides a balanced benchmark comprising 18 Wikidata-derived relation types across diverse domains. The dataset is built through a distant supervision pipeline that links Wikidata triples with Portuguese Wikipedia sentences and enhanced by a Natural Language Inference (NLI)-based… See the full description on the dataset page: https://huggingface.co/datasets/NLP-CISUC/RelEx-PT.CISA_Enrichment
CISA Known Exploited Vulnerabilities Catalog Enrichment
The CISA recently started to publish the Known Exploited Vulnerabilities Catalog Enrichment to help federal agencies keep up with exploited vulnerabilities.
The data they provide is minimal, so I have built this jupyter notebook to enrich the data using the CIRCL public CVE API to add the following data points:
CWE
CVE Published Date
CVE Modified Date
Reference URLs
CPE 2.3 Data
A Github Action runs every 6 hours and updates… See the full description on the dataset page: https://huggingface.co/datasets/cvelist/CISA_Enrichment.cisnes-b3CISA-testsetCISA-testset from İbrahim Temo's Memoir: Cross-Individual Sentiment Analysis Test Dataset for Historical Turkish
This test dataset is specifically designed for evaluating Cross-Individual Sentiment Analysis (CISA) performance on historical Turkish texts from İbrahim Temo's memoirs.
📚 Dataset Description
This test dataset contains 200 sentences extracted from the first 66 pages of İbrahim Temo's original memoirs "İttihad ve Terakki Cemiyetinin Teşekkülü ve Hidematı Vataniye ve İnkılâbı Milliye… See the full description on the dataset page: https://huggingface.co/datasets/dbbiyte/CISA-testset.cisco_cli_commandsYoutubeCommentsForProjectdefendable-pain-cisa-advisory-pain-v0.1
CISA Advisory Pain Receipt
"the advisory" — Mr. Defendable
A free pain-receipt dataset from the DefendableOS ecosystem. 4 rows · ready to read · all cited or graded · CC-BY-4.0.
Part of the 100-pack — 100 free pain-receipt datasets dropped from the Defendable Bakery to the open AI-trust community. Different theme per dataset. Same operator voice across all of them.
Tribunal begins before training. No proof, no honey. To the shed.
What's in here
4 pain receipts… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-cisa-advisory-pain-v0.1.cisco_catalyst_9800_wireless_controller_cli_commandshomesecurityzonesSECHOMEold-datasetYoutubeCommentSentimentFinalcisco_system_messageCISS_TRAINtest-oldcisdemoCIS_4190_Final_Project
