datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pentesting-explanations
Pentesting Explanations - Adversarial Reasoning & Vulnerability Research
A high-quality supervised fine-tuning dataset for penetration testing expertise, red team tradecraft, and - as the dataset matures - novel vulnerability research and zero-day reasoning. The dataset is structured to teach models how to think like offensive security practitioners, not merely recall labels or technique names.
The long-term goal of this dataset is to train models capable of genuine adversarial… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/pentesting-explanations.snowball-5.7t-sft-eval-artifacts
Snowball 5.7T cold-start SFT evaluation artifacts
This dataset archives the evaluation records, sampled traces, resolved launch configurations,
analysis inputs, and derived tables for
marin-community/marin#8225.
The experiment compares the 5.7T-token Snowball cooldown with its Chat, Thinking, and
Nemotron-Terminal SFT descendants. It also includes the corresponding 2T-token cooldown cohort.
The top-level EVAL_RESULTS.csv in the experiment record is generated from the durable… See the full description on the dataset page: https://huggingface.co/datasets/penfever/snowball-5.7t-sft-eval-artifacts.penicillin
Penicillin dataset
Paper: https://arxiv.org/abs/2510.25771
Note
This dataset is a processed combination of existing benchmark datasets. The licensing and data rights belong to the original dataset owners. Please refer to the original datasets for more information.
ov-kit-docpenicillin_plus
Penicillin-Plus dataset
Paper: https://arxiv.org/abs/2510.25771
Note
This dataset is a processed combination of existing benchmark datasets. The licensing and data rights belong to the original dataset owners. Please refer to the original datasets for more information.
task149_afs_argument_quality_death_penalty
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task149_afs_argument_quality_death_penalty
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task149_afs_argument_quality_death_penalty.pentesting-explanations
Pentesting Explanations - Adversarial Reasoning & Vulnerability Research
A high-quality supervised fine-tuning dataset for penetration testing expertise, red team tradecraft, and - as the dataset matures - novel vulnerability research and zero-day reasoning. The dataset is structured to teach models how to think like offensive security practitioners, not merely recall labels or technique names.
The long-term goal of this dataset is to train models capable of genuine adversarial… See the full description on the dataset page: https://huggingface.co/datasets/me-aas/pentesting-explanations.PENS
Dataset Card for PENS: PErsonalized News headlineS
PENS is an English dataset for Personalized News Headline Generation Research. It contains two parts for training and test individually. The training set was collected from anonymized user impressions logs of Microsoft News website, and the test set is manually-created by hundreds of native speakers to enable a fair testbed for evaluating models in an offline mode.
PENS contains about 113k English news articles whose topics are… See the full description on the dataset page: https://huggingface.co/datasets/THEATLAS/PENS.medium-web-pentesting
Medium Web Pentesting Articles
Dataset Description
A curated collection of 357 Medium articles focused on web penetration testing, scraped from Medium's search results for the query web pentesting. Each record includes article metadata and the opening snippet of the article body.
This dataset is useful for NLP tasks such as topic modeling, text classification, content recommendation, and summarization within the cybersecurity domain.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/shaikat005/medium-web-pentesting.Synthetic_PenTest_ReportsThe full CJ Jones' synthetic dataset catalog is available at:
https://datadeveloper1.gumroad.com
Want more? 🚀 Get the AI Startup Bundle from Gumroad.
📄 100 Samples of Synthetic Automated Penetration Test Reports
This dataset contains 100+ realistic, synthetic penetration testing reportsstructured to simulate professional internal security assessments. Each record models the full flow of a pentest engagement, including:
Reconnaissance / Discovery Phase
Vulnerability Assessment… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Synthetic_PenTest_Reports.pentestr1-chat
Pentest-R1 Chat — TCO Fine-Tuning Dataset
Converted from KHenryAegis/Pentest-R1 into Gemma chat-template JSONL for fine-tuning autonomous penetration testing agents with Unsloth Studio.
Dataset summary
Field
Value
Conversations
535
Total messages
28,731
Format
JSONL — one {"messages": [...]} per line
Chat template
Gemma (unsloth/gemma-3n-E4B-it)
Token p50 / p95 / p99 / max
2,988 / 6,111 / 7,843 / 8,178
Recommended context_length
8192… See the full description on the dataset page: https://huggingface.co/datasets/supersamdev/pentestr1-chat.task145_afs_argument_similarity_death_penalty
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task145_afs_argument_similarity_death_penalty
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task145_afs_argument_similarity_death_penalty.Mille-Pensees-Dataset
Mille-Pensées-Dataset
Dataset Summary
The Mille-Pensées-Dataset is a math reasoning dataset with a 50% french / 50% english composition used to train the Mille-Pensées french reasoning model.
The source data comes from the following english math reasoning datasets:
s1K-1.1
OpenThoughts3-1.2M
OpenR1-Math-220k
OpenMathReasoning
Nemotron-Post-Training-Dataset-v1
LIMO-v2
DeepMath-103K
AM-DeepSeek-R1-0528-Distilled
The french reasoning data was obtained by translating the… See the full description on the dataset page: https://huggingface.co/datasets/GLauzza/Mille-Pensees-Dataset.mirror-pentesting-explanations
Pentesting Explanations - Adversarial Reasoning & Vulnerability Research
A high-quality supervised fine-tuning dataset for penetration testing expertise, red team tradecraft, and - as the dataset matures - novel vulnerability research and zero-day reasoning. The dataset is structured to teach models how to think like offensive security practitioners, not merely recall labels or technique names.
The long-term goal of this dataset is to train models capable of genuine adversarial… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-pentesting-explanations.task1167_penn_treebank_coarse_pos_tagging
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1167_penn_treebank_coarse_pos_tagging
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1167_penn_treebank_coarse_pos_tagging.code-pensions-civiles-militaires-retraite
Code des pensions civiles et militaires de retraite, non-instruct (2025-03-10)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-pensions-civiles-militaires-retraite.Pentesting_Datasetpentest
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/mikoube/pentest.code-penal
Code pénal, non-instruct (2025-05-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models based… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-penal.PreferenceTravelPlanner
PreferenceTravelPlanner Dataset
PreferenceTravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints and preferences. For more details, see our paper. It is created by augmenting TravelPlanner (See paper for more details) with several common type of preferences under various preference paradigms.
Introduction
In PreferenceTravelPlanner, for a given query, language agents are expected to formulate a… See the full description on the dataset page: https://huggingface.co/datasets/pensieves/PreferenceTravelPlanner.code-procedure-penale
Code de procédure pénale, non-instruct (2025-03-10)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-procedure-penale.code-justice-penale-mineurs
Code de la justice pénale des mineurs, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-justice-penale-mineurs.anonymized_data
PencilCode Program Traces
Paper: Modeling Student Learning with 3.8 Million Program TracesCode: meghabyte/pencilcode-publicContact: megha@cs.stanford.edu, alexisro@mit.edu, jjb@eng.ufl.edu, jda@mit.edu
Gated Dataset — Manual Approval Required.Access is granted pending review by the Pencil Code team. Please submit your request above with a brief description of your intended research use. Any requests will be reviewed by Jeremiah Blanchard at jjb@eng.ufl.edu
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/pencil-code/anonymized_data.cantonesewiki_doyouknow
Cantonese Question Dataset from Yue Wiki
A collection of questions in Cantonese, extracted from Yue Wiki. This dataset contains a variety of questions covering different topics and domains.
Disclaimer
The content and opinions expressed in this dataset do not represent the views, beliefs, or positions of the dataset creators, contributors, or hosting organizations. This dataset is provided solely for the purpose of improving AI systems' understanding of the Cantonese… See the full description on the dataset page: https://huggingface.co/datasets/pendingremove32894/cantonesewiki_doyouknow.food-penguin-v1
Food Penguin v1
Food Penguin v1 is a compact, domain-specific text-generation dataset from the SakThai family. It provides conversation turns paired with tool definitions and expected function calls in JSON format, optimized for instruction-following, food-domain QA, and lightweight tool-use model training.
Dataset Summary
Food Penguin v1 is designed for:
Instruction-following models
Food & restaurant domain QA
Tool-use and function-calling training… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/food-penguin-v1.offences_and_penalties_in_general_2018_datasetpentadrive-v1
TrueHuman PentaDrive
Dataset summary
TrueHuman PentaDrive is a compact behavioral matrix for affective and conversational modeling. It organizes human-relevant motivational structure into five drives (coded S, K, A, M, G), each implemented as a set of nodes (stable behavioral motifs). Every node defines three phases—anticipation, release, and block—with:
Markers: short textual cues for lightweight classification or retrieval
Response kernels: structured hints for… See the full description on the dataset page: https://huggingface.co/datasets/datamarketinglabs/pentadrive-v1.MM-MathInstruct-longest-20k-solutions-with-images
MM-MathInstruct Longest 20K Solutions with Images
This dataset contains the top 20,000 samples from MathLLMs/MM-MathInstruct selected by solution length, filtered to include only samples with valid images.
Dataset Structure
Usage
Source
This dataset is derived from MathLLMs/MM-MathInstruct by selecting the 20,000 samples with the longest solution text.
License
Apache 2.0 (inherited from source dataset)
code-disciplinaire-penal-marine-marchande
Code disciplinaire et pénal de la marine marchande, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-disciplinaire-penal-marine-marchande.code-penitentiaire
Code pénitentiaire, non-instruct (2025-03-10)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-penitentiaire.
