datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Penguin-Recap-I
Penguin-Recap-I
Penguin-Recap-I publishes recap metadata only. The repository does not contain
image binaries.
Included subsets
subset
collection
local source roots
expected records
datacomp_coyo_penguin
DataComp + COYO Penguin recap
datamultimodal/IMAGE/datacomp_1b, datamultimodal/IMAGE/coyo_700m
57,618,155
sa1b_penguin
SA-1B Penguin recap
datamultimodal/IMAGE/SA-1B
9,254,501
openimages_penguin
OpenImages Penguin recap… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Penguin-Recap-I.shotplan
ShotPlan Training Dataset
Multi-shot video training data for ShotPlan: Cinematic Video Generation with Learnable Planning Token.
💻 Code: https://github.com/Pensioner-11/ShotPlan
🤖 Models: ShotPlan-Wan2.1-T2V-14B · ShotPlan-Wan2.2-T2V-A14B-HighNoise
Contents
Path
Description
data/train_meta_16fps.json
6,404 training samples (metadata + captions)
data/videos/V*_16fps.mp4
549 source videos, re-encoded to 16 fps
Each sample is an 80-frame (5 s @… See the full description on the dataset page: https://huggingface.co/datasets/Pensioner/shotplan.multimath-300kmusique-anslm-eval-results-penfever-Llama-3-8B-tulu-human-v2-private
Dataset Card for Evaluation run of penfever/Llama-3-8B-tulu-human-v2
Dataset automatically created during the evaluation run of model penfever/Llama-3-8B-tulu-human-v2
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-penfever-Llama-3-8B-tulu-human-v2-private.pentest-agent-dataset-chatml
Pentest Agent Dataset - ChatML
This dataset is part of the Pentest Agent project and contains cybersecurity data formatted for fine-tuning language models.
Data Sources
Based on real CVEs from MITRE/NVD
Enriched with CVSS impact metrics
Linked to exploit code when available
Includes real-world pentesting scenarios
Contains command logic and execution steps
Includes red team techniques with MITRE references
Generated in ChatML format
Structure
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/7h3-R3v3n4n7/pentest-agent-dataset-chatml.Penguin-Recap-V
Penguin-Recap-V
Penguin-Recap-V provides Multi-granularity video annotation. This figure illustrates the alignment between visual content and textual descriptions across three temporal scales: Dense time-level, Paragraph-level, and Video-level.
Included subsets
subset
source collection
videos / clips
expected rows
source jsonl
sharegpt4video
ShareGPT4Video
40,145
120,435
sharegpt4video/predictions_process_relative.jsonl
shortvideo
ShortVideo
147,326
441,978… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Penguin-Recap-V.bug-bounty-pentest-en
Bug Bounty & Pentesting Methodologies
Methodologies (OWASP, PTES), checklists by app type, attack techniques, platforms, report templates and tools.
Links
French version
AYI NEDJIMI Consultants
pentest-agent-dataset-alpaca
Pentest Agent Dataset - Alpaca
This dataset is part of the Pentest Agent project and contains cybersecurity data formatted for fine-tuning language models.
Data Sources
Based on real CVEs from MITRE/NVD
Enriched with CVSS impact metrics
Linked to exploit code when available
Includes real-world pentesting scenarios
Contains command logic and execution steps
Includes red team techniques with MITRE references
Generated in Alpaca format
Structure
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/7h3-R3v3n4n7/pentest-agent-dataset-alpaca.PentestingCommandLogicEste dataset busca profundizar en la toma de decisiones respecto a comandos a ejecutar, es solo un test de octubre
lm-eval-results-penfever-Llama-3-8B-NuminaCoT-private
Dataset Card for Evaluation run of penfever/Llama-3-8B-NuminaCoT
Dataset automatically created during the evaluation run of model penfever/Llama-3-8B-NuminaCoT
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-penfever-Llama-3-8B-NuminaCoT-private.pentabrid-reproducibility
Pentabrid 27B: reproducibility package
Everything required to recompute the results of a controlled evaluation of fine-tuning
configurations for medical question answering. Openly available with no access
restrictions.
Contents
Path
Description
per_item/medxpertqa_*.jsonl
Per-item predictions for all six checkpoints on 2,450 MedXpertQA-Text items. Fields: id, gold, extracted_answer, correct, explicit_marker_present, n_markers, response_chars… See the full description on the dataset page: https://huggingface.co/datasets/Clinical-Reasoning-Hub/pentabrid-reproducibility.sec-edgar-filing-risks
Penumbra AI SEC 10-K Risk Factor Disclosures (S&P 500, FY2025-2026)
Dataset Summary
This dataset contains a sample of 2,683 individual risk factor disclosures from the most recent Form 10-K filing of 50 S&P 500 companies (fiscal year ends ranging from May 2025 to February 2026). Each record is a single risk driver — one discrete point a company disclosed under "Risk Factors" — together with the verbatim disclosure text and a set of machine generated categorical… See the full description on the dataset page: https://huggingface.co/datasets/PenumbraAI/sec-edgar-filing-risks.Medical-Entity-JSON-Extractionps4mas_penguin_scenery_200
ps4mas_penguin_scenery_200
200 stratified PENGUIN cases for the PS4MAS 0919 plan.
Schema (jsonl)
Each line:
{
"case_id": "c3085276",
"scenario_id": "c3085276",
"question": "How do I evaluate whether to hold or sell investments after a market crash?",
"profile": "{"Age": "55-64 years", "Gender": "Female", ...}",
"domain": "PENGUIN",
"scenario": "Cryptocurrency Crash"
}
Provenance
Pulled from… See the full description on the dataset page: https://huggingface.co/datasets/yinita/ps4mas_penguin_scenery_200.ps4mas_penguin_train_v2
ps4mas_penguin_train_v2
Disjoint training split for the PS4MAS V5 PPO experiments.
Source: wick1d/Personalized_Safety_Data (8098 rows)
Filter: 5 rows missing >=1 of 9 core attrs, 200 eval queries, dedup
Size: 7893 rows
Eval excluded: any query appearing in
yinita/ps4mas_penguin_scenery_200
Overlap with eval: 0 (verified)
Domain: PENGUIN
Schema (JSONL)
case_id, scenario_id, question, profile (JSON str), domain, scenario
pentest-findings
Pentest Security Findings Dataset
Cybersecurity and penetration testing training data for fine-tuning LLMs.
Sources
CVE Records Training Dataset
Cybersecurity Dataset Fenrir v2.0 (OWASP, MITRE, NIST)
Cybersecurity Instruction Tuning Dataset
Vulnerability CWE Patch Database
Exploit Database
MITRE ATT&CK Tactics and Techniques
Security Tools Pentesting
HackerOne Disclosed Reports
Format
Each example is a 3-message chat conversation:
system: Penetration testing… See the full description on the dataset page: https://huggingface.co/datasets/SkywardNomad92/pentest-findings.pentestr1-chat
Pentest-R1 Chat — TCO Fine-Tuning Dataset
Converted from KHenryAegis/Pentest-R1 into Gemma chat-template JSONL for fine-tuning autonomous penetration testing agents with Unsloth Studio.
Dataset summary
Field
Value
Conversations
535
Total messages
28,731
Format
JSONL — one {"messages": [...]} per line
Chat template
Gemma (unsloth/gemma-3n-E4B-it)
Token p50 / p95 / p99 / max
2,988 / 6,111 / 7,843 / 8,178
Recommended context_length
8192… See the full description on the dataset page: https://huggingface.co/datasets/supersamdev/pentestr1-chat.pentest-findings-v2
Pentest Security Findings Dataset v2 (Balanced)
Balanced cybersecurity and penetration testing training data for fine-tuning LLMs.
Key Improvements over v1
Balanced sources: CVE capped at 50K (was 240K+), underrepresented sources upsampled
Quality filter: Stricter minimum assistant response length for capped sources
Target: ~80-100K examples with better source diversity
Source Distribution
cve: 50,000 (46.8%)
fenrir: 30,000 (28.1%)
hackerone: 14,936 (14.0%)… See the full description on the dataset page: https://huggingface.co/datasets/SkywardNomad92/pentest-findings-v2.FIRE
Vision language models (VLMs) have achieved impressive progress in diverse applications, becoming a prevalent research direction. In this paper, we build FIRE, a feedback-refinement dataset, consisting of 1.1M multi-turn conversations that are derived from 27 source datasets, empowering VLMs to spontaneously refine their responses based on user feedback across diverse tasks. To scale up the data collection, FIRE is collected in two components: FIRE-100K and FIRE-1M, where FIRE-100K is… See the full description on the dataset page: https://huggingface.co/datasets/PengxiangLi/FIRE.lm-eval-results-penfever-Amber-7B-000-tulu-v2-private
Dataset Card for Evaluation run of penfever/Amber-7B-000-tulu-v2
Dataset automatically created during the evaluation run of model penfever/Amber-7B-000-tulu-v2
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-penfever-Amber-7B-000-tulu-v2-private.pentest-agent-dataset-alpaca
Pentest Agent Dataset - Alpaca
This dataset is part of the Pentest Agent project and contains cybersecurity data formatted for fine-tuning language models.
Data Sources
Based on real CVEs from MITRE/NVD
Enriched with CVSS impact metrics
Linked to exploit code when available
Includes real-world pentesting scenarios
Contains command logic and execution steps
Includes red team techniques with MITRE references
Generated in Alpaca format
Structure
Each… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/pentest-agent-dataset-alpaca.PenguinScrolls
PenguinScrolls: A User-Aligned Fine-Grained Benchmark for Long-Context Language Model Evaluation
Introduction
PenguinScrolls (企鹅卷轴) is a comprehensive benchmark designed to evaluate and enhance the long-text processing capabilities of large language models (LLMs).
Current benchmarks for evaluating long-context language models often rely on synthetic tasks that fail to adequately reflect real user needs, leading to a weak correlation between benchmark scores and actual… See the full description on the dataset page: https://huggingface.co/datasets/Penguin-Scrolls/PenguinScrolls.state-osha-plan-penalty-maximums
State OSHA plan civil penalty maximums by violation type compared with federal
Canonical, always-current version: https://referencesource.org/state-osha-plan-penalty-maximums/
Machine-readable: https://referencesource.org/state-osha-plan-penalty-maximums/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-15
Stale after: 2027-02-11 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 25
Maximum civil… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/state-osha-plan-penalty-maximums.PentestingCommandLogicbug-bounty-pentest-fr
Bug Bounty & Méthodologies de Pentest
Méthodologies (OWASP, PTES), checklists par type d app, techniques d attaque, plateformes, templates de rapports et outils.
Links
Version anglaise
AYI NEDJIMI Consultants
Pentesting_Datasetguodegang-penggenlangchain-MRKL-finetunelm-eval-results-penfever-Mistral-7B-tulu-v2-private
Dataset Card for Evaluation run of penfever/Mistral-7B-tulu-v2
Dataset automatically created during the evaluation run of model penfever/Mistral-7B-tulu-v2
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-penfever-Mistral-7B-tulu-v2-private.
