datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
the-secrets-of-ceos-book-2k
The-Secrets-Of-Ceos-Book-2k
Made with ❤️ using 🦥 Unsloth Studio
Beta2x was generated with Unsloth Recipe Studio. It contains 2,000 generated records.
🚀 Quick Start
from datasets import load_dataset
# Load the main dataset
dataset = load_dataset("filipwx/the-secrets-of-ceos-book-2k", "data", split="train")
df = dataset.to_pandas()
📊 Dataset Summary
📈 Records: 2,000
📋 Columns: 4
📋 Schema & Statistics
Column
Type
Column Type
Unique… See the full description on the dataset page: https://huggingface.co/datasets/filipwx/the-secrets-of-ceos-book-2k.prowl-secrets-corpus
Prowl secrets corpus
A labeled corpus for training and evaluating secret detectors (credentials, API keys, tokens,
database URIs, private keys, and passwords) across code, Jira tickets, Confluence pages, chat,
and logs, in multiple languages. It is the training data behind
Prowl and its stage-3
encoder.
503,027 records: 181,530 positive, 321,497 negative.
No live credentials. Every value is synthetic, format-preserving-obfuscated, or drawn from a public
test fixture. The… See the full description on the dataset page: https://huggingface.co/datasets/Podric/prowl-secrets-corpus.Scott-Schmitz-Real-Estate-CRM-Secretspstu-synthetic-secrets
PSTU Synthetic Secrets Dataset
Synthetic secrets benchmark for evaluating LLM memorization and unlearning, from the paper:
Not All Secrets Are Equal: Type-Aware Unlearning for Language Model Secret Removal
Hoda Fakhar — ECML PKDD 2026
Dataset Description
175 synthetic secrets across 25 types, each paired with 100 structurally similar decoys for computing the Carlini exposure metric.
All data is synthetically generated. No real credentials, PII, or sensitive information… See the full description on the dataset page: https://huggingface.co/datasets/Hodfa71/pstu-synthetic-secrets.curatorkit-testrun-Secrets
curatorkit-testrun-Secrets
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
curation
Backend
—
Model
—
Formats
alpaca
Artifact
dataset
Published
2026-08-30 05:55 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Secrets", "alpaca")
Dark-Psychology-Secretskl3m-data-dotgov-www.secretservice.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.secretservice.gov.TOP_SECRETShayleys-secretskl3m-filter-data-dotgov-www.secretservice.gov
