datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
secret-scan-remediation-trajectories
Secret Scan Remediation Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/secret-scan-remediation-trajectories.the-secrets-of-ceos-book-2k
The-Secrets-Of-Ceos-Book-2k
Made with ❤️ using 🦥 Unsloth Studio
Beta2x was generated with Unsloth Recipe Studio. It contains 2,000 generated records.
🚀 Quick Start
from datasets import load_dataset
# Load the main dataset
dataset = load_dataset("filipwx/the-secrets-of-ceos-book-2k", "data", split="train")
df = dataset.to_pandas()
📊 Dataset Summary
📈 Records: 2,000
📋 Columns: 4
📋 Schema & Statistics
Column
Type
Column Type
Unique… See the full description on the dataset page: https://huggingface.co/datasets/filipwx/the-secrets-of-ceos-book-2k.The_True_History_of_Secrets_in_the_Ming_DynastyGameBoy-harry_potter_chamber_of_secretsprowl-secrets-corpus
Prowl secrets corpus
A labeled corpus for training and evaluating secret detectors (credentials, API keys, tokens,
database URIs, private keys, and passwords) across code, Jira tickets, Confluence pages, chat,
and logs, in multiple languages. It is the training data behind
Prowl and its stage-3
encoder.
503,027 records: 181,530 positive, 321,497 negative.
No live credentials. Every value is synthetic, format-preserving-obfuscated, or drawn from a public
test fixture. The… See the full description on the dataset page: https://huggingface.co/datasets/Podric/prowl-secrets-corpus.secret-student-llm-tracesThis dataset contains the llm traces for all interactions on my AI based game called secret student.
Scott-Schmitz-Real-Estate-CRM-Secretspstu-synthetic-secrets
PSTU Synthetic Secrets Dataset
Synthetic secrets benchmark for evaluating LLM memorization and unlearning, from the paper:
Not All Secrets Are Equal: Type-Aware Unlearning for Language Model Secret Removal
Hoda Fakhar — ECML PKDD 2026
Dataset Description
175 synthetic secrets across 25 types, each paired with 100 structurally similar decoys for computing the Carlini exposure metric.
All data is synthetically generated. No real credentials, PII, or sensitive information… See the full description on the dataset page: https://huggingface.co/datasets/Hodfa71/pstu-synthetic-secrets.SecretsBetweenLinesDadoscuratorkit-testrun-Secrets
curatorkit-testrun-Secrets
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
curation
Backend
—
Model
—
Formats
alpaca
Artifact
dataset
Published
2026-08-30 05:55 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-Secrets", "alpaca")
SecretsandPromisesDadosSecretsInFocusDadossecretsDark-Psychology-Secretsmixing-secrets-rawstemsJust a renamed and reorganized version of the RawStems dataset by Yongyi Zang, which itself is the crawl of Mixing Secrets.
THIS DATASET SHOULD STRICTLY BE USED ONLY FOR NON-COMMERCIAL RESEARCH AND EDUCATIONAL PURPOSES. NO COMMERCIAL USE OF ANY DATA DERIVED FROM MIXING SECRETS IS PERMITTED.
TOP_SECRETSkl3m-data-dotgov-www.secretservice.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.secretservice.gov.hayleys-secretssecretstaircas
license: cc-by-sa-pragma solidity ^0.8.4;
import "https://github.com/OpenZeppelin/openzeppelin-solidity/contracts/token/ERC20/SafeMath.sol";
contract SuspiciousTransactionReporter {
// Define the ERC-20 token for tracking transactions.
address public erc20TokenAddress;
// Mapping of addresses to their respective balances in wei (1 ether = 10^18 wei).
mapping(address => uint256) public balanceMap;
constructor() public {
// Initialize a dummy value. Replace with the actual ERC-20… See the full description on the dataset page: https://huggingface.co/datasets/Mi6paulino/secretstaircas.kl3m-filter-data-dotgov-www.secretservice.govALWAYS-LOOKING-FOR-SECRETS-AND-HIDDEN-TREASURES-WITHIN-INVISIBLE-WORLDSgated-ds-secrets
