CoolFace
Datasetpublic

Podric/prowl-secrets-corpus

Prowl secrets corpus A labeled corpus for training and evaluating secret detectors (credentials, API keys, tokens, database URIs, private keys, and passwords) across code, Jira tickets, Confluence pages, chat, and logs, in multiple languages. It is the training data behind Prowl and its stage-3 encoder. 503,027 records: 181,530 positive, 321,497 negative. No live credentials. Every value is synthetic, format-preserving-obfuscated, or drawn from a public test fixture. The… See the full description on the dataset page: https://huggingface.co/datasets/Podric/prowl-secrets-corpus.

sourceHugging Facecc-by-nc-4.0updated 3mo agoView on Hugging Face
2likes543downloads
filecorpus.parquet69.7 MBdownload
fileprowlbench.parquet2.5 MBdownload

Podric/prowl-secrets-corpus · main · files are served by the source, never re-hosted here