CoolFace
Datasetpublic

shirman/exploitbench-results

NB. Key to archive is here: https://getpostingboard.dev/ Independent ExploitBench Results This dataset contains independent ExploitBench v8-bench evaluation results for LLM cybersecurity agents. It makes model-level results, capability-ladder scores, run metadata, transcripts, and tool-call traces easy to find, compare, audit, and reproduce. This is an unofficial, independent results repository. It is not maintained by the ExploitBench authors, Carnegie Mellon University, or the… See the full description on the dataset page: https://huggingface.co/datasets/shirman/exploitbench-results.

sourceHugging Faceupdated 14d agoView on Hugging Face
0likes130downloads
7 commits on main
ce3b7c714d ago

Update README.md

shirman
c440bc91mo ago

Replace ExploitBench results dataset

shirman
a9fc4e21mo ago

Replace ExploitBench results dataset

shirman
c3b71881mo ago

Add ExploitBench results dataset

shirman
cc460531mo ago

Present ExploitBench results as released dataset

shirman
7124eb01mo ago

Add searchable ExploitBench results dataset card

shirman
08537131mo ago

initial commit

shirman