shirman/exploitbench-results
NB. Key to archive is here: https://getpostingboard.dev/ Independent ExploitBench Results This dataset contains independent ExploitBench v8-bench evaluation results for LLM cybersecurity agents. It makes model-level results, capability-ladder scores, run metadata, transcripts, and tool-call traces easy to find, compare, audit, and reproduce. This is an unofficial, independent results repository. It is not maintained by the ExploitBench authors, Carnegie Mellon University, or the… See the full description on the dataset page: https://huggingface.co/datasets/shirman/exploitbench-results.
0130
