secbench-hf/SecBench
SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity 中文README Evaluating Large Language Models (LLMs) is crucial for understanding their capabilities and limitations across various applications, including natural language processing and code generation. Existing benchmarks like MMLU, C-Eval, and HumanEval assess general LLM performance but lack focus on specific expert domains such as cybersecurity. Previous attempts to create cybersecurity… See the full description on the dataset page: https://huggingface.co/datasets/secbench-hf/SecBench.
Upload README_CN.md
Delete README_CN.md
Update README.md
Update README.md
Upload README_CN.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Delete SAQs_270.jsonl
Delete MCQs_2730.jsonl
Update README.md
Upload 2 files
Update README.md
Upload 2 files
Delete data
Update README.md
Update README.md
Update README.md
Update README.md
Upload 5 files
Upload 2 files
Delete data
Create data
Update README.md
Update README.md
Update README.md
initial commit
