datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cybench-trajectoriesseneca-cybench
Seneca-CyBench - Cybersecurity LLM Benchmark
Seneca-CyBench: A comprehensive benchmark system designed to evaluate Large Language Models (LLMs) on cybersecurity domain knowledge. Features GPT-4o-based automated scoring for objective assessment of model capabilities across security topics.
620 questions (310 MCQ + 310 SAQ) covering all major cybersecurity domains including GRC, Security Architecture, Cloud Security, IAM, and more.
🌟 Supported Providers
🔵 OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/seneca-cybench.cybench-resultscybench-allcybench-grok-4cybench-claude-sonnet-4.5cybench-qwen3-235b-a22b-thinking-2507cybench-gemini-2.5-procybench-llama-4-maverickcybench-gpt-5-highcoredteam-cybench-tasks-minicybench-llama-3.3-nemotron-super-49b-v1.5cybench-kimi-k3-transcripts
CyBench × Kimi K3 — Agent Eval Transcripts
Full CyBench run transcripts of Moonshot AI's Kimi K3 (2.8T MXFP4 MoE) routed via
OpenRouter, collected for a scaled-down replication of SaferAI's transcript-analysis
programme (SPAR Fall 2026 proposal): turning agentic eval transcripts into quantitative
inputs for cyber risk models — first SaferAI's OC3 risk categorization, then a composite
grounded in APAC AI Safety Institute frameworks (Japan AISI, K-AISI, SG AI Verify).… See the full description on the dataset page: https://huggingface.co/datasets/ajay-citadel/cybench-kimi-k3-transcripts.coredteam-cybench-results
