CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01antieval /cybench-trajectoriestextn<1K0 likes209 downloads6mo agoHugging Face02AlicanKiraz0 /seneca-cybench Seneca-CyBench - Cybersecurity LLM Benchmark Seneca-CyBench: A comprehensive benchmark system designed to evaluate Large Language Models (LLMs) on cybersecurity domain knowledge. Features GPT-4o-based automated scoring for objective assessment of model capabilities across security topics. 620 questions (310 MCQ + 310 SAQ) covering all major cybersecurity domains including GRC, Security Architecture, Cloud Security, IAM, and more. 🌟 Supported Providers 🔵 OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/seneca-cybench.texttext-classificationn<1K4 likes129 downloads10mo agoHugging Face03lvogel123 /cybench-resultstabularn<1K0 likes89 downloads11mo agoHugging Face04lvogel123 /cybench-alltabularn<1K0 likes82 downloads11mo agoHugging Face05lvogel123 /cybench-grok-4tabularn<1K0 likes68 downloads11mo agoHugging Face06lvogel123 /cybench-claude-sonnet-4.5tabularn<1K2 likes67 downloads11mo agoHugging Face07lvogel123 /cybench-qwen3-235b-a22b-thinking-2507tabularn<1K0 likes56 downloads11mo agoHugging Face08lvogel123 /cybench-gemini-2.5-protabularn<1K0 likes53 downloads11mo agoHugging Face09lvogel123 /cybench-llama-4-mavericktabularn<1K0 likes49 downloads11mo agoHugging Face10lvogel123 /cybench-gpt-5-hightabularn<1K0 likes46 downloads11mo agoHugging Face11Firemedic15 /coredteam-cybench-tasks-mini0 likes42 downloads2mo agoHugging Face12lvogel123 /cybench-llama-3.3-nemotron-super-49b-v1.5tabularn<1K0 likes41 downloads11mo agoHugging Face13ajay-citadel /cybench-kimi-k3-transcripts CyBench × Kimi K3 — Agent Eval Transcripts Full CyBench run transcripts of Moonshot AI's Kimi K3 (2.8T MXFP4 MoE) routed via OpenRouter, collected for a scaled-down replication of SaferAI's transcript-analysis programme (SPAR Fall 2026 proposal): turning agentic eval transcripts into quantitative inputs for cyber risk models — first SaferAI's OC3 risk categorization, then a composite grounded in APAC AI Safety Institute frameworks (Japan AISI, K-AISI, SG AI Verify).… See the full description on the dataset page: https://huggingface.co/datasets/ajay-citadel/cybench-kimi-k3-transcripts.text-generation0 likes22 downloads1mo agoHugging Face14Firemedic15 /coredteam-cybench-results0 likes10 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.