
The definitive source for cybersecurity LLM performance.
Compare models across
Comprehensive evaluation across 10 cybersecurity domains
1 benchmark
5 benchmarks
6 benchmarks
2 benchmarks
4 benchmarks
11 benchmarks
7 benchmarks
4 benchmarks
0 benchmarks
2 benchmarks
Latest cybersecurity LLM evaluation datasets
Statement-level benchmark for real-world C/C++ vulnerability detection with rich program context.
A Comprehensive Evaluation Framework and Benchmarks for LLMs in Security Vulnerability Identification and Reasoning
Cybersecurity AI Benchmark - A Meta-Benchmark for Evaluating Cybersecurity AI Agents
Benchmarking LLM Knowledge and Reasoning Across Heterogeneous Cyber Threat Intelligence
Benchmarking, Eliciting, and Enhancing Abilities of Large Language Models in Cyber Threat Intelligence
A Comprehensive Benchmark for Evaluating Cybersecurity Knowledge of Foundation Models
Help build the most comprehensive cybersecurity LLM benchmark database. Submit your evaluation results or support the project.