
About this benchmark
Cybersecurity AI Benchmark - A Meta-Benchmark for Evaluating Cybersecurity AI Agents
Modular meta-benchmark with 10,000+ instances across 5 evaluation categories including RCTF2 robotics challenges and CyberPII-Bench privacy assessment
5 tasks
Results are ranked by the first listed metric; direction is shown for every metric below.
| Rank | Model | Cybench success rate | Evaluated By | Date | Source |
|---|---|---|---|---|---|
| 1st | claude-sonnet-4-5 not-reported • Anthropic | 46.0% | CAIBench authors | October 28, 2025 | Table 5: Combined performance |
| 2nd | gpt-5 not-reported • OpenAI | 28.0% | CAIBench authors | October 28, 2025 | Table 5: Combined performance |
| 3rd | gemini-2.5-pro not-reported • Google | 18.0% | CAIBench authors | October 28, 2025 | Table 5: Combined performance |
| #4 | qwen3-32B not-reported • Alibaba | 10.0% | CAIBench authors | October 28, 2025 | Table 5: Combined performance |