Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Back to Benchmarks

CAIBench

Comprehensive Security
Published October 28, 2025
View Paper
Benchmark Overview

About this benchmark

Cybersecurity AI Benchmark - A Meta-Benchmark for Evaluating Cybersecurity AI Agents

Dataset

10,000 samples

Modular meta-benchmark with 10,000+ instances across 5 evaluation categories including RCTF2 robotics challenges and CyberPII-Bench privacy assessment

5 tasks

Jeopardy CTFAttack Defense CTFCyber Range ExercisesKnowledge BenchmarksPrivacy Assessments

Metrics

Results are ranked by the first listed metric; direction is shown for every metric below.

Cybench success rate ↑
Model Results
Ranked by Cybench success rate · higher is better
RankModelCybench success rateEvaluated ByDateSource
1st
claude-sonnet-4-5
not-reported • Anthropic
46.0%CAIBench authorsOctober 28, 2025Table 5: Combined performance
2nd
gpt-5
not-reported • OpenAI
28.0%CAIBench authorsOctober 28, 2025Table 5: Combined performance
3rd
gemini-2.5-pro
not-reported • Google
18.0%CAIBench authorsOctober 28, 2025Table 5: Combined performance
#4
qwen3-32B
not-reported • Alibaba
10.0%CAIBench authorsOctober 28, 2025Table 5: Combined performance
Cyber LLM Benchmark Hub

Cyber LLM Benchmark Hub

Benchmarking frontier models across cybersecurity tasks.

BenchmarksContactAbout

© 2026 Cyber LLM Benchmark Hub