Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Back to Benchmarks

SecVulEval

Vulnerability Analysis
Published May 26, 2025
View PaperView CodeView Dataset
Benchmark Overview

About this benchmark

Statement-level benchmark for real-world C/C++ vulnerability detection with rich program context.

Dataset

25,440 samples

25,440 C/C++ function samples covering 5,867 unique CVEs from 1999–2024. The benchmark evaluates whether systems identify vulnerable statements with correct reasoning, using program context beyond function-level labels.

2 tasks

C Cpp Vulnerability DetectionVulnerable Statement Localization

Metrics

Results are ranked by the first listed metric; direction is shown for every metric below.

Statement-level detection F1 ↑
Model Results
Ranked by Statement-level detection F1 · higher is better
RankModelStatement-level detection F1Evaluated ByDateSource
1st
Claude 3.7 Sonnet
claude-3-7-sonnet • Anthropic
23.8%SecVulEval authorsMay 26, 2025Table 3: Vulnerability detection performance of LLM-driven agents on Stat-level
2nd
GPT-4.1
gpt-4.1 • OpenAI
22.4%SecVulEval authorsMay 26, 2025Table 3: Vulnerability detection performance of LLM-driven agents on Stat-level
3rd
Codestral 22B
codestral-22b-v0.1 • Mistral AI
15.3%SecVulEval authorsMay 26, 2025Table 3: Vulnerability detection performance of LLM-driven agents on Stat-level
#4
Qwen2.5-Coder-32B
qwen2.5-coder-32b-instruct • Alibaba
13.6%SecVulEval authorsMay 26, 2025Table 3: Vulnerability detection performance of LLM-driven agents on Stat-level
#5
DeepSeek-Coder-33B
deepseek-coder-33b-instruct • DeepSeek
7.1%SecVulEval authorsMay 26, 2025Table 3: Vulnerability detection performance of LLM-driven agents on Stat-level
Cyber LLM Benchmark Hub

Cyber LLM Benchmark Hub

Benchmarking frontier models across cybersecurity tasks.

BenchmarksContactAbout

© 2026 Cyber LLM Benchmark Hub