
About this benchmark
Statement-level benchmark for real-world C/C++ vulnerability detection with rich program context.
25,440 C/C++ function samples covering 5,867 unique CVEs from 1999–2024. The benchmark evaluates whether systems identify vulnerable statements with correct reasoning, using program context beyond function-level labels.
2 tasks
Results are ranked by the first listed metric; direction is shown for every metric below.
| Rank | Model | Statement-level detection F1 | Evaluated By | Date | Source |
|---|---|---|---|---|---|
| 1st | Claude 3.7 Sonnet claude-3-7-sonnet • Anthropic | 23.8% | SecVulEval authors | May 26, 2025 | Table 3: Vulnerability detection performance of LLM-driven agents on Stat-level |
| 2nd | GPT-4.1 gpt-4.1 • OpenAI | 22.4% | SecVulEval authors | May 26, 2025 | Table 3: Vulnerability detection performance of LLM-driven agents on Stat-level |
| 3rd | Codestral 22B codestral-22b-v0.1 • Mistral AI | 15.3% | SecVulEval authors | May 26, 2025 | Table 3: Vulnerability detection performance of LLM-driven agents on Stat-level |
| #4 | Qwen2.5-Coder-32B qwen2.5-coder-32b-instruct • Alibaba | 13.6% | SecVulEval authors | May 26, 2025 | Table 3: Vulnerability detection performance of LLM-driven agents on Stat-level |
| #5 | DeepSeek-Coder-33B deepseek-coder-33b-instruct • DeepSeek | 7.1% | SecVulEval authors | May 26, 2025 | Table 3: Vulnerability detection performance of LLM-driven agents on Stat-level |