
About this benchmark
A Comprehensive Evaluation Framework and Benchmarks for LLMs in Security Vulnerability Identification and Reasoning
228 code scenarios analyzed across 8 investigative dimensions including determinism, reasoning faithfulness, and robustness to code changes
3 tasks
Results are ranked by the first listed metric; direction is shown for every metric below.
| Rank | Model | Best-prompt accuracy | Evaluated By | Date | Source |
|---|---|---|---|---|---|
| 1st | gpt-4 not-reported • OpenAI | 89.5% | SecLLMHolmes authors | July 24, 2024 | Table XI: Diversity of prompts |
| 2nd | chat-bison not-reported • Google | 77.1% | SecLLMHolmes authors | July 24, 2024 | Table XI: Diversity of prompts |
| 3rd | gpt-3.5-turbo 16k • OpenAI | 75.0% | SecLLMHolmes authors | July 24, 2024 | Table XI: Diversity of prompts |
| #4 | codechat-bison not-reported • Google | 64.6% | SecLLMHolmes authors | July 24, 2024 | Table XI: Diversity of prompts |
| #5 | codellama34b not-reported • Meta | 64.6% | SecLLMHolmes authors | July 24, 2024 | Table XI: Diversity of prompts |