Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Back to Benchmarks

SecLLMHolmes

Vulnerability Analysis
Published December 19, 2023
View PaperView Code
Benchmark Overview

About this benchmark

A Comprehensive Evaluation Framework and Benchmarks for LLMs in Security Vulnerability Identification and Reasoning

Dataset

228 samples

228 code scenarios analyzed across 8 investigative dimensions including determinism, reasoning faithfulness, and robustness to code changes

3 tasks

Vulnerability IdentificationSecurity ReasoningCode Analysis

Metrics

Results are ranked by the first listed metric; direction is shown for every metric below.

Best-prompt accuracy ↑
Model Results
Ranked by Best-prompt accuracy · higher is better
RankModelBest-prompt accuracyEvaluated ByDateSource
1st
gpt-4
not-reported • OpenAI
89.5%SecLLMHolmes authorsJuly 24, 2024Table XI: Diversity of prompts
2nd
chat-bison
not-reported • Google
77.1%SecLLMHolmes authorsJuly 24, 2024Table XI: Diversity of prompts
3rd
gpt-3.5-turbo
16k • OpenAI
75.0%SecLLMHolmes authorsJuly 24, 2024Table XI: Diversity of prompts
#4
codechat-bison
not-reported • Google
64.6%SecLLMHolmes authorsJuly 24, 2024Table XI: Diversity of prompts
#5
codellama34b
not-reported • Meta
64.6%SecLLMHolmes authorsJuly 24, 2024Table XI: Diversity of prompts
Cyber LLM Benchmark Hub

Cyber LLM Benchmark Hub

Benchmarking frontier models across cybersecurity tasks.

BenchmarksContactAbout

© 2026 Cyber LLM Benchmark Hub