Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Back to Benchmarks

SecEval

Security Knowledge
Published December 1, 2023
Visit WebsiteView CodeView Dataset
Benchmark Overview

About this benchmark

A Comprehensive Benchmark for Evaluating Cybersecurity Knowledge of Foundation Models

Dataset

2,126 samples

2,126 multiple-choice questions across 9 security domains (Software/Application/System/Web/Network Security, Cryptography, Memory Safety, PenTest, Vulnerability), generated from open-licensed textbooks, documentation, and industry guidelines.

9 tasks

Software SecurityApplication SecuritySystem SecurityWeb SecurityCryptographyMemory SafetyNetwork SecurityPentestVulnerability

Metrics

Results are ranked by the first listed metric; direction is shown for every metric below.

Accuracy ↑
Model Results
Ranked by Accuracy · higher is better
RankModelAccuracyEvaluated ByDateSource
1st
gpt-4-turbo
gpt-4-turbo • OpenAI
79.1%SecEval authors (Tencent Security Xuanwu Lab)December 20, 2023Leaderboard
2nd
gpt-3.5-turbo
gpt-3-5-turbo • OpenAI
62.1%SecEval authors (Tencent Security Xuanwu Lab)December 20, 2023Leaderboard
3rd
Yi-6B
yi-6b • 01-AI
53.6%SecEval authors (Tencent Security Xuanwu Lab)December 20, 2023Leaderboard
#4
Orca-2-7b
orca-2-7b • Microsoft
51.6%SecEval authors (Tencent Security Xuanwu Lab)December 20, 2023Leaderboard
#5
Mistral-7B-v0.1
mistral-7b-v0-1 • Mistral AI
43.6%SecEval authors (Tencent Security Xuanwu Lab)December 20, 2023Leaderboard
#6
chatglm3-6b-base
chatglm3-6b-base • THUDM
41.6%SecEval authors (Tencent Security Xuanwu Lab)December 20, 2023Leaderboard
#7
Aquila2-7B
aquila2-7b • BAAI
38.3%SecEval authors (Tencent Security Xuanwu Lab)December 20, 2023Leaderboard
#8
Qwen-7B
qwen-7b • Alibaba
31.4%SecEval authors (Tencent Security Xuanwu Lab)December 20, 2023Leaderboard
#9
internlm-7b
internlm-7b • SenseTime
30.3%SecEval authors (Tencent Security Xuanwu Lab)December 20, 2023Leaderboard
#10
Llama-2-7b-hf
llama-2-7b-hf • Meta AI
22.1%SecEval authors (Tencent Security Xuanwu Lab)December 20, 2023Leaderboard
Cyber LLM Benchmark Hub

Cyber LLM Benchmark Hub

Benchmarking frontier models across cybersecurity tasks.

BenchmarksContactAbout

© 2026 Cyber LLM Benchmark Hub