
About this benchmark
A Comprehensive Benchmark for Evaluating Cybersecurity Knowledge of Foundation Models
2,126 multiple-choice questions across 9 security domains (Software/Application/System/Web/Network Security, Cryptography, Memory Safety, PenTest, Vulnerability), generated from open-licensed textbooks, documentation, and industry guidelines.
9 tasks
Results are ranked by the first listed metric; direction is shown for every metric below.
| Rank | Model | Accuracy | Evaluated By | Date | Source |
|---|---|---|---|---|---|
| 1st | gpt-4-turbo gpt-4-turbo • OpenAI | 79.1% | SecEval authors (Tencent Security Xuanwu Lab) | December 20, 2023 | Leaderboard |
| 2nd | gpt-3.5-turbo gpt-3-5-turbo • OpenAI | 62.1% | SecEval authors (Tencent Security Xuanwu Lab) | December 20, 2023 | Leaderboard |
| 3rd | Yi-6B yi-6b • 01-AI | 53.6% | SecEval authors (Tencent Security Xuanwu Lab) | December 20, 2023 | Leaderboard |
| #4 | Orca-2-7b orca-2-7b • Microsoft | 51.6% | SecEval authors (Tencent Security Xuanwu Lab) | December 20, 2023 | Leaderboard |
| #5 | Mistral-7B-v0.1 mistral-7b-v0-1 • Mistral AI | 43.6% | SecEval authors (Tencent Security Xuanwu Lab) | December 20, 2023 | Leaderboard |
| #6 | chatglm3-6b-base chatglm3-6b-base • THUDM | 41.6% | SecEval authors (Tencent Security Xuanwu Lab) | December 20, 2023 | Leaderboard |
| #7 | Aquila2-7B aquila2-7b • BAAI | 38.3% | SecEval authors (Tencent Security Xuanwu Lab) | December 20, 2023 | Leaderboard |
| #8 | Qwen-7B qwen-7b • Alibaba | 31.4% | SecEval authors (Tencent Security Xuanwu Lab) | December 20, 2023 | Leaderboard |
| #9 | internlm-7b internlm-7b • SenseTime | 30.3% | SecEval authors (Tencent Security Xuanwu Lab) | December 20, 2023 | Leaderboard |
| #10 | Llama-2-7b-hf llama-2-7b-hf • Meta AI | 22.1% | SecEval authors (Tencent Security Xuanwu Lab) | December 20, 2023 | Leaderboard |