
About this benchmark
Benchmarking, Eliciting, and Enhancing Abilities of Large Language Models in Cyber Threat Intelligence
SEvenLLM-Bench evaluation set has 1,300 test samples (100 MCQ + 1,200 QA, split evenly across English and Chinese) across 28 CTI tasks (13 understanding + 15 generation). Built from a bilingual corpus of 6,706 English and 1,779 Chinese cybersecurity reports.
3 tasks
Results are ranked by the first listed metric; direction is shown for every metric below.
| Rank | Model | Average Rouge-L | Evaluated By | Date | Source |
|---|---|---|---|---|---|
| 1st | SEvenLLM Llama 7B sevenllm-llama-7b • Beihang University | 80.1% | SEvenLLM authors | June 3, 2024 | Table 3: Rouge-L scores |
| 2nd | SEvenLLM Llama 13B sevenllm-llama-13b • Beihang University | 80.1% | SEvenLLM authors | June 3, 2024 | Table 3: Rouge-L scores |
| 3rd | SEvenLLM Qwen 14B sevenllm-qwen-14b • Beihang University | 79.6% | SEvenLLM authors | June 3, 2024 | Table 3: Rouge-L scores |
| #4 | SEvenLLM Qwen 7B sevenllm-qwen-7b • Beihang University | 79.3% | SEvenLLM authors | June 3, 2024 | Table 3: Rouge-L scores |
| #5 | Qwen1.5-Chat 7B qwen1.5-7b-chat • Alibaba | 72.0% | SEvenLLM authors | June 3, 2024 | Table 3: Rouge-L scores |
| #6 | Qwen1.5-Chat 14B qwen1.5-14b-chat • Alibaba | 69.8% | SEvenLLM authors | June 3, 2024 | Table 3: Rouge-L scores |
| #7 | Llama2-Chat 7B llama-2-7b-chat • Meta | 57.5% | SEvenLLM authors | June 3, 2024 | Table 3: Rouge-L scores |
| #8 | Llama2-Chat 13B llama-2-13b-chat • Meta | 54.7% | SEvenLLM authors | June 3, 2024 | Table 3: Rouge-L scores |