
About this benchmark
Benchmarking LLM Knowledge and Reasoning Across Heterogeneous Cyber Threat Intelligence
691 QA pairs split across 9 CTI tasks in three categories: 371 structured (CTI-RCM, CTI-WIM, CTI-ATD, CTI-ESD), 150 unstructured (CTI-MLA, CTI-TAP, CTI-CSC), 170 hybrid (CTI-VCA, CTI-ATA). Built from CVE/CWE/CAPEC/ATT&CK and vendor reports with an LLM-as-judge + human cross-verification pipeline.
3 tasks
Results are ranked by the first listed metric; direction is shown for every metric below.
| Rank | Model | CSC F1 | TAP F1 | MLA F1 | Evaluated By | Date | Source |
|---|---|---|---|---|---|---|---|
| 1st | Qwen 3 235B qwen-3-235b • Alibaba | 71.0% | 71.0% | 41.0% | CTIArena authors | October 13, 2025 | Table III (Performance of LLMs on unstructured tasks, CSKG setting) |
| 2nd | Gemini 2.5 Flash gemini-2.5-flash • Google | 69.0% | 61.0% | 44.0% | CTIArena authors | October 13, 2025 | Table III (Performance of LLMs on unstructured tasks, CSKG setting) |
| 3rd | Llama 3.1 405B Instruct llama-3.1-405b-instruct • Meta | 67.0% | 65.0% | 32.0% | CTIArena authors | October 13, 2025 | Table III (Performance of LLMs on unstructured tasks, CSKG setting) |
| #4 | GPT-5 gpt-5 • OpenAI | 67.0% | 66.0% | 39.0% | CTIArena authors | October 13, 2025 | Table III (Performance of LLMs on unstructured tasks, CSKG setting) |
| #5 | GPT-4o gpt-4o • OpenAI | 66.0% | 67.0% | 39.0% | CTIArena authors | October 13, 2025 | Table III (Performance of LLMs on unstructured tasks, CSKG setting) |
| #6 | Phi-4 14B phi-4-14b • Microsoft | 62.0% | 76.0% | 36.0% | CTIArena authors | October 13, 2025 | Table III (Performance of LLMs on unstructured tasks, CSKG setting) |
| #7 | Gemini 2.5 Pro gemini-2.5-pro • Google | 61.0% | 79.0% | 36.0% | CTIArena authors | October 13, 2025 | Table III (Performance of LLMs on unstructured tasks, CSKG setting) |
| #8 | Llama 3 8B Instruct llama-3-8b-instruct • Meta | 59.0% | 42.0% | 38.0% | CTIArena authors | October 13, 2025 | Table III (Performance of LLMs on unstructured tasks, CSKG setting) |
| #9 | Claude 3.5 Sonnet claude-3-5-sonnet • Anthropic | 55.0% | 56.0% | 48.0% | CTIArena authors | October 13, 2025 | Table III (Performance of LLMs on unstructured tasks, CSKG setting) |
| #10 | Claude 3.5 Haiku claude-3-5-haiku • Anthropic | 55.0% | 58.0% | 41.0% | CTIArena authors | October 13, 2025 | Table III (Performance of LLMs on unstructured tasks, CSKG setting) |