Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Back to Benchmarks

CTIArena

Threat Intelligence
Published October 13, 2025
View Paper
Benchmark Overview

About this benchmark

Benchmarking LLM Knowledge and Reasoning Across Heterogeneous Cyber Threat Intelligence

Dataset

691 samples

691 QA pairs split across 9 CTI tasks in three categories: 371 structured (CTI-RCM, CTI-WIM, CTI-ATD, CTI-ESD), 150 unstructured (CTI-MLA, CTI-TAP, CTI-CSC), 170 hybrid (CTI-VCA, CTI-ATA). Built from CVE/CWE/CAPEC/ATT&CK and vendor reports with an LLM-as-judge + human cross-verification pipeline.

3 tasks

Structured CTI AnalysisUnstructured CTI AnalysisHybrid CTI Analysis

Metrics

Results are ranked by the first listed metric; direction is shown for every metric below.

CSC F1 ↑TAP F1 ↑MLA F1 ↑
Model Results
Ranked by CSC F1 · higher is better
RankModelCSC F1TAP F1MLA F1Evaluated ByDateSource
1st
Qwen 3 235B
qwen-3-235b • Alibaba
71.0%71.0%41.0%CTIArena authorsOctober 13, 2025Table III (Performance of LLMs on unstructured tasks, CSKG setting)
2nd
Gemini 2.5 Flash
gemini-2.5-flash • Google
69.0%61.0%44.0%CTIArena authorsOctober 13, 2025Table III (Performance of LLMs on unstructured tasks, CSKG setting)
3rd
Llama 3.1 405B Instruct
llama-3.1-405b-instruct • Meta
67.0%65.0%32.0%CTIArena authorsOctober 13, 2025Table III (Performance of LLMs on unstructured tasks, CSKG setting)
#4
GPT-5
gpt-5 • OpenAI
67.0%66.0%39.0%CTIArena authorsOctober 13, 2025Table III (Performance of LLMs on unstructured tasks, CSKG setting)
#5
GPT-4o
gpt-4o • OpenAI
66.0%67.0%39.0%CTIArena authorsOctober 13, 2025Table III (Performance of LLMs on unstructured tasks, CSKG setting)
#6
Phi-4 14B
phi-4-14b • Microsoft
62.0%76.0%36.0%CTIArena authorsOctober 13, 2025Table III (Performance of LLMs on unstructured tasks, CSKG setting)
#7
Gemini 2.5 Pro
gemini-2.5-pro • Google
61.0%79.0%36.0%CTIArena authorsOctober 13, 2025Table III (Performance of LLMs on unstructured tasks, CSKG setting)
#8
Llama 3 8B Instruct
llama-3-8b-instruct • Meta
59.0%42.0%38.0%CTIArena authorsOctober 13, 2025Table III (Performance of LLMs on unstructured tasks, CSKG setting)
#9
Claude 3.5 Sonnet
claude-3-5-sonnet • Anthropic
55.0%56.0%48.0%CTIArena authorsOctober 13, 2025Table III (Performance of LLMs on unstructured tasks, CSKG setting)
#10
Claude 3.5 Haiku
claude-3-5-haiku • Anthropic
55.0%58.0%41.0%CTIArena authorsOctober 13, 2025Table III (Performance of LLMs on unstructured tasks, CSKG setting)
Cyber LLM Benchmark Hub

Cyber LLM Benchmark Hub

Benchmarking frontier models across cybersecurity tasks.

BenchmarksContactAbout

© 2026 Cyber LLM Benchmark Hub