
Cyber LLM Benchmark Hub is an open, source-linked catalog for tracking and comparing language model and agent evaluations across specialized cybersecurity tasks.
Current Version
v0.1.0
Catalog Release
August 2026
Verification
Schema Validated
How we maintain benchmark accuracy, transparency, and data integrity.
Every evaluation result connects directly to its peer-reviewed paper, primary code repository, or official leaderboard source.
Scores retain their original published units (accuracy, pass@k, score ranges) without forcing lossy conversions.
Automated catalog validation continuously verifies link health, metadata structure, and evaluation source fidelity.
Benchmarks are classified into specialized categories including vulnerability analysis, CTF challenges, code auditing, and threat intelligence.
Help expand the catalog by submitting missing publications, peer-reviewed evaluations, or dataset updates.
Historical log of catalog enhancements, schema revisions, and dataset additions.