Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Cyber LLM Benchmark Hub Logo
Cyber LLM Benchmark Hub
  • Home
  • Benchmarks
  • Contact
  • About
Support
Version v0.1.0Catalog Status: Active

Cybersecurity Benchmark Data You Can Trace

Cyber LLM Benchmark Hub is an open, source-linked catalog for tracking and comparing language model and agent evaluations across specialized cybersecurity tasks.

Current Version

v0.1.0

Catalog Release

August 2026

Verification

Schema Validated

Core Architectural Principles

How we maintain benchmark accuracy, transparency, and data integrity.

Source-backed Traceability

Every evaluation result connects directly to its peer-reviewed paper, primary code repository, or official leaderboard source.

Native Metrics Preservation

Scores retain their original published units (accuracy, pass@k, score ranges) without forcing lossy conversions.

Continuous Data Integrity

Automated catalog validation continuously verifies link health, metadata structure, and evaluation source fidelity.

Domain-Focused Taxonomy

Benchmarks are classified into specialized categories including vulnerability analysis, CTF challenges, code auditing, and threat intelligence.

Have benchmark data or updates to share?

Help expand the catalog by submitting missing publications, peer-reviewed evaluations, or dataset updates.

Suggest a Correction

Version & Changelog

Historical log of catalog enhancements, schema revisions, and dataset additions.

Latest: v0.1.0
v0.1.0Released August 2026
Current Release
  • FeatureLaunched initial source-linked catalog of cybersecurity benchmark evaluations for LLMs and autonomous agents.
  • FeatureAdded benchmark details view, model leaderboards, and side-by-side comparison matrices with source attribution.
  • InfrastructureImplemented automated catalog validation script (scripts/validate-data.js) and link check test suites.
  • DataCurated initial benchmark data across vulnerability exploitation, code auditing, and security Q&A tasks.
Cyber LLM Benchmark Hub

Cyber LLM Benchmark Hub

Benchmarking frontier models across cybersecurity tasks.

BenchmarksContactAbout

© 2026 Cyber LLM Benchmark Hub