CIO Influence
CIO Influence News Machine Learning Security

Collinear AI Launches CWE-bench to Test Frontier Coding Agents on Defensive Cybersecurity Capabilities

Collinear AI Launches CWE-bench to Test Frontier Coding Agents on Defensive Cybersecurity Capabilities

AWS Marketplace: Collinear AI

Held-out cybersecurity benchmark spans 54 weakness types; the leading agent passes less than 50% of tasks, and 18 remain unsolved.

Collinear AI has launched CWE-bench, a held-out benchmark that tests whether frontier coding agents can defend real software vulnerabilities.

AI systems are now finding and exploiting security flaws across real infrastructure, largely on their own. If agents can discover and exploit cyber weaknesses, we need to know how well they can also find and repair them. CWE-bench accurately measures whether they can.

To ensure broad vulnerability coverage, CWE-bench is built around MITRE’s CWE taxonomy. The benchmark of 100 agentic tasks currently spans 54 weakness types and all 10 OWASP Top 10 2025 categories, with the aim of further expanding coverage across MITRE’s catalog.

Also Read: CIO Influence Interview with John Elliott, Cybersecurity Author Fellow at Pluralsight

The tasks are designed so that memorizing published fixes is not enough. In one example, three of four leading agents fixed the publicly documented token-revocation paths but missed a newly introduced path, earning zero credit. They recognized the known version of the vulnerability but failed to reason through how the weakness appeared elsewhere in the code.

“Cybersecurity is one of the toughest remaining hill climbs in coding. We built CWE-bench to make that climb faster with hard but fair environments that expose useful failures,” said Nazneen Rajani, CEO of Collinear AI, who previously led post-training at Hugging Face. “CWE-bench applies all known vulnerabilities to known open-source code bases. All frontier models have the knowledge of these vulnerabilities, and these codebases are already in their training data but the leading model still scores less than 50%. The benchmark gives model builders a trusted signal about what to improve next.”

Initial results include:

  • Fable 5 leads the current leaderboard with a 47% pass@1 score at maximum reasoning.
  • 18 of the 100 tasks remain unsolved by every model tested.
  • No agent tested successfully repairs a majority of the benchmark.
  • Performance varies across weakness types, exposing specific areas where defensive software reasoning still needs to improve.
  • Models on the Pareto front of performance vs. cost include Fable 5, Gemini 3.8 Flash Cyber, and GPT-5.6 Sol — all at high reasoning.

Catch more CIO Insights: How Are CIOs Aligning Technology with Workforce Agility?

[To share your insights with us, please write to psen@itechseries.com ]

Related posts

FocalPointK12 Earns SOC 2 Type II Certification, Demonstrating Commitment to Data Security

PR Newswire

SilverEdge DC Launches with M4 Corridor Data Centre

CIO Influence News Desk

ESET to Showcase Advanced Cybersecurity Solutions at Black Hat MEA 2025

EIN Presswire