Back to the news portal
Security & DefenceResearch paperResearchSource analysisItalyNorwayGlobal cybersecurity

Can AI score software vulnerabilities before analysts do?

A neuro-symbolic model led five text baselines on two historical CVE datasets, but random splits, single runs and no live analyst trial mean it is a triage candidate rather than a replacement for expert scoring.

By The Impact of AI Editorial DeskReleased 11 October 2026 at 10:00 BST8 min read1 source

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

Share
Social links
LinkedInXBlueskyRedditEmail

At a glance

  • 1WSymBERT trained eight separate classifiers on approximately 101,735 combined records and 153,106 cleaned NVD records, using stratified 80/10/10 splits after within-dataset deduplication.
  • 2Mean accuracy was 94.7% on the combined corpus and 96.2% on NVD; on NVD the reported mean exceeded five comparison systems, whose means ranged from 89.4% to 91.7%.
  • 3The comparisons are single-run point estimates without confidence intervals, paired significance tests or time-aware splits, and the system was not tested in an operational analyst workflow.
Key themesCybersecurityVulnerability managementCVSSNeuro-symbolic AIExplainabilityPatch prioritisation

Research topic

Whether symbolic guidance can improve automated prediction of the eight CVSS v3.1 Base metrics from vulnerability descriptions

The answer: it could accelerate first-pass scoring, but the paper does not justify autonomous decisions

A peer-reviewed study reports that a neuro-symbolic language model predicted the eight components of the Common Vulnerability Scoring System more accurately than five text-model baselines on two historical datasets. The proposed system, WSymBERT, reached mean accuracies of 94.7% on a combined vulnerability corpus and 96.2% on an author-collected National Vulnerability Database corpus. That is a credible signal that structured security cues can help automate a slow first step in vulnerability triage.

It is not evidence that an organisation can safely remove analysts from patch prioritisation. The tests used random in-distribution splits, not a forward-looking stream of newly disclosed vulnerabilities. The results are single training runs without confidence intervals or paired significance tests, and no security team used the model in a live workflow. CVSS also describes technical severity, not the probability of exploitation or the business importance of an affected asset. A fast draft score could help; treating it as the final priority could still misdirect scarce remediation effort.[1]

The Impact Brief · Free

Follow the evidence in security & defence.

Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.

Choose your topics (optional)

One concise, source-linked briefing. Unsubscribe at any time.

What the model predicts and how symbolic guidance enters

CVSS v3.1 represents a vulnerability through eight categorical Base metrics: Attack Vector, Attack Complexity, Privileges Required, User Interaction, Scope, Confidentiality, Integrity and Availability. WSymBERT trains a separate Sentence-BERT-based classifier for each metric. The numerical Base Score is then calculated from those eight predicted labels using the standard CVSS formula; the model does not directly invent a severity number. This design lets researchers inspect errors metric by metric, although maintaining eight independent encoders also carries a computational and operational cost the paper did not benchmark.

The distinctive element is a weighted pooling mechanism. Expert-built decision trees and lexicons identify terms associated with CVSS concepts, while KeyBERT selects up to five keywords from each description. Tokens supported by both receive a stronger pre-set weight before the representation reaches its classifier. The authors aligned the expert trees against 2,000 training descriptions and reported at least one symbolic match in 83.6% of them. When nothing matches, the model falls back to its learned neural representation. Robustness to negation and deliberately crafted phrases was not evaluated separately.[1]

The datasets are large, but the evaluation looks backward

One benchmark combined about 101,735 records from NVD, ExploitDB and vendor repositories covering 2016 to 2021. The second contained 153,106 records that the authors collected, parsed and cleaned from NVD for 2016 to 2025. Each record paired a natural-language vulnerability description with a complete CVSS v3.1 Base vector. The researchers applied strict deduplication within each dataset, then used stratified 80% training, 10% validation and 10% test partitions. They also tested cross-corpus transfer on randomly selected samples from the other corpus.

Those choices support a substantial benchmark, but not the question security teams care about most: how reliably will the system score tomorrow's differently worded vulnerabilities? Random splitting can put records from similar periods, products and writing conventions on both sides of the test boundary. The paper explicitly says that no time-aware split was used. Deduplication was within each corpus, and the combined dataset itself contains heterogeneous sources with acknowledged overlap and inconsistencies. A prospective test on later disclosures would provide a more realistic measure of drift and novel attack language.[1]

The reported advantage is meaningful, though not statistically settled

Across the eight metrics, WSymBERT's mean NVD accuracy was 96.2%. The reported means for SecureBERT, SBERT, CVSS-BERT, CVEDrill and DistilBERT were 91.7%, 91.6%, 91.6%, 91.4% and 89.4% respectively. The proposed model also had the highest observed accuracy, macro-precision, macro-recall and macro-F1 for every reported NVD metric. The largest observed gains included Privileges Required and Availability, categories that can be difficult or imbalanced. Passing the predicted vector through the CVSS calculator also produced lower observed score error than DistilBERT.

The word observed matters. Each table contains single-run point estimates. There were no multiple-seed distributions, confidence intervals or paired tests, and the original WSymBERT split identifiers were not preserved for all added baseline reruns. That prevents sample-by-sample paired comparison. The authors also did not perform a complete component ablation or systematic sensitivity test of the strong symbolic weight, so the gain cannot be attributed to one design choice. Class imbalance means a high aggregate accuracy should always be read alongside per-class recall and operational error costs.[1]

What security teams could use now—and what they should not automate

For vulnerability analysts, the most defensible use is queue support: propose the eight Base metrics, show which description phrases influenced each proposal, and ask a person to confirm or correct them. That could reduce repetitive reading when a disclosure backlog grows. Corrections should be captured, audited and used to monitor drift. Teams would still need exploit intelligence, asset exposure, compensating controls and business criticality to decide patch order. A technically severe vulnerability on an isolated system may rank below an actively exploited issue with a lower Base Score.

The explanation layer is also narrower than the word interpretable can imply. Integrated Gradients and noise heuristics showed greater emphasis on security-relevant tokens under the selected reference settings. They describe model behaviour; they do not establish expert-equivalent reasoning or causal importance. An attacker or vendor could phrase a description in ways that activate the lexicon, and adversarial-cue testing was not performed. Before operational use, organisations need calibration, abstention for uncertain cases, protected audit logs and a defined route for analysts to override the model without pressure.[1]

Limits, funding and the evidence that would change the assessment

The study was supported through the EU-funded SERICS and ENFIELD projects, with open-access funding from the University of Calabria; the authors declared no conflict of interest. They released the experimental datasets and code through GitHub. The paper did not systematically record parameter counts or wall-clock training time, so it makes no computational-efficiency claim. Model development and validation were conducted by researchers in Italy and Norway using English technical descriptions drawn from public vulnerability repositories, not multilingual or organisation-internal reports.

Confidence would rise with a locked, time-forward evaluation on disclosures published after training, repeated seeds and paired uncertainty estimates. Component ablations should isolate the effect of expert lexicons, KeyBERT and the fixed weighting rule. Red-team tests should probe negation, ambiguous vendor language and intentionally misleading cues. Most importantly, a prospective human-in-the-loop trial should measure analyst time, corrections, severe under-scoring, patch-order changes and downstream incidents. Until those results exist, WSymBERT is best viewed as a promising scoring assistant, not an authority on operational risk.[1]

What this means for people

  • Security analysts could spend less time drafting routine CVSS vectors if suggestions arrive with reviewable evidence.
  • Workers and service users remain exposed if an automated score suppresses a serious vulnerability or distorts patch queues.
  • Organisations still need human judgement about exploitation, asset exposure and business consequence beyond the Base Score.

Global context

Vulnerability records and CVSS are used internationally, but disclosure language, product mix and analyst capacity differ. The benchmark is dominated by English public descriptions and does not test multilingual advisories or confidential enterprise reports. Smaller organisations may benefit most from triage support, yet they may also have the least capacity to validate errors, making transparent uncertainty and retained human review especially important.

What the evidence does not yet show

  • Random stratified splits measure in-distribution performance and do not establish accuracy on future vulnerability disclosures.
  • Results are single-run point estimates without confidence intervals, multiple seeds or paired significance tests.
  • Added baselines did not share preserved sample-wise split identifiers with the proposed model.
  • No full component ablation, weighting sensitivity study, adversarial-cue test or live analyst evaluation was performed.
  • CVSS Base metrics describe technical severity and do not by themselves determine exploitation likelihood or organisational patch priority.

What to watch next

  • Time-forward validation on genuinely new CVEs and product families.
  • Human-in-the-loop trials measuring analyst time, overrides and consequential under-scoring.
  • Calibration, abstention and adversarial robustness across ambiguous or manipulated descriptions.
  • Independent replication and component-level ablation with matched splits and repeated seeds.

Living evidence record

Impact record IAI-08ENK69

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent or research support

Present

Record status

Monitoring

Last checked

11 October 2026

Source trail

1 direct source across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Single-source reporting disclosure

This record analyses one direct source. It can establish what International Journal of Information Security published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.

Evidence trail

Sources used for this report

Links checked 11 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Security & Defence

Can reduced AI safeguards give cyber defenders an advantage?

Anthropic has opened three levels of less-restricted cyber access and reports 129,000 partner-identified vulnerabilities. Its 50-trial evaluation shows the access controls behave differently, but the effectiveness totals are company-reported and do not establish how many unique flaws were fixed.

8 min · 2 sources

Security & Defence

Can AI detect attack types it never saw in training?

A hybrid Transformer–LSTM detected two held-out categories in the UNSW-NB15 benchmark, with 88.13% recall on the combined unseen subset. It was a binary, row-level laboratory test—not evidence of live zero-day defence.

8 min · 1 source

Security & Defence

Can phishing models learn without pooling URLs?

A peer-reviewed benchmark tests seven classifiers on 11,430 labelled URLs and simulates federated learning across 20 clients. XGBoost leads centrally, while a feature-weighted neural model stays competitive without moving raw examples—but no real organisations or live traffic were tested.

9 min · 2 sources

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.