Back to the news portal
Security & DefenceResearch paperResearchSource analysisVietnamCzech RepublicInternational

Can page content expose malicious websites?

A peer-reviewed benchmark of 689,556 webpages finds that analysing visible page content alongside the URL sharply reduced errors compared with URL-only models. The random stratified split, previously detected threats and missing per-page language labels leave live-world drift and novel attacks unresolved.

By The Impact of AI Security & Defence DeskReleased 4 October 2026 at 11:57 BST7 min read1 source

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesPhishingMalicious websitesMultilingual AICybersecurityBERTWeb safety

Research topic

Whether visible webpage content provides more reliable malicious-site detection than a URL alone across a large multilingual benchmark

The Impact of AI research cover asking whether page content can expose malicious websites, with conceptual text layers passing through a multilingual security shield.
AI-generated editorial illustration. The page, multilingual shield and suspicious content markers are conceptual; they do not depict a real website, measured output, named provider or successful breach.

At a glance

  • 1The final dataset contained 689,556 webpages from four sources, split 80:10:10 by class label after filtering and undersampling; language was estimated rather than verified per page.
  • 2XLM-RoBERTa with URL and content reached 99.01% accuracy, while a hybrid ensemble reached 99.06% with a 0.70% false-positive rate and 1.20% false-negative rate on the held-out benchmark.
  • 3The data contain already reported threats and a random split rather than a forward-in-time deployment test, so the headline accuracy does not establish performance against new attacks, domain drift or adversarial adaptation.

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Living evidence record

Impact record IAI-1OH4PR3

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Present

Record status

Monitoring

Last checked

4 October 2026

Source trail

1 direct source across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Single-source reporting disclosure

This record analyses one direct source. It can establish what Discover Computing published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.

A harmless-looking URL can hide a malicious page

Many malicious-site detectors start with the address: domain length, punctuation, suspicious words, subdomains or known reputation. Attackers can choose a clean-looking path, shorten a link, compromise a trusted host or copy a familiar brand. The paper tests a simple proposition: a system should inspect what the page is trying to persuade the user to do, not only the string in the browser bar.

The researchers convert raw HTML into a compact Markdown-like representation that keeps visible text and lightweight structure while removing scripts, styles and media data. Average document length fell by more than 96%. The URL and processed content were then supplied to multilingual transformer models, traditional text classifiers and hybrid ensembles. This is a classification benchmark, not a browser product or evidence that every malicious action is visible on-page.[1]

The dataset contains 689,556 filtered webpages

The final corpus combines 559,077 pages from PhreshPhish, 61,771 from a Mendeley phishing benchmark, 49,262 from MTLP and 19,446 from ChongLuaDao, a Vietnamese community reporting initiative. Each record included a URL, HTML and a safe-or-malicious label. The pipeline removed invalid and duplicate URLs, social-media and file-sharing cases, empty content and generic error pages, then undersampled to reduce class imbalance.

The authors used an 80:10:10 stratified split by class label. A balanced 20,000-record subset of the training partition supported exploratory tuning, while final scores came from the held-out test partition. The source datasets lacked verified language labels, so the researchers estimated language from processed text: roughly three quarters appeared to be English, with smaller Japanese, Vietnamese, French, Spanish, German and Chinese shares. Aggregate multilingual performance should not be read as equal performance in each language.[1]

Content produced the larger gain than model scale

The best content-aware transformer, XLM-RoBERTa Base, reached 99.01% accuracy, which the paper describes as a 62.2% error reduction over its URL-only counterpart. A hybrid ensemble reached 99.06% accuracy, a 0.70% false-positive rate and a 1.20% false-negative rate. Those rates are small, but at internet scale even fractions of a percent can block legitimate pages or miss many harmful ones.

A notable baseline was TF-IDF plus logistic regression using URL and content. It reached 97.82% accuracy and outperformed every tested URL-only transformer, including much larger models. The result suggests that obtaining the right evidence can matter more than adding parameters. It does not mean simple models are universally sufficient: the benchmark rewards patterns present in the collected corpus and may not capture rapidly changing attacker behaviour.[1]

Random splits can overstate readiness for a changing threat

A stratified random split preserves class balance but does not reproduce deployment, where tomorrow's phishing campaigns, brands and page templates differ from yesterday's training data. Pages from related campaigns or the same source can share language and structure across partitions even after exact URL deduplication. The paper does not report a domain-grouped or forward-in-time split, so some of the 99% headline may depend on stable source patterns rather than general resistance to drift.

Selection bias also runs in the opposite direction: malicious examples were already detected and reported to systems such as PhishTank, OpenPhish and community lists. Sophisticated attacks that escaped those pipelines are less likely to appear. Accuracy on known reported threats cannot measure recall against the unseen threat population. A prospective test on newly arriving pages, frozen before labels are known, would provide a stronger estimate of operational performance.[1]

The errors show why outside evidence remains necessary

The paper's false negatives included high-quality brand impersonation, complete fraudulent storefronts and sparse redirect pages. A copied page can look entirely legitimate if the model does not know which domain officially belongs to the brand. The authors argue that domain-registration history, certificates, reputation feeds and brand-to-domain mappings may add more value than simply scaling the language encoder.

Content-only reasoning also misses harm completed through email, phone, social media or an external destination. Community labels can encode regional legal judgements, while legitimate services with machine-generated subdomains may resemble attacker infrastructure. For users, a detector should therefore be one layer in a broader defence: safe browsing, identity and reputation checks, sandboxed content retrieval, user warnings, reporting and human investigation for disputed decisions.[1]

What would change the assessment

The decisive next test is temporal and adversarial. Researchers should freeze a model, evaluate later-arriving sites from sources not used in training and report performance by language, attack family, domain age and brand. Domain-grouped splits and near-duplicate detection would reduce campaign leakage. Red-team studies could test harmless-looking URLs, copied content, image-only phishing, deliberate multilingual obfuscation and pages designed to fool the Markdown conversion.

Operational trials should measure latency, bandwidth and the safety of fetching a potentially hostile page, not only classifier accuracy. They should also report appeals and harm from false positives, because blocking a legitimate service can affect access, commerce and speech. The study reports no external funding and no competing interests, and its datasets and methods are described in detail. It provides strong evidence that visible content adds useful signal within the benchmark; it does not yet prove durable protection on the live web.[1]

What this means for people

  • Users could receive better warnings when a benign-looking address hosts credential theft or a fraudulent storefront.
  • False positives can block legitimate services, so high aggregate accuracy does not remove the need for review and appeals.
  • Security teams may gain more from retrieving and analysing page content than from scaling an URL-only model, but doing so safely adds operational cost and risk.

Global context

The research team is based in Vietnam and the Czech Republic, and the corpus combines international benchmarks with a Vietnamese community dataset. That widens coverage beyond English-only URLs, but roughly three quarters of pages were estimated to be English and language labels were not verified. Attackers, brands and legal categories vary across jurisdictions, making country- and language-specific validation essential before deployment.

What the evidence does not yet show

  • The evaluation uses a random stratified split rather than a forward-in-time or domain-grouped deployment test.
  • Malicious examples were previously detected and reported, which may underrepresent novel attacks that evade existing systems.
  • Language composition was estimated and results were reported in aggregate, not with verified per-language performance.
  • The model cannot reliably detect harms completed through external channels or pages with little visible content.
  • High-quality brand impersonation and legitimate machine-generated subdomains remain important sources of false decisions.

What to watch next

  • Prospective evaluation on later-arriving webpages and sources absent from training.
  • Domain-grouped and campaign-aware splits with near-duplicate checks.
  • Per-language performance and testing against image-only, obfuscated and adversarial pages.
  • Production evidence covering safe page retrieval, latency, false-positive appeals and harm avoided.

Evidence trail

Sources used for this report

Links checked 4 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Security & Defence

Is AI changing cyberattacks—or speeding up familiar tactics?

Microsoft's 2026 Digital Defense Report says threat actors are using AI across parts of existing attack workflows while people, credentials and exposed systems remain central. Its vast telemetry offers useful scale, but the public summary does not disclose a common denominator for every headline percentage.

9 min · 2 sources

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.