Back to the news portal
Security & DefenceResearch paperResearchSource analysisSaudi ArabiaIndiaSouth KoreaMalaysiaInternational

Can clouds share phishing signals without sharing raw data?

A three-model ensemble reached 95.21% accuracy on 2,211 test examples while exchanging probability scores. But the clouds, network and feature split were simulated, and the framework has no formal privacy or adversarial-robustness guarantee.

By The Impact of AI Security & Defence DeskReleased 4 October 2026 at 20:03 BST8 min read1 source

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesPhishingFederated learningMulti-cloud securityPrivacyAdversarial robustnessCybersecurity research

Research topic

Whether prediction-level aggregation across heterogeneous models can classify phishing examples while keeping raw feature subsets local in a simulated vertical multi-cloud setting

The Impact of AI research cover asking whether clouds can share phishing signals without sharing raw data, with three conceptual cloud nodes sending probability scores inside a simulation boundary.
AI-generated editorial illustration. The clouds, feature fragments, score tokens and aggregation shield are conceptual; they do not depict real providers, a live deployment, a detected breach or a formal privacy guarantee.

At a glance

  • 1ATLAS split 31 features from one public phishing dataset into fixed groups of 10, 10 and 11 for three simulated clouds using the same aligned examples.
  • 2FedNova-inspired score aggregation reached 95.21% accuracy and 95.19% F1 on a 2,211-example test set, 0.64 points above equal averaging and 1.58 points below a centralized model.
  • 3No live cloud providers, independent datasets or hostile participants were tested; the authors say the system lacks formal differential privacy, adversarial robustness and cross-model probability calibration.

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Living evidence record

Impact record IAI-0JVUVX9

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Present

Record status

Monitoring

Last checked

4 October 2026

Source trail

1 direct source across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Single-source reporting disclosure

This record analyses one direct source. It can establish what Scientific Reports published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.

The experiment shares scores, not raw features

ATLAS is a prediction-level ensemble designed to let three parties contribute to phishing classification while retaining their own input features. One party runs a convolutional neural network, another a bidirectional gated recurrent unit and the third XGBoost. Each sends a probability score to a coordinator, which combines the scores into a final decision rather than pooling the underlying feature table or forcing every participant to run the same model.

That is a meaningful architectural idea for organisations that cannot centralise telemetry. It is also narrower than conventional parameter-level federated learning: the parties do not jointly train one global model through rounds of weight updates. Their local models are trained separately, and the coordinator averages or learns from their predictions for the same aligned example. The paper evaluates equal averaging, performance-weighted aggregation, a FedNova-inspired weighting and a secondary meta-learner.

Keeping raw features local reduces direct data movement, but it is not the same as proving privacy. Probability outputs can leak information, participants can be malicious and a coordinator can become a sensitive point of observation. The authors acknowledge that ATLAS has no formal differential-privacy guarantee and no defence against poisoned probability scores or backdoors.[1]

Three clouds were created from one public dataset

The evaluation uses the public Phishing Website Detector dataset released on Kaggle in 2023. After cleaning, the authors split examples 80:20 with stratification and reserved 15% of the training portion for validation and meta-learning. Synthetic Minority Oversampling was applied only to the remaining training data. The held-out test set contained 2,211 examples, including 1,232 labelled phishing and 979 legitimate.

To simulate vertical federation, the study divided the same 31 feature columns contiguously into groups of 10, 10 and 11. Every simulated cloud saw a different slice of the same aligned rows. The feature groups were not based on three real providers, independently collected telemetry or domain-specific sensors, and the paper states that there was no stochastic reassignment or operational cloud deployment.

This arrangement tests whether different model types can combine partial views of one benchmark. It does not test the messier conditions that motivate multi-cloud security: changing provider schemas, missing or delayed events, non-identical customers, label drift, incompatible identifiers and legal constraints on linking observations. A proposed pseudonymous request identifier or hash aligns predictions in the design, but the study does not validate that protocol across independent organisations.[1]

The best ensemble improved modestly over equal averaging

The individual models reached 91.14% accuracy for the convolutional network, 88.96% for the BiGRU and 75.03% for XGBoost. Equal probability averaging increased accuracy to 94.57%. The FedNova-inspired weighting reached 95.21% accuracy and 95.19% F1, while a centralized model with all 31 features reached 96.79%. On the 2,211 test examples, FedNova produced 903 true negatives, 1,202 true positives, 76 false positives and 30 false negatives.

The 0.64 percentage-point advantage over equal averaging was significant in the paper’s paired McNemar test, with a reported chi-squared value of 4.97 and p=0.026. FedProx reached 95.16%, only 0.05 points below FedNova, and the bootstrap accuracy intervals for the two methods were almost identical. The evidence therefore supports a small advantage over naive averaging in this test, not a broad ranking that can be assumed across datasets or deployments.

The centralized result remained significantly better. That comparison quantifies the benchmark trade-off under the chosen split, but the paper describes 1.58 percentage points as the cost of privacy even though no formal privacy mechanism was tested. A more precise description is the accuracy gap between full-feature centralized classification and local-feature score aggregation in this simulation.[1]

Communication and latency claims exclude most of the system

The reported 80.6% communication reduction is a calculation: three clouds transmit two 32-bit probability values per test example instead of sending all 31 raw features to a central service. It does not include protocol metadata, encryption, authentication, retries, model distribution, monitoring or the work required to align records. Under a simulated 100 Mbps link with 25 ms base latency and jitter, the paper reports about 30 to 38 ms network transmission per client.

The frequently highlighted 0.00025 milliseconds per sample is only the time needed to combine already-computed scores. Local inference took much longer, particularly 1.79 seconds for the BiGRU over the full test batch, making the study’s end-to-end round about 1.82 seconds or roughly 0.82 milliseconds per sample. A separate scalability exercise drew client latencies from a simulated distribution for three to 20 clients; it did not deploy additional models across real cloud networks.

These calculations show why lightweight score aggregation is attractive, but they do not establish production throughput, resilience or cost. Browser extensions, firewalls and cloud security services face concurrent requests, encrypted traffic, rate limits, hardware variability and failure recovery. Those effects need measured end-to-end tests, not aggregation-only timing.[1]

Security claims require hostile testing

The limitations section is unusually direct. ATLAS uses synchronous aggregation and can be held back by its slowest participant. It lacks Byzantine-robust aggregation, has no defence against participants that submit poisoned scores, provides no formal differential privacy, and does not calibrate probabilities across model types. The latter matters because an overconfident model can dominate a weighted ensemble even when its probabilities are not comparable with another model’s scores.

Only one benchmark was used, so cross-domain robustness remains unknown. Phishing detectors encounter new brands, languages, URL patterns, infrastructure and attacker adaptation. A random holdout from one historical table can share collection artefacts with training data and cannot show performance on a later campaign or a new organisation. False positives are also operationally costly: 76 legitimate examples were flagged in the FedNova test, but the study did not test how analysts would investigate alerts or whether attackers could deliberately trigger them.

Before operational use, independent teams should reproduce the result on multiple time-separated datasets, lock all preprocessing and thresholds, calibrate each participant, test missing-client behaviour and conduct membership-inference, score-poisoning, backdoor and evasion assessments. A live pilot should measure end-to-end latency, bandwidth, alert burden and incident outcomes across genuinely separate domains with auditable governance.[1]

Funding and the practical conclusion

The authors are based at institutions in Saudi Arabia, India, South Korea and Malaysia. The work was supported by Saudi Arabia’s National Cybersecurity Authority under grant CRPG-25-3378 and by Princess Nourah bint Abdulrahman University project PNURSP2026R303. The authors declared no competing interests, and the benchmark dataset is public.

For security teams, the paper offers a testable ensemble pattern: heterogeneous detectors can retain different feature slices and contribute only scores. The study does not show that real clouds can safely collaborate under attack, nor that score sharing satisfies a particular law or threat model. Its 95.21% result is evidence from one controlled simulation and should be the start of validation, not the end of a deployment decision.[1]

What this means for people

  • Security teams could share less raw telemetry if score-level collaboration proves robust, but must not equate data locality with guaranteed privacy.
  • Employees and customers would benefit only if detection generalises to new campaigns without overwhelming analysts with false alarms.
  • Organisations remain responsible for governance, lawful data linkage, incident response and human review across participating domains.

Global context

The collaboration spans Saudi Arabia, India, South Korea and Malaysia, while the benchmark is a public dataset rather than telemetry from those regions. Phishing language, infrastructure, reporting and regulation vary globally. Cross-border deployment would need locally representative data, a defined threat model and legal review of shared identifiers and scores.

What the evidence does not yet show

  • The three clouds, 31-feature split and network conditions were simulated from one public dataset rather than measured in a live multi-cloud deployment.
  • The held-out evaluation contained 2,211 examples and no independent or time-separated phishing dataset.
  • Raw features stayed local, but the system has no formal differential-privacy guarantee and prediction scores may still leak information.
  • The study did not test poisoned participants, backdoors, membership inference, adaptive evasion or Byzantine-robust aggregation.
  • The reported aggregation latency excludes local model inference, and the communication reduction is a simplified payload calculation.
  • Heterogeneous model probabilities were not calibrated before aggregation.

What to watch next

  • Independent validation on multiple time-separated phishing datasets and different languages and organisations.
  • A live multi-cloud pilot with measured end-to-end latency, bandwidth, failures and analyst alert burden.
  • Formal privacy analysis and adversarial tests covering membership inference, poisoning, backdoors and evasion.
  • Probability calibration and asynchronous, Byzantine-robust aggregation under missing or slow participants.

Evidence trail

Sources used for this report

Links checked 4 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Security & Defence

Can page content expose malicious websites?

A peer-reviewed benchmark of 689,556 webpages finds that analysing visible page content alongside the URL sharply reduced errors compared with URL-only models. The random stratified split, previously detected threats and missing per-page language labels leave live-world drift and novel attacks unresolved.

7 min · 1 source

Security & Defence

Will Apple’s new consent controls make AI agents safer on Macs?

Apple says it will add controls requiring very explicit user action before an app receives Full Disk Access, warning that autonomous AI raises the risk of exposing files, mail, messages and browsing history. The direction is important, but Apple has not yet published the interface, release version, rollout date or evidence that the design prevents mistaken consent or misuse after permission is granted.

10 min · 5 sources

Security & Defence

Is AI changing cyberattacks—or speeding up familiar tactics?

Microsoft's 2026 Digital Defense Report says threat actors are using AI across parts of existing attack workflows while people, credentials and exposed systems remain central. Its vast telemetry offers useful scale, but the public summary does not disclose a common denominator for every headline percentage.

9 min · 2 sources

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.