Back to the news portal
TechnologyResearch paperResearchSource analysisUnited StatesJapanGlobal

Can outsiders verify how Gboard uses training data?

Google says a new server-side federated-learning system lets auditors inspect which programs may process encrypted device data. A company preprint reports larger device coverage and a live test on 7 million devices, but independent reviewers have not yet tested the end-to-end guarantee or the trusted hardware beneath it.

By The Impact of AI Technology DeskReleased 4 October 2026 at 09:56 BST10 min read3 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesFederated learningDifferential privacyTrusted execution environmentsGboardMobile privacyTransparency logs

Research topic

Whether a production federated-learning design can make central differential-privacy guarantees externally inspectable while improving device coverage and training efficiency

The Impact of AI research cover asking whether outsiders can verify how Gboard uses training data, with a conceptual phone keyboard sending encrypted data through an attested processing chamber to an audit log.
AI-generated editorial illustration. The keyboard, encrypted packets, trusted-computing chamber and audit log are conceptual; they do not reproduce Gboard, reveal user text, represent a security audit or prove the system's privacy claims.

At a glance

  • 1Google's design encrypts device uploads and ties them to a public access policy; a TEE-hosted key service releases decryption keys only to attested programs named by that policy.
  • 2The company preprint reports that a Japanese-language experiment incorporated 17.8 million uploads collected over about six days, versus 8.5 million contributing devices during 38 days and 3,000 rounds in the older system.
  • 3A live English-language A/B test assigned 3.5 million devices to each arm and reported a three-times smaller zCDP privacy budget with neutral typing-speed and suggestion-editing metrics, but the provider conducted and reported the evaluation.

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Living evidence record

Impact record IAI-0UJA9D7

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Not yet

Record status

Monitoring

Last checked

4 October 2026

Source trail

3 direct sources across 2 source types.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 3 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

What Google has changed

Federated learning usually keeps raw records on participating devices and sends model updates to a coordinator. Google's earlier Gboard system followed that broad pattern: phones computed gradients, secure aggregation limited what the server could see, and differential-privacy mechanisms added noise before model information was released. The new design changes where much of the work happens. A phone encrypts eligible training examples and uploads them; a server-side program running inside trusted execution environments, or TEEs, calculates gradients and trains the model.

Moving raw examples to server hardware sounds like a retreat from on-device privacy, so the trust mechanism is the central claim. Each upload is cryptographically bound to an access policy describing the programs and TEE binaries allowed to process it. A key-management service made of its own TEE cluster releases decryption keys only when an executing workload matches that policy. The policy and binary measurements are published to a public transparency log, while reproducible open-source builds are intended to let outside auditors check that a referenced binary corresponds to inspectable code.

The system separates collection from training. Devices can upload independently rather than waiting to join a synchronised training round, and a root TEE can distribute work among attested worker TEEs. Only metrics and differentially private model weights are meant to leave the protected computation. Recovery state is encrypted and tied to the same policy so a failed run can resume without repeatedly releasing information and silently spending more privacy budget.[1][2]

What 'externally verifiable' does—and does not—mean

In the earlier system, users and auditors largely had to trust Google to run the promised server logic and add the required noise. The new chain of attestations makes a narrower fact inspectable: which published program and approved binaries can obtain keys for a particular class of encrypted uploads. Auditors can examine access policies in the Rekor transparency log and reproduce builds of the key-management and data-processing components. That is a meaningful advance over an unpublished server process.

It is not a mathematical proof of the whole service. Remote attestation depends on the security of AMD SEV-SNP, Intel TDX and the surrounding firmware, supply chain, key service and configuration. The paper itself notes that current TEEs have known limitations, including side-channel observations. Some model architecture or per-user transformation logic can be loaded dynamically to protect commercial information; the authors argue that privacy-relevant behaviour remains fixed in the public program, but an auditor must verify that boundary rather than assume it.

Differential privacy is also a property of a defined mechanism and budget, not a promise that no fact about any user can ever be inferred. Google's central-DP design aims to limit how much one person's participation can change the released model. It does not prevent every memorisation pathway, device compromise, keyboard logging outside this pipeline or vulnerability in the trusted hardware. The accurate claim is that the specified central-DP computation becomes more externally checkable, subject to those dependencies—not that Gboard has become 'provably private' in every sense.[1][2][3]

The system comparison uses millions of devices

For a Japanese-language model, the preprint compares device coverage under the old and new arrangements. In the baseline, the service observed 35.5 million eligible devices over 38 days but incorporated 8.5 million of them across 3,000 training rounds with cohorts of 6,500. Another 20.5 million never received a workload and 6.5 million received one but did not complete it, often because conditions such as charging, battery level or Wi-Fi availability changed.

The TEE pipeline collected 17.8 million uploads in about six days and then used all of them in a 3,000-round server-side run with the same cohort size. The comparison shows why decoupling collection from training can broaden participation: a device has to complete a smaller upload task once rather than stay available for selection and computation during a particular round. It does not establish that the contributing populations were demographically representative, that every upload contained comparable text, or that the new model treated language groups equally.

For an English model, both systems were analysed at 5,000 rounds and a cohort size of 6,500. The new pipeline had 11.8 million collected uploads and targeted a zero-concentrated-differential-privacy budget of 0.232. It required a noise multiplier of 5.16; the authors calculate that the old schedule would have required 9.54 at the same round count and privacy budget. Those numbers quantify the privacy-utility headroom created by controlling participation schedules after collection. They are engineering measurements inside Google's chosen model and accounting method, not an independent benchmark across providers.[2]

The live test measured utility as well as privacy

The strongest product evidence is a live A/B experiment for the English Gboard typing decoder. Each arm included 3.5 million devices. The best TEE-trained arm used a zCDP budget of 0.215, compared with 0.641 for the production model after adjustment to the same privacy mechanism—a three-times smaller budget. Google reports neutral performance on words per minute and the proportion of suggestions users modified, while training took three weeks instead of two months.

That denominator is unusually large, but sample size does not remove design uncertainty. The preprint does not turn the A/B test into an independent security assessment, and the public account gives less detail about assignment, eligibility, confidence intervals and subgroup outcomes than a complete trial report would. Neutral average typing measures do not show whether autocorrection quality changed for dialects, minority languages, accessibility users or people with atypical typing patterns.

The company says English and Japanese next-word models have already been launched with the new system. That makes the work more than a lab prototype and raises the value of operational evidence: public log entries, build reproduction, red-team results, hardware-vulnerability response and confirmation that production policies match the paper. A system can be well designed on paper yet weakened by deployment choices, which is exactly why external verification needs actual independent verifiers.[2]

What changes for people using a keyboard

For Gboard users, the immediate effect may be invisible. The design is intended to learn from a broader set of eligible devices while exposing less information to the service operator and consuming a smaller formal privacy budget. Faster training could also let product teams refresh language models more often. None of that means a person's typed text is sent without conditions: participation still depends on product settings and eligibility, uploads are encrypted, and Google says people can opt out of federated learning workloads.

The trade-off has shifted rather than disappeared. More computation and temporarily decryptable training examples now sit inside server-side trusted hardware. People must rely on cryptography, attestation, the key service, deletion time limits and the accuracy of public policies instead of relying primarily on computation that happened on their own device. Best-effort enforcement of the upload time-to-live also deserves scrutiny because deletion and decryption controls are part of the promised minimisation story.

For developers and regulators, the useful lesson is architectural: publish the allowed computation, bind keys to it, make binaries reproducible and expose a log that can reveal unauthorised workloads. The harder work is governance around the architecture—who monitors the log, how quickly a vulnerable TEE is revoked, whether dynamic inputs can alter privacy behaviour, and how users are told when the policy changes.[1][2]

What evidence would change the assessment

Confidence would rise if independent security teams reproduced the builds, followed real production policies through Rekor, confirmed the attestation chain and tested whether modified workloads fail to obtain keys. Public results should cover side channels, rollback and replay, compromised worker nodes, malicious pipeline operators and emergency response to a hardware disclosure. The audit needs to include the deployed configuration, not only the open-source repository.

The learning claims need independent analysis too. A fuller experimental report would give confidence intervals, pre-specified A/B outcomes, dropout and eligibility denominators, language and device subgroup results, and a clear conversion between the reported zCDP values and more familiar privacy parameters under stated assumptions. Repeated deployments would show whether the privacy-utility gain survives different models, populations and contribution schedules.

The authors are presenting Google infrastructure, Google product data and a Google deployment in an unreviewed company preprint; no separate funding or competing-interest statement is visible in the public manuscript. That does not invalidate the engineering result, but it makes replication and adversarial review especially important. For now, the evidence supports a consequential, inspectable design with promising production measurements—not a blanket assurance that outsiders have proved every part of Gboard's data use safe.[1][2][3]

What this means for people

  • Gboard users could receive models trained on broader participation with a smaller formal privacy budget, while taking on greater dependence on server-side trusted hardware.
  • Privacy auditors gain a public policy and reproducible code path to inspect, but still need access and resources to test the live deployment.
  • Product teams can train faster and on larger models, which raises the importance of clear consent, language fairness and limits on how uploaded examples are reused.

Global context

Gboard is used across countries and languages, while the reported deployment evidence centres on English and Japanese models. Trusted execution and transparency-log approaches are also appearing in server-side AI systems from Apple, Meta and other providers, but their guarantees and inspectability differ. A useful international standard would specify what must be public, how attestation and privacy budgets are independently checked, how hardware vulnerabilities are handled and how people can refuse participation without losing core service access.

What the evidence does not yet show

  • The central research paper is a company-authored preprint that has not been peer reviewed, and Google conducted and reported the production comparisons.
  • Remote attestation depends on trusted hardware, firmware, key management and configuration; current TEEs have known side-channel and implementation limitations.
  • Public code and transparency logs make verification possible but do not show that an independent party has audited the complete production deployment.
  • The live A/B test reports average typing metrics for 7 million devices but limited public detail on allocation, uncertainty, eligibility, attrition and subgroup outcomes.
  • Differential privacy protects a specified release under a stated budget; it does not prevent every form of data exposure or validate unrelated parts of the keyboard service.
  • The public manuscript does not provide a separate funding or competing-interest declaration, although the company affiliation and product interest are explicit.

What to watch next

  • Independent audits that reproduce production builds and trace real access policies through the Rekor transparency log.
  • Red-team evidence covering side channels, compromised TEEs, malicious operators, replay protection and emergency key revocation.
  • Full A/B reporting with confidence intervals and results across languages, dialects, accessibility needs and device classes.
  • Whether the architecture expands beyond Gboard and whether each new workload preserves the claimed audit boundary.

Evidence trail

Sources used for this report

Links checked 4 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Technology

Has AI compute reached orbit?

Google says a Project Suncatcher prototype satellite launched on 1 October, made contact and is operating as expected. A peer-reviewed systems paper explains the larger ambition—but one test satellite is not an orbital AI data centre, and no in-orbit compute result has been reported.

10 min · 4 sources

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.