Can outsiders verify how Gboard uses training data?
Google says a new server-side federated-learning system lets auditors inspect which programs may process encrypted device data. A company preprint reports larger device coverage and a live test on 7 million devices, but independent reviewers have not yet tested the end-to-end guarantee or the trusted hardware beneath it.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
Research topic
Whether a production federated-learning design can make central differential-privacy guarantees externally inspectable while improving device coverage and training efficiency

At a glance
- 1Google's design encrypts device uploads and ties them to a public access policy; a TEE-hosted key service releases decryption keys only to attested programs named by that policy.
- 2The company preprint reports that a Japanese-language experiment incorporated 17.8 million uploads collected over about six days, versus 8.5 million contributing devices during 38 days and 3,000 rounds in the older system.
- 3A live English-language A/B test assigned 3.5 million devices to each arm and reported a three-times smaller zCDP privacy budget with neutral typing-speed and suggestion-editing metrics, but the provider conducted and reported the evaluation.
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Living evidence record
Impact record IAI-0UJA9D7
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent support
Not yet
Record status
Monitoring
Last checked
4 October 2026
Source trail
3 direct sources across 2 source types.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 3 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
What Google has changed
Federated learning usually keeps raw records on participating devices and sends model updates to a coordinator. Google's earlier Gboard system followed that broad pattern: phones computed gradients, secure aggregation limited what the server could see, and differential-privacy mechanisms added noise before model information was released. The new design changes where much of the work happens. A phone encrypts eligible training examples and uploads them; a server-side program running inside trusted execution environments, or TEEs, calculates gradients and trains the model.
Moving raw examples to server hardware sounds like a retreat from on-device privacy, so the trust mechanism is the central claim. Each upload is cryptographically bound to an access policy describing the programs and TEE binaries allowed to process it. A key-management service made of its own TEE cluster releases decryption keys only when an executing workload matches that policy. The policy and binary measurements are published to a public transparency log, while reproducible open-source builds are intended to let outside auditors check that a referenced binary corresponds to inspectable code.
The system separates collection from training. Devices can upload independently rather than waiting to join a synchronised training round, and a root TEE can distribute work among attested worker TEEs. Only metrics and differentially private model weights are meant to leave the protected computation. Recovery state is encrypted and tied to the same policy so a failed run can resume without repeatedly releasing information and silently spending more privacy budget.[1][2]
What 'externally verifiable' does—and does not—mean
In the earlier system, users and auditors largely had to trust Google to run the promised server logic and add the required noise. The new chain of attestations makes a narrower fact inspectable: which published program and approved binaries can obtain keys for a particular class of encrypted uploads. Auditors can examine access policies in the Rekor transparency log and reproduce builds of the key-management and data-processing components. That is a meaningful advance over an unpublished server process.
It is not a mathematical proof of the whole service. Remote attestation depends on the security of AMD SEV-SNP, Intel TDX and the surrounding firmware, supply chain, key service and configuration. The paper itself notes that current TEEs have known limitations, including side-channel observations. Some model architecture or per-user transformation logic can be loaded dynamically to protect commercial information; the authors argue that privacy-relevant behaviour remains fixed in the public program, but an auditor must verify that boundary rather than assume it.
Differential privacy is also a property of a defined mechanism and budget, not a promise that no fact about any user can ever be inferred. Google's central-DP design aims to limit how much one person's participation can change the released model. It does not prevent every memorisation pathway, device compromise, keyboard logging outside this pipeline or vulnerability in the trusted hardware. The accurate claim is that the specified central-DP computation becomes more externally checkable, subject to those dependencies—not that Gboard has become 'provably private' in every sense.[1][2][3]
The system comparison uses millions of devices
For a Japanese-language model, the preprint compares device coverage under the old and new arrangements. In the baseline, the service observed 35.5 million eligible devices over 38 days but incorporated 8.5 million of them across 3,000 training rounds with cohorts of 6,500. Another 20.5 million never received a workload and 6.5 million received one but did not complete it, often because conditions such as charging, battery level or Wi-Fi availability changed.
The TEE pipeline collected 17.8 million uploads in about six days and then used all of them in a 3,000-round server-side run with the same cohort size. The comparison shows why decoupling collection from training can broaden participation: a device has to complete a smaller upload task once rather than stay available for selection and computation during a particular round. It does not establish that the contributing populations were demographically representative, that every upload contained comparable text, or that the new model treated language groups equally.
For an English model, both systems were analysed at 5,000 rounds and a cohort size of 6,500. The new pipeline had 11.8 million collected uploads and targeted a zero-concentrated-differential-privacy budget of 0.232. It required a noise multiplier of 5.16; the authors calculate that the old schedule would have required 9.54 at the same round count and privacy budget. Those numbers quantify the privacy-utility headroom created by controlling participation schedules after collection. They are engineering measurements inside Google's chosen model and accounting method, not an independent benchmark across providers.[2]
The live test measured utility as well as privacy
The strongest product evidence is a live A/B experiment for the English Gboard typing decoder. Each arm included 3.5 million devices. The best TEE-trained arm used a zCDP budget of 0.215, compared with 0.641 for the production model after adjustment to the same privacy mechanism—a three-times smaller budget. Google reports neutral performance on words per minute and the proportion of suggestions users modified, while training took three weeks instead of two months.
That denominator is unusually large, but sample size does not remove design uncertainty. The preprint does not turn the A/B test into an independent security assessment, and the public account gives less detail about assignment, eligibility, confidence intervals and subgroup outcomes than a complete trial report would. Neutral average typing measures do not show whether autocorrection quality changed for dialects, minority languages, accessibility users or people with atypical typing patterns.
The company says English and Japanese next-word models have already been launched with the new system. That makes the work more than a lab prototype and raises the value of operational evidence: public log entries, build reproduction, red-team results, hardware-vulnerability response and confirmation that production policies match the paper. A system can be well designed on paper yet weakened by deployment choices, which is exactly why external verification needs actual independent verifiers.[2]
What changes for people using a keyboard
For Gboard users, the immediate effect may be invisible. The design is intended to learn from a broader set of eligible devices while exposing less information to the service operator and consuming a smaller formal privacy budget. Faster training could also let product teams refresh language models more often. None of that means a person's typed text is sent without conditions: participation still depends on product settings and eligibility, uploads are encrypted, and Google says people can opt out of federated learning workloads.
The trade-off has shifted rather than disappeared. More computation and temporarily decryptable training examples now sit inside server-side trusted hardware. People must rely on cryptography, attestation, the key service, deletion time limits and the accuracy of public policies instead of relying primarily on computation that happened on their own device. Best-effort enforcement of the upload time-to-live also deserves scrutiny because deletion and decryption controls are part of the promised minimisation story.
For developers and regulators, the useful lesson is architectural: publish the allowed computation, bind keys to it, make binaries reproducible and expose a log that can reveal unauthorised workloads. The harder work is governance around the architecture—who monitors the log, how quickly a vulnerable TEE is revoked, whether dynamic inputs can alter privacy behaviour, and how users are told when the policy changes.[1][2]
What evidence would change the assessment
Confidence would rise if independent security teams reproduced the builds, followed real production policies through Rekor, confirmed the attestation chain and tested whether modified workloads fail to obtain keys. Public results should cover side channels, rollback and replay, compromised worker nodes, malicious pipeline operators and emergency response to a hardware disclosure. The audit needs to include the deployed configuration, not only the open-source repository.
The learning claims need independent analysis too. A fuller experimental report would give confidence intervals, pre-specified A/B outcomes, dropout and eligibility denominators, language and device subgroup results, and a clear conversion between the reported zCDP values and more familiar privacy parameters under stated assumptions. Repeated deployments would show whether the privacy-utility gain survives different models, populations and contribution schedules.
The authors are presenting Google infrastructure, Google product data and a Google deployment in an unreviewed company preprint; no separate funding or competing-interest statement is visible in the public manuscript. That does not invalidate the engineering result, but it makes replication and adversarial review especially important. For now, the evidence supports a consequential, inspectable design with promising production measurements—not a blanket assurance that outsiders have proved every part of Gboard's data use safe.[1][2][3]
What this means for people
- Gboard users could receive models trained on broader participation with a smaller formal privacy budget, while taking on greater dependence on server-side trusted hardware.
- Privacy auditors gain a public policy and reproducible code path to inspect, but still need access and resources to test the live deployment.
- Product teams can train faster and on larger models, which raises the importance of clear consent, language fairness and limits on how uploaded examples are reused.
Global context
Gboard is used across countries and languages, while the reported deployment evidence centres on English and Japanese models. Trusted execution and transparency-log approaches are also appearing in server-side AI systems from Apple, Meta and other providers, but their guarantees and inspectability differ. A useful international standard would specify what must be public, how attestation and privacy budgets are independently checked, how hardware vulnerabilities are handled and how people can refuse participation without losing core service access.
What the evidence does not yet show
- The central research paper is a company-authored preprint that has not been peer reviewed, and Google conducted and reported the production comparisons.
- Remote attestation depends on trusted hardware, firmware, key management and configuration; current TEEs have known side-channel and implementation limitations.
- Public code and transparency logs make verification possible but do not show that an independent party has audited the complete production deployment.
- The live A/B test reports average typing metrics for 7 million devices but limited public detail on allocation, uncertainty, eligibility, attrition and subgroup outcomes.
- Differential privacy protects a specified release under a stated budget; it does not prevent every form of data exposure or validate unrelated parts of the keyboard service.
- The public manuscript does not provide a separate funding or competing-interest declaration, although the company affiliation and product interest are explicit.
What to watch next
- Independent audits that reproduce production builds and trace real access policies through the Rekor transparency log.
- Red-team evidence covering side channels, compromised TEEs, malicious operators, replay protection and emergency key revocation.
- Full A/B reporting with confidence intervals and results across languages, dialects, accessibility needs and device classes.
- Whether the architecture expands beyond Gboard and whether each new workload preserves the claimed audit boundary.
Evidence trail
Sources used for this report
Links checked 4 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Technology
Has AI compute reached orbit?
Google says a Project Suncatcher prototype satellite launched on 1 October, made contact and is operating as expected. A peer-reviewed systems paper explains the larger ambition—but one test satellite is not an orbital AI data centre, and no in-orbit compute result has been reported.
10 min · 4 sources
Technology
Google adds a live visual avatar to its real-time Gemini model
Google introduced Gemini 3.8 Live with Live Avatar, combining spoken conversation with a real-time visual presence and pointing to a more embodied form of consumer and workplace AI interaction.
4 min · 1 source
Technology
Cloud offloading could change the performance and battery trade-off for robots
Microsoft Research reports that moving some physical-AI inference from a robot's onboard processor to edge or cloud GPUs can improve task success, efficiency and the complexity of workloads the machine can handle.
4 min · 1 source
Reader discussion
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.