Can a small AI run a smart home without the cloud?
A peer-reviewed benchmark put one 3-billion-parameter model through 90 simulated smart-home commands and synthetic preference changes. AdaHome beat reimplemented baselines, but no people, homes or physical devices were tested—and responses still took roughly 10 to 14 seconds.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
Research topic
Whether a locally run small language model can interpret varied smart-home requests and adapt to simulated preferences without a cloud-scale reasoning system

At a glance
- 1AdaHome and four comparison systems used the same Llama 3.2 3B model, a standardised 12-device environment and 90 direct, indirect or ambiguous commands repeated three times.
- 2AdaHome scored 86.7% on direct and indirect commands and 88.9% on ambiguous commands, while its simulated preference mechanism was more stable and adaptive than a retrieval-augmented baseline.
- 3The evidence is synthetic: no resident, physical home, speech input, appliance failure or safety-critical deployment was tested. Average responses still took about 10.5 to 14.5 seconds.
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Living evidence record
Impact record IAI-0Y132L4
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent support
Present
Record status
Monitoring
Last checked
4 October 2026
Source trail
2 direct sources across 2 source types.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
The attraction is local control without a cloud-sized model
Smart-home assistants face a deceptively difficult language problem. “Turn on the kitchen light” is explicit. “It is hard to read in here” requires a system to infer a useful action from context. “Make the room comfortable” may depend on a resident’s earlier preferences, the time of day and which devices exist. Many recent designs send such requests to large cloud models or run elaborate reasoning chains, introducing network dependence, cost and privacy questions.
Researchers at the University of Southampton propose AdaHome as a smaller, local alternative. The system routes a command according to its estimated intent: direct requests receive a short path, requests needing interpretation use a compact chain-of-draft process, and ambiguous requests can trigger a clarification. A separate preference memory updates from feedback without retraining the language model or adding an ever-growing history to every prompt. The paper’s central question is therefore practical: how much smart-home reasoning can be retained when the same 3-billion-parameter model runs locally on modest hardware?[1][2]
Ninety constructed commands form the main benchmark
The authors built a standardised environment with 12 simulated devices and assembled 90 commands: 30 direct, 30 indirect and 30 ambiguous. Forty came from the Sasha benchmark and 50 from SAGE, with manual balancing across the three categories. Each system processed the set three times. AdaHome was compared with reimplementations of Sasha, SAGE and Harmony under the same Llama 3.2 3B model and greedy decoding, using Ollama on a machine with an AMD Ryzen 5 220 processor and 16GB of memory.
Direct requests were scored by exact match against an expected action. Indirect and ambiguous responses were judged by Qwen2.5-14B in three passes with majority voting. To test that automated judge, members of the research team labelled 48 examples and obtained a Cohen’s kappa of 0.834 against it. That is useful agreement evidence, but it is not a blinded assessment by residents or independent annotators, and an LLM judge can favour outputs that resemble its own preferred formulation.[2]
AdaHome led the simulated command tests, with noticeable waits
AdaHome returned the expected direct action in 86.7% of trials, compared with 63.3% for SAGE, 56.7% for Sasha and 41.1% for Harmony. It also scored 86.7% on indirect requests, versus roughly 63% to 68% for the comparators. On ambiguous requests, Harmony scored highest at 92.2%, AdaHome reached 88.9%, SAGE 84.4% and Sasha 30%. The authors’ separate intent classifier was correct on 75.6% of commands, so routing itself remained a material source of error.
Local does not mean instantaneous. Depending on command type, AdaHome’s average response time ranged from 10.48 to 14.47 seconds. It was faster than the reimplemented reasoning-heavy baselines—sometimes by close to threefold—but a ten-second pause may still feel slow for a light, lock or thermostat. The study measured software latency, not end-to-end time through speech recognition, a home network and an appliance, and it did not report energy use. Those omissions matter when local inference is promoted as efficient.[2]
The preference experiment is a scripted stress test, not a household trial
For personalisation, the researchers created 30 synthetic sequences of eight turns. Scenarios tested stable preferences, temporary deviation and a genuine shift. AdaHome retained the intended preference in 87.5% of stable cases, recovered from a one-off deviation 80% of the time and adapted in every simulated preference-shift sequence, taking an average 2.6 turns. A retrieval-augmented generation baseline achieved 52.5% stability, 10% recovery and 30% adaptation success, with successful changes taking two turns.
These numbers reveal behaviour under the authors’ scripted rules, not how a family negotiates conflicting preferences or changes its mind. The test had no children, guests, carers or shared rooms; no noisy feedback; and no mistaken or adversarial correction. Perfect adaptation success can therefore coexist with a weak understanding of real preference formation. A home assistant also needs to know when not to adapt—for example, when one unusual command should not overwrite a long-standing accessibility setting.[2]
Keeping data local reduces one exposure route, not every risk
Processing commands on a household machine can reduce the need to transmit intimate routines to a cloud provider. That is a meaningful design advantage, but the paper does not measure end-to-end privacy, security or data retention. A compromised hub, insecure device protocol or poorly protected preference store could still expose when people are home, what rooms they use and how they live. The benchmark also assumes the system’s device descriptions and current state are accurate.
Wrong actions can have consequences that a text benchmark obscures. Turning on a lamp unnecessarily is different from unlocking a door, switching on heat or disabling an alarm. AdaHome asks for clarification when its classifier identifies ambiguity, yet that classifier was wrong on nearly one command in four. A deployment would need action-specific permissions, confirmation thresholds, safe defaults, audit logs and reliable fallbacks when the local model or network fails.[2]
What would move the result from benchmark to home evidence
The next useful study would run prospectively in diverse occupied homes, compare AdaHome with a conventional non-generative controller and a cloud assistant, and preregister safety and usability outcomes. It should measure task completion, inappropriate actions, clarifications, response time, resident corrections, privacy expectations, energy use and failures over weeks rather than eight scripted turns. Household members should be able to inspect, contest and reset learned preferences.
Independent replication across other small models, processors, languages and device ecosystems would show whether the architecture—not one implementation—drives the result. The accessible paper did not include a specific funding acknowledgement or competing-interest declaration. Its evidence supports a promising local design and a controlled benchmark advantage. It does not yet establish that a small model can safely, privately or pleasantly run a real home.[1][2]
What this means for people
- Residents could keep more voice and routine data inside the home, but only if the full device stack and preference store are secured.
- People with disabilities may benefit from adaptable control, while slow or mistaken actions could create new barriers or safety risks.
- Households need ways to resolve conflicting preferences and stop one person’s feedback from silently changing shared settings.
Global context
The study comes from a UK university and uses English-language command benchmarks. Smart-home infrastructure, household composition, electricity constraints, privacy law and language vary widely. Local processing may be especially valuable where connectivity is poor or cloud services are unaffordable, but the tested desktop-class hardware and simulated devices do not establish performance on low-power hubs or multilingual homes.
What the evidence does not yet show
- All 90 commands, device states and preference sequences were simulated or constructed; no residents, occupied homes or physical devices were tested.
- The same small model and one hardware configuration were used, and comparator systems were reimplemented with some original functionality constrained.
- Indirect and ambiguous results depended on a larger language-model judge validated on 48 examples by the research team, not independent blinded raters.
- The study did not measure speech recognition, network and device latency, energy use, cybersecurity, privacy leakage or safety-critical consequences.
- The accessible paper did not state a specific funding source or competing-interest declaration.
What to watch next
- Prospective, preregistered trials in occupied homes with diverse residents and shared preferences.
- Safety tests that separate low-risk convenience actions from locks, heat, alarms and other consequential controls.
- Independent replication across devices, languages, small models and low-power local hardware.
- Transparent controls for viewing, correcting, pausing and deleting learned household preferences.
Evidence trail
Sources used for this report
Links checked 4 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Technology
Why is AI making enterprise modernisation more complex, not less?
A new survey of 2,000 senior business and technology leaders finds that AI is pushing organisations to expand several infrastructure types at once. The pattern matters, but the vendor-funded, self-reported study cannot show that AI or agentic tools caused better results.
9 min · 2 sources
Technology
Meta’s Muse gains downloads while Amazon challenges its shopping access
A third-party estimate puts the agent at 2.8 million downloads in its first 12 days. Amazon says the service was not authorised to access its store, raising questions about consent and credentials.
6 min · 2 sources
Technology
Qualcomm opens a Japan robotics initiative around on-device AI
Qualcomm announced a long-term robotics investment programme and a Japan Robotics Center intended to connect chip design with local manufacturers, researchers and automation specialists.
4 min · 1 source
Reader discussion
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.