Can a social robot improve older adults’ wellbeing?
A six-week Japanese pilot found higher average wellbeing scores in all four small groups, including older adults speaking with a robot. With only 18 completers, no no-intervention control and no between-group tests, it cannot show that the robot or its encouragement caused the change.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1The project began with a 55-person survey, then randomised 20 eligible adults into four human-or-robot and high-or-low encouragement groups; two withdrew, leaving 18 completers and only four or five people per cell.
- 2Mean WHO-5 and life-satisfaction scores rose in every group over six weeks. The robot/high-encouragement group had the largest WHO-5 increase, but it started lower and finished almost level with the robot/low group.
- 3The researchers deliberately did not run between-group significance tests because the cells were so small. Without a no-conversation control, the study cannot isolate an effect of the robot, encouragement, attention, time or regression to the mean.
Research topic
Whether six weeks of encouraging conversations with a human or a social robot change subjective wellbeing among older Japanese adults
The direct answer: the pilot is encouraging, not conclusive
This study cannot establish that a social robot improved older adults' wellbeing. Average wellbeing scores rose during six weeks of conversation in all four study groups—whether participants spoke with a person or a RoBoHoN robot, and whether the partner used a strategy judged more or less encouraging. The robot group receiving the higher-efficacy strategy showed the largest average increase on one scale, but only five people completed that condition and they began with lower scores than the comparison robot group.
The result is best understood as feasibility evidence. Some older adults were willing to continue structured conversations with a robot and described useful experiences, which supports larger testing. It is not evidence that a device treats depression, prevents loneliness or can replace family, community workers or clinicians. For services considering social robots, the immediate question is how to test them safely as a supplement while preserving human contact and routes to professional help.[1]
The Impact Brief · Free
Follow the evidence in health & life sciences.
Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.
Two phases, with very different denominators
In the first phase, 55 Japanese adults aged 65 to 74 rated imagined encouraging responses in an online scenario. The researchers used those ratings to select 'offering specific actions' as a high-perceived-efficacy strategy and 'distraction' as a lower-efficacy strategy. This phase identified candidate conversation styles; participants did not interact physically with the robot, and their age range differed from the later experimental group.
For the six-week phase, 21 people were recruited, one did not pass screening and 20 were randomised equally to four cells: human partner with higher-efficacy encouragement, human with lower-efficacy encouragement, robot with higher-efficacy encouragement and robot with lower-efficacy encouragement. Two participants withdrew. The final analysis included 18 people—eight men and 10 women—aged 62 to 90, with a mean age of 74.7. Final cell sizes were four, four, five and five.
Participants held weekly conversations. Early sessions focused on relationship-building and reminiscence; weeks three to six invited discussion of worries and the assigned form of encouragement. The robot was Sharp's RoBoHoN. The study combined scripted structure with participant diaries, interviews and open-ended responses. That mixed approach is useful for learning how an intervention feels, but it does not compensate statistically for four or five observations in each comparison cell.[1]
What changed on the two wellbeing measures
Researchers measured affective wellbeing with the five-item World Health Organization index, WHO-5, and cognitive evaluation of life with the Satisfaction With Life Scale, SWLS. Scores were recorded before the intervention and at weeks two, four and six. These are recognised self-report measures, but they are outcomes reported by participants who knew whether they were talking to a person or a robot; the study was not blinded.
On WHO-5, the average pre-to-week-six change was plus 2.3 points for human/high encouragement, plus 1.8 for human/low, plus 4.2 for robot/high and plus 1.4 for robot/low. Yet the robot/high group began at 15.4, compared with 18.4 in robot/low, and finished at 19.6 versus 19.8. The apparently larger rise may therefore reflect its lower starting point rather than a superior intervention.
On SWLS, the corresponding mean changes were plus 3.3, plus 1.3, plus 2.2 and plus 2.0 points. Every group moved upward on average. The researchers appropriately did not report between-group inferential tests, because the sample was too small for credible comparisons. These descriptive trajectories can generate hypotheses, but they cannot tell us whether any difference is larger than sampling noise or meaningful to an individual older person.[1]
Why the missing control changes the interpretation
There was no group receiving no weekly conversation, usual activity alone or neutral listening without encouragement. As a result, improvement could reflect being heard, expecting benefit, repeated measurement, unrelated life events or a return toward typical scores after a low baseline. The study also cannot cleanly separate the identity of the conversation partner from the assigned encouragement strategy with such small cells.
Human delivery was uneven. The two human partners were both women, and one conducted seven participants' sessions while the other conducted one. Their warmth, pacing and skill may have mattered. On the technology side, results from one compact robot and one dialogue design do not establish a general effect for social robots or conversational AI. A different voice, embodiment, error rate or privacy practice could change both engagement and harm.
The participants were generally living in the community and many were in comparatively good health or already involved in social activities. People with cognitive impairment, severe depression, hearing difficulties, limited Japanese fluency or acute social-care needs may respond differently and may require safeguards not tested here. No clinical diagnosis or treatment outcome was evaluated.[1]
What the qualitative material can—and cannot—add
Diaries and interviews can reveal why a conversation felt supportive or awkward. They may help developers understand whether specific suggestions invite action, whether distraction feels dismissive, and when repetition or limited responses expose the robot's constraints. In an exploratory study, that information is valuable for refining scripts and deciding what a larger trial should measure.
Qualitative enthusiasm is not a clinical effect size. Participants may value novelty, attention from researchers or the ritual of a weekly appointment. People who withdrew may differ from those who completed the study. Because the paper explores sustained interaction over six weeks rather than a one-off encounter, it moves beyond a laboratory demonstration, but six weeks is still too short to establish durable wellbeing gains, reduced service use or protection against isolation.[1]
Implications for families, clinicians and public services
A cautious deployment would position a robot as an optional conversational aid, not a substitute for a visitor, therapist, nurse or social worker. Older adults should be able to decline it without losing services. They need a plain explanation of what is recorded, where data go, who can review conversations and what happens when the system misunderstands. A device that elicits worries also needs an escalation path for distress, abuse, self-harm, medication problems or urgent health symptoms.
Clinicians and commissioners should demand outcomes beyond engagement minutes. Useful trials would measure loneliness, mood, quality of life, adverse experiences, dropout, human-contact displacement and caregiver burden, with prespecified thresholds for meaningful change. Costs, maintenance, connectivity and staff time matter too. A robot that appears inexpensive but shifts monitoring work onto family members or fails in homes with poor connectivity may widen rather than reduce inequality.
For families, the study offers no basis to reduce contact because a device is present. The more constructive use is to test whether a tool helps a person prepare topics, remember an activity or bridge gaps between human visits. The person—not the device or provider—should retain control over when interaction starts and stops.[1]
Funding, disclosure and what would change the assessment
The study was supported by JST SPRING grant JPMJSP2128 and a Waseda University laboratory work fee. The authors declared no commercial or financial conflicts. Waseda University's ethics committee approved the work under reference 2024-603. The paper reports that GPT-5.2 was used for English grammar editing. These disclosures are useful, but ethical approval and peer review do not turn exploratory descriptive changes into causal evidence.
Confidence would change with a preregistered, adequately powered trial that includes a no-intervention or attention-matched control, balances human partners, masks outcome assessment where possible and reports both absolute scores and clinically meaningful change. It should recruit people with varied health, housing, income and digital experience, follow them for longer and record adverse effects and displacement of human contact. Until then, this pilot supports further research and careful co-design—not a claim that social robots improve older adults' wellbeing.[1]
What this means for people
- Older adults should retain a real choice and human routes to care; this study does not support replacing visits or clinical support with a robot.
- Families should treat a social robot as a possible bridge between human contacts, not evidence that less contact is safe.
- Clinicians and public services need escalation, privacy and adverse-event safeguards before inviting people to disclose worries to a device.
Global context
Japan's ageing population and robotics ecosystem make it an important setting for this work, but one small Waseda-led study cannot establish effects elsewhere. Expectations of robots, household structure, languages, care systems, connectivity and privacy law vary widely. Larger international trials should test whether benefits survive those differences and whether technology supplements rather than displaces human support.
What the evidence does not yet show
- Only 18 participants completed the six-week experiment, leaving four or five people in each cell and no credible basis for between-group significance testing.
- There was no no-intervention, usual-care or attention-matched control, so changes cannot be attributed to the robot or encouragement strategy.
- The robot/high group started with lower mean WHO-5 scores and ended almost level with robot/low, making its larger change vulnerable to baseline and regression-to-the-mean effects.
- Two female human partners delivered sessions unevenly, and one robot model limits generalisation across people and technologies.
- Self-reported wellbeing, a short follow-up and a relatively healthy community sample do not establish clinical benefit, reduced loneliness or durable effects.
What to watch next
- A preregistered trial powered for between-group comparisons with no-intervention and attention-matched controls.
- Longer follow-up measuring clinically meaningful change, loneliness, adverse events and whether robot use displaces human contact.
- Independent replication with varied robot designs, human facilitators, living arrangements and health needs.
- Clear privacy, consent and escalation protocols for conversations that surface health or safeguarding concerns.
- Cost and equity evidence, including connectivity, maintenance, caregiver workload and access for people with disabilities.
Living evidence record
Impact record IAI-1JSS933
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
9 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what Frontiers in Robotics and AI published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 9 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Do mental-health chatbots help university students?
A preregistered review found small short-term symptom advantages over waitlists or low-intensity materials, based on three depression and four anxiety comparisons. Two active conversational-control trials found no reliable added benefit, and certainty was low to very low.
8 min · 1 source
Health & Life Sciences
Can AI flag a longer ICU stay?
A peer-reviewed model separated longer-stay risk across 5,281 retrospective sepsis admissions in one US development cohort and two Korean external cohorts. Its AUROC fell from 0.848 internally to 0.781 at the smaller Korean site, and no prospective study tested whether alerts improve care or capacity.
7 min · 2 sources
Health & Life Sciences
Did the AI write a better clinic note?
A prospective single-centre study of 62 adults with abdominal pain found that AI-generated notes were more complete than clinicians' notes but scored worse on factuality. The assistant also produced only 50% diagnostic accuracy in the prospective cohort, so the result supports supervised documentation research—not autonomous clinical use.
9 min · 2 sources
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.