Back to the news portal
Health & Life SciencesResearch paperResearchSource analysisChinaAsiaInternational

Does a digital human improve AI stress counselling?

Seventy-one Beijing university students rated an avatar-based CBT system easier to use than the same text chatbot after one session. The study measured experience—not symptom improvement, long-term safety or therapeutic effectiveness.

By The Impact of AI Health & Life Sciences DeskReleased 4 October 2026 at 16:57 BST7 min read1 source

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesDigital mental healthUniversity studentsCognitive behavioural therapyMulti-agent systemsUser experienceAI safety

Research topic

Whether a multimodal digital-human interface changes cognitive load, perceived social presence, usability and intended reuse compared with the same CBT multi-agent system delivered as text

The Impact of AI research cover asking whether a digital human improves AI stress counselling, with a conceptual student-facing avatar and a one-session evidence boundary.
AI-generated editorial illustration. The student, avatar, dialogue and safety boundary are conceptual; they do not depict a study participant, therapy session, clinical outcome, crisis intervention or approved mental-health service.

At a glance

  • 1Ninety-seven Beijing Normal University students were screened; 71 aged 18–25 with mild-to-moderate anxiety or depression and willingness to seek support were randomly assigned to an avatar condition (36) or the same backend as text only (35).
  • 2After one session, the digital-human group reported lower intrinsic and extraneous cognitive load, higher usability and stronger intention to reuse; the social-presence difference did not reach the conventional significance threshold.
  • 3The researchers did not report a post-session change in depression, anxiety or academic stress, no human-therapist control was used, and subtle crisis signals could escape the system’s explicit-risk screen.

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Living evidence record

Impact record IAI-1I0BZAD

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Present

Record status

Monitoring

Last checked

4 October 2026

Source trail

1 direct source across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Single-source reporting disclosure

This record analyses one direct source. It can establish what Scientific Reports published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.

The trial isolates presentation more than therapy

DHCBT combines a large-language-model workflow organised around cognitive behavioural therapy with an animated digital human that speaks, displays text and supplies visual cues. The authors argue that specialised agents can divide rapport, assessment, cognitive reframing and action planning into a coordinated process, while an anthropomorphic interface may make support feel more socially present than an ordinary chat window.

The user study compared that avatar, voice and text interface with a text-only chatbot running the same CBT multi-agent backend. This is a useful design choice because the generated counselling architecture was intended to be shared between groups. The experiment therefore mainly asks whether multimodal embodiment changes perceived effort and experience during one session. It does not compare the AI with a therapist, usual care, a non-CBT conversation or no intervention.[1]

Seventy-one students completed the randomised comparison

The team recruited 97 students aged 18 to 25 from Beijing Normal University. Following a pre-test, 71 who met predefined criteria were included: they had mild-to-moderate anxiety or depression and scored above five on willingness to seek psychological support. Thirty-six were assigned to DHCBT and 35 to the text-only condition. The gender counts were closely balanced within both groups, although the paper does not establish representativeness beyond this single university.

Participants gave written informed consent and the Beijing Normal University ethics committee approved the study under protocol 202406210122. After their assigned single session, students completed scales covering cognitive load, social presence, system usability and intention to keep using the technology, then joined a short interview. The PHQ-9 and GAD-7 were part of screening, but the results section does not report a pre-to-post symptom comparison. That omission prevents claims that either system reduced depression, anxiety or academic stress.[1]

The avatar improved experience scores, not a clinical endpoint

On seven-point cognitive-load scales, the digital-human group reported lower intrinsic load, with a mean of 3.67 versus 4.76 for text, and lower extraneous load, 3.04 versus 4.34. The reported standardised differences were large and the two-sided tests had p values below 0.001. System Usability Scale scores averaged 73.48 for DHCBT and 62.78 for text; technology-use intention averaged 4.67 versus 3.23, again with reported p values below 0.001.

Social-presence scores were higher with the avatar, 7.12 versus 6.82, but the reported p value was 0.062. Under the usual 0.05 convention, that comparison is not statistically significant. The paper’s abstract and discussion sometimes group social presence with the system’s advantages, so readers should keep the actual result in view. Self-reported intent to reuse is also not observed continued use, and usability is not evidence that advice is accurate or improves mental health.

Interview responses were generally positive: 80.5% were described as favourable overall, 86.1% reported substantial help and 52.8% mentioned response speed as an area to improve. These percentages are perceptions gathered immediately after the encounter. The publication does not describe blinded qualitative coding, a preregistered interview analysis or follow-up to determine whether positive impressions persisted.[1]

Benchmark results broaden the evaluation but bring new biases

Separate from the user experiment, the multi-agent backend was tested on three CBT-BENCH classification sets containing 146, 184 and 112 cases. It was also compared with DeepSeekV3-1, GLM-4.5 and CBT-LLM on 60 academic-stress prompts drawn from 120 de-identified posts collected on two public Chinese counselling platforms. Nine psychology-counselling students—three undergraduates, three master’s students and three doctoral students—rated anonymised responses for safety, CBT process, problem comprehension and empathy.

The proposed system had the highest aggregated human ratings across all four dimensions and overall inter-rater reliability of 0.750. Reliability for safety alone was only 0.475, showing substantial disagreement on the dimension with the greatest potential consequences. Automatic ratings used GPT-5 and Claude Opus 4.1, creating a risk that language models reward styles similar to their own. Neither benchmark demonstrated safe behaviour in unscripted, prolonged or acute real-world use.[1]

The safety boundary is narrower than many crises

DHCBT screens inputs for offensive or illegal material and explicit expressions of suicide or self-harm. When it detects explicit high risk, the routine flow stops and the user is directed to a crisis hotline or immediate professional support. Generated replies also undergo a compliance review and can be regenerated. Those are meaningful design controls, but the authors state that the current mechanism does not reliably identify subtler crisis signals such as hopelessness without direct self-harm language.

That limitation is especially important because the enrolled students had mild-to-moderate symptoms, not a verified absence of evolving risk. An engaging human-like avatar may increase disclosure and trust, but it may also make a system’s limitations less visible. Any deployment would need validated risk escalation, local emergency information, clinician oversight where appropriate, clear disclosure that the system is not a person, privacy controls for sensitive speech and transcripts, and monitoring for harmful or overconfident advice.[1]

What would establish benefit rather than appeal

The next trial should be preregistered, larger and multi-site, with allocation concealment, attrition reporting and outcome assessors insulated from condition where possible. It should measure academic stress, anxiety, depression, functioning, adverse events and help-seeking over weeks or months—not only experience immediately after one session. Comparators should include ordinary support or clinician-guided digital CBT where ethically appropriate, while factorial designs could separate voice, facial animation, embodiment and multi-agent workflow.

The authors are affiliated with Beijing University of Posts and Telecommunications and China Mobile Group Device. Chinese national research funds supported the work, and the authors declared no competing interests. The dataset is available only from the corresponding author on reasonable request. The study supports a preliminary claim that one avatar interface felt easier and more usable than the same text backend for selected students. It does not yet show that a digital human delivers effective therapy or safely substitutes for qualified care.[1]

What this means for people

  • Students may find an embodied interface easier to use, but a more human-like presentation can also encourage trust beyond the system’s evidence or competence.
  • Universities should not treat immediate usability or reuse intentions as proof of mental-health benefit or a replacement for counselling staff.
  • People in crisis need reliable routes to qualified and emergency support; subtle warning signs remain a known gap in this prototype.

Global context

The participants were selected young adults at one Beijing university and the response materials came from Chinese-language counselling contexts. Cultural expectations about avatars, help-seeking and mental-health disclosure may shape the results. Replication should include different ages, institutions, languages, disability needs and health systems, while localising crisis escalation and privacy safeguards rather than assuming one interface or workflow transfers globally.

What the evidence does not yet show

  • The user study included 71 selected students from one Beijing university and lasted one session.
  • It reported experience measures, not improvement in anxiety, depression, academic stress, functioning or longer-term help-seeking.
  • There was no human-therapist, usual-care or no-intervention comparator, and the component effects of voice, avatar and workflow were not separated.
  • Automatic response ratings used other language models, while human inter-rater reliability for safety was only 0.475.
  • The system was designed around explicit crisis language and may miss indirect or subtle signs of acute risk.

What to watch next

  • Preregistered multi-site trials with symptom, functioning, adverse-event and follow-up outcomes.
  • Independent safety testing using indirect crisis language, adversarial inputs and prolonged conversations.
  • Factorial studies that separate the effects of avatar, voice, facial cues, CBT workflow and knowledge retrieval.
  • Clear governance for consent, sensitive conversation data, escalation to human help and disclosure of system limitations.

Evidence trail

Sources used for this report

Links checked 4 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Health & Life Sciences

Can AI estimate survival after a Parkinson’s diagnosis?

A peer-reviewed Chinese registry study compared four survival models in 3,148 people with Parkinson’s disease. A transparent Cox model matched the machine-learning alternatives, but validation stayed within the same registry and no clinical-impact study was performed.

9 min · 2 sources

Health & Life Sciences

Can an ICU glucose digital twin be trusted after testing on ten patients?

A peer-reviewed study reports that a pretrained transformer can produce 15- and 30-minute glucose forecasts for ten septic ICU patients within seconds. The result is technically useful, but it is a retrospective forecast test—not evidence that the system improves insulin decisions or patient outcomes.

10 min · 1 source

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.