Back to the news portal
Society & MediaNew analysis today · source 7 October 2026Research paperResearchSource analysisSouth KoreaUnited KingdomGlobal

Can AI carry emotion across languages?

Partly, in one English-to-Korean experiment. GPT-4o-generated text and images produced responses aligned with intended emotions above chance, but same-language reference ratings were closer for 11 of 18 constructs.

By The Impact of AI Editorial DeskReleased 8 October 2026 at 04:42 BST8 min read1 source

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

Share
Social links
LinkedInXBlueskyRedditEmail

At a glance

  • 1The study combined affect ratings from 547 English-speaking participants with GPT-4o to create 180 text and 18 image stimuli, then collected ratings from 91 Korean-speaking participants and a separate 76-person Korean reference cohort.
  • 2Responses to the generated material aligned with the intended emotion better than a non-corresponding-emotion null for all 18 constructs and both formats, but no construct exceeded the authors' stricter semantic-distinctness threshold.
  • 3Same-language Korean reference ratings were closer than the English-derived reference for 11 of 18 constructs in at least one format, so the experiment supports measurable but incomplete transfer—not universal emotional equivalence.
Key themesCross-lingual AIAffective computingGenerative mediaMental health researchCultural contextModel evaluation

Research topic

Whether LLM-generated text and images preserve intended emotional structure across English-speaking and Korean-speaking populations

The Impact of AI cover showing conceptual emotion cards passing through an AI prism, with recognisable but imperfectly aligned responses on the other side.
AI-generated editorial illustration. The emotion cards and AI prism are conceptual; they do not represent clinical diagnosis, a real interface or a measured participant response.

The direct answer: emotional direction travelled, but nuance did not transfer intact

A Korean research team found that GPT-4o could transmit a recognisable emotional pattern from English-derived ratings into generated text and images, but it did not erase the gap between language communities. For all 18 tested emotion constructs, responses from Korean-speaking participants were more aligned with the intended construct than with a bootstrapped reference made from non-corresponding emotions. That is evidence of non-random transfer. Under the authors' stricter test, however, none of the constructs in either text or image format reached the threshold at which emotion representations became difficult to distinguish within the source cohort.

The clearest real-world comparison used a separate Korean-speaking cohort as a same-language reference. For 11 of the 18 constructs, alignment to that Korean reference exceeded alignment to the English-speaking source in at least one format. Three exceptions—surprise in images, and righteous anger and moral disgust in text—favoured the cross-lingual comparison, but the overall pattern was residual divergence. The practical answer is therefore not that AI translates emotion accurately or inaccurately in the abstract. It is that the model preserved some structure in this controlled task while leaving systematic room for linguistic and cultural context.[1]

How the three cohorts and 198 generated stimuli fit together

The source data came from 547 English-speaking participants recruited through Prolific. They rated 18 constructs—including six basic emotions and additional states such as satisfaction, relaxation, shame, guilt, pride, gratitude and moral disgust—on four dimensions: valence, arousal, self-versus-other focus and dominance. The researchers averaged those responses to create one four-dimensional reference vector for each construct. They then supplied the labels and empirical distributions to the GPT-4o-2024-08-06 model snapshot with temperature and top-p both set to 1.0.

For each emotion, the model produced 50 candidate text passages. Two raters knew the target label but were blinded to its numerical affect coordinates; they scored relevance and selected ten texts per construct, producing 180 passages. The image process produced one selected image per construct, for 18 images. A Korean-speaking evaluation cohort of 91 adults viewed each image and text stimulus for ten seconds in randomised blocks, without seeing the emotion label, and rated their own response on the same four dimensions. A separate Korean-speaking validation cohort of 76 rated the 18 emotion labels directly, without seeing AI-generated material, to provide the same-language benchmark.[1]

The analysis separated 'better than chance' from 'close enough'

The researchers used cosine similarity to compare the direction of four-dimensional affect vectors while analysing response magnitude separately. First, they constructed an empirical null by repeatedly sampling vectors from the other 17 emotions and applied false-discovery-rate correction across the 18 constructs. All observed alignments for text and images cleared that non-randomness test. This result matters, but its scope is narrow: it shows that participants' rating patterns corresponded to the intended emotions better than shuffled alternatives; it does not show accurate translation of every word, equivalent subjective experience or clinical usefulness.

Second, the team created a semantic-distinctness threshold from the 75th percentile of all 153 pairwise similarities among the source cohort's emotion vectors. No cross-lingual comparison exceeded that threshold. The authors also caution that failing to find a significant difference from a threshold is not proof of equivalence, and they report sensitivity analyses at other percentiles. Third, paired comparisons tested alignment against the English source and the independent Korean reference for the same participant responses. That same-language benchmark is what exposed the remaining gap for 11 constructs. The layered design is valuable precisely because a simple above-chance result would have sounded much stronger than the full analysis supports.[1]

A mental-health association emerged, but it is not a diagnostic result

Participants in the Korean evaluation cohort also completed six psychological measures. Linear mixed-effects models examined whether those scores were associated with the magnitude and redundancy of responses to the generated text. In models considering measures separately, interoceptive awareness and depressive symptoms produced some associations or trends. After the predictors were entered together, sensitivity to prosocial emotions on the PFQ-2 was the only measure associated with lower response magnitude and less redundant responses, with the redundancy association surviving the authors' multiple-testing correction and the magnitude result reaching only trend level.

That pattern should not be turned into a screening claim. The outcome was a response to researcher-selected generated stimuli, not a diagnosis, symptom improvement or comparison with standard clinical assessment. The cohort was recruited online rather than from a diagnostically diverse clinical population. The authors themselves call for independent replication and ask whether behavioural responses to generated material predict anything beyond existing self-report measures. Two authors have a patent pending that is assigned to KAIST, which is relevant when interpreting proposed future applications. The work was supported by the Korean Neuropsychiatric Research Foundation and Korea's Ministry of Science and ICT.[1]

What this changes for localisation and affect-sensitive systems

For translated health messages, educational material, social robots and affect-sensitive media, the study supplies a useful warning: preserving topic and grammar does not guarantee preservation of emotional structure. A team should validate generated material with people from the target linguistic and cultural setting, and compare it with a same-language baseline rather than treating an English-derived prompt as ground truth. The result also argues for reporting which model snapshot, generation parameters, human-selection process and language pair were used. A later model or different curation rule could produce different stimuli and different response patterns.

The experiment remains a proof of concept. It tested one English-to-Korean direction, one proprietary model snapshot, subjective ratings rather than physiological responses, and short researcher-curated stimuli rather than natural conversation. The two raters selected high-relevance outputs while knowing each target label, which can improve the final set relative to unattended generation. Stronger evidence would replicate the design across language families, include clinical and non-clinical populations, compare several models and professionally curated material, add autonomic or neural measures, and pre-specify an equivalence margin. Until then, the paper establishes a measurable channel with boundaries—not a universal emotional translator or a new mental-health test.[1]

What this means for people

  • People receiving translated or generated emotional content may recognise its broad intent while experiencing important nuance differently from the source population.
  • Mental-health and education teams should not assume that emotionally targeted material validated in one language remains equivalent after AI-mediated localisation.
  • Developers can use same-language reference groups and transparent model settings to detect some of the loss before deploying affect-sensitive systems.

Global context

The study links an English-speaking online source cohort with Korean-speaking evaluation and validation cohorts, making cultural and linguistic context part of the measured phenomenon rather than an afterthought. Its results cannot establish a single global hierarchy of emotional fidelity: Korean and English differ in vocabulary, grammar, social conventions and cultural expression, while other language pairs may be closer or further apart. Globally deployed systems should therefore treat affective localisation as a validation problem for each target population, not a one-time translation step.

What the evidence does not yet show

  • The experiment covered one translation direction—English-derived source ratings to Korean-speaking respondents—and may not generalise to other language or cultural pairs.
  • Only the GPT-4o-2024-08-06 snapshot was tested, using one generation configuration and a human-curated subset of outputs.
  • Outcomes were self-reported affective ratings; no physiological, behavioural or clinical outcome validated equivalent emotional experience.
  • The Korean evaluation cohort contained 91 participants and the separate same-language reference cohort 76; larger and diagnostically diverse replication is needed.
  • The study did not compare generated stimuli directly with a professionally curated matched stimulus set.
  • Two authors disclosed a patent pending and assigned to KAIST. Funding came from the Korean Neuropsychiatric Research Foundation and Korea's Ministry of Science and ICT.

What to watch next

  • Independent replication across additional language pairs, cultures and model families with pre-specified equivalence margins.
  • Direct comparisons between LLM-generated stimuli and expert-curated material matched on affective coordinates.
  • Studies adding behavioural, autonomic or neural measures rather than relying only on subjective ratings.
  • Evidence from clinical populations showing whether generated-stimulus responses add useful information beyond established questionnaires.
  • Whether results hold without researchers selecting the highest-rated outputs from a larger candidate pool.

Living evidence record

Impact record IAI-1QLCH5T

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent or research support

Present

Record status

Monitoring

Last checked

8 October 2026

Source trail

1 direct source across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Single-source reporting disclosure

This record analyses one direct source. It can establish what Scientific Reports published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.

Evidence trail

Sources used for this report

Links checked 8 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Society & Media

What is changing in ChatGPT for teens?

OpenAI says a US College Planner for grades 10–12 is coming, alongside flashcards, easier quizzes and multi-photo note capture. The company also released large usage counts, but no study protocol, denominators for several comparisons or evidence that the tools improve learning or admissions outcomes.

7 min · 1 source

Society & Media

When should newsrooms disclose AI use?

A peer-reviewed case study based on 13 interviews with 12 Financial Times managers and 28 internal documents finds that AI disclosure is treated as a spectrum shaped by oversight, risk and context. It describes one newsroom's practice and does not test whether labels improve audience trust.

7 min · 1 source

Society & Media

What do people actually ask image-upload AI to do?

An unreviewed Microsoft-led study analysed 42,617 de-identified Copilot image-upload sessions and checked its taxonomy against 23,413 ChatGPT sessions. Users often chained recognition into writing, code and data tasks—but private, model-generated summaries limit what outsiders can verify.

8 min · 2 sources

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.