Are human–AI digital twins ready for real factories?
A peer-reviewed systematic review screened 569 records and included 82 studies of human–AI digital twins in manufacturing. Most evidence remained in laboratories or simulations; only six studies reached industrially relevant validation, and worker-data ethics were rarely addressed.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
Research topic
How mature the evidence is for human–AI-assisted digital twins in manufacturing, including industrial validation, worker impact and data-governance coverage

At a glance
- 1The authors screened 569 unique records from three databases and included 82 full-text publications after applying explicit manufacturing, digital-twin and human-collaboration criteria.
- 2Laboratory and simulation studies dominated. Only six studies reached industrially relevant validation, and no reviewed system reached the highest technology-readiness levels.
- 3The review found weak reproducibility, few head-to-head comparisons and sparse treatment of consent, privacy or ethics for worker-state data; it does not show that digital twins already improve jobs or production at scale.
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Living evidence record
Impact record IAI-0AQ25BO
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent support
Present
Record status
Monitoring
Last checked
4 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what The International Journal of Advanced Manufacturing Technology published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
The review asks whether the human side of a digital twin is operational
A manufacturing digital twin links a physical process to a changing virtual representation. The ambitious version does more than mirror a machine: it uses sensor data, simulation and AI to help a worker, robot or supervisory system decide what should happen next. Industry 5.0 language adds a further promise—that these systems will be human-centred, resilient and sustainable rather than simply more automated.
The new review tests how much published evidence supports that promise. Its focus is the junction between digital twins and human–AI or human–robot collaboration: task allocation, shared autonomy, operator-state sensing, extended-reality interfaces and AI reasoning. That makes it directly relevant to workers whose movements, fatigue, comfort or decisions may become inputs to an industrial model, as well as to managers deciding whether a laboratory prototype is ready for a production line.[1]
The denominator is 82 studies from 569 screened records
The authors followed PRISMA 2020 and searched Scopus, Web of Science Core Collection and IEEE Xplore, with a common July 2026 cut-off. They screened 569 unique records and excluded 423 at title and abstract stage. Of 146 reports sought for retrieval, 38 were not taken through full-text assessment. A further 26 of the 108 assessed reports were excluded, leaving 82 publications in the final synthesis.
Inclusion required an explicit digital-twin implementation, framework, evaluation or analysis; structured integration between physical and virtual entities; a human–AI or human–robot collaboration element; and a manufacturing or industrial-production setting. Each included paper was coded against a seven-dimension quality rubric, a nine-level technology-readiness scale, validation environment, AI-method family and ethics coverage. The authors provide per-paper classifications in supplementary material, making the review more auditable than a narrative overview.[1]
Most systems remain close to the laboratory
Laboratory validation at technology readiness level four was the dominant category. Seventy-eight of the 82 papers received a readiness level; four reviews or syntheses did not. Only six studies crossed into industrially relevant validation at level six or higher, comprising four industrial pilots and two systems combining laboratory work with an operational setting. No reviewed study reached levels eight or nine, which would indicate a completed and proven operational system.
That distribution matters because shop-floor constraints can invalidate an attractive prototype. Operators may not have time to monitor and correct an AI agent within a production cycle. Wearable or biometric sensors that work in a study may be intrusive, unreliable or unacceptable during a full shift. Simulator-trained controllers may encounter edge cases that were absent from training. The review finds that these conditions are discussed more often than they are tested longitudinally in real workplaces.[1]
Reproducibility and comparison are weaker than the vision
On the authors' zero-to-three rubric, problem-definition clarity was the strongest dimension, averaging 2.67. Reproducibility was the weakest at 1.68, followed by comparative analysis at 1.80. Few papers supplied both code and data, and shared benchmarks were uncommon. Conference papers scored lower on average than journal articles—11.9 versus 16.8 out of 21—although publication format and reporting space may partly explain that difference.
The practical gap is not simply whether a model can produce a prediction. Manufacturers need to compare systems on safety, cycle time, workload, recovery from error and maintenance cost under the same conditions. The review reports no rigorous head-to-head comparison of deep learning, reinforcement learning, knowledge graphs and other approaches under industrial constraints. Without common tasks and reporting, a strong result in one laboratory cannot tell a buyer which architecture is safer or more useful on another line.[1]
Worker benefit and worker surveillance cannot be separated
Human-state-aware twins could help adjust a workstation, predict fatigue, guide maintenance or share tasks with a robot. Those benefits depend on collecting data about bodies and behaviour. The review says fewer than one in ten studies explicitly addressed informed consent, privacy or institutional review. It also found little evidence on whether workers would accept continuous sensing or whether performance estimates remain fair across demographic groups and job roles.
For employees, the distinction between support and surveillance is therefore central. A system that adapts work to reduce strain may use the same signals that an employer could use to rank pace or infer vulnerability. Useful deployment needs clear purpose limits, worker participation, access controls, retention rules and a route to contest automated inferences. The review does not measure job satisfaction, injury reduction, displacement or bargaining outcomes, so it cannot resolve whether the human-centred label is deserved in practice.[1]
What would make a factory claim credible
A stronger evidence base would include preregistered industrial pilots across multiple sites, representative worker samples and comparisons against the existing process rather than against a weak technical baseline. Studies should report failures, near misses, operator overrides, production downtime, data quality, cyber risk and total cost. Longitudinal follow-up is needed because novelty effects and intensive researcher support can make a short pilot look better than routine use.
The review itself has limits. Thirty-eight potentially relevant reports were not fully retrieved, four non-English full texts were excluded and the search covered three databases, which may bias the corpus toward accessible English-language engineering work. Classification involves judgement, and the field has moved since the July cut-off. The authors disclosed European Union and Swedish Knowledge Foundation support, no competing financial interests, and use of generative AI for drafting and language work that they say was checked against sources. The result is a useful maturity map, not proof that a particular digital twin works.[1]
What this means for people
- Factory workers could gain safer, more adaptive tools, but the same data can enable intrusive monitoring if governance is weak.
- Engineers and buyers gain a clearer warning that laboratory accuracy is not operational readiness.
- Employers and unions need evidence on workload, autonomy, fairness and job quality before treating Industry 5.0 claims as worker benefit.
Global context
The authors are based in Sweden, while the 82-paper corpus includes manufacturing research across Europe and Asia. Industrial digital twins are promoted globally, but only six reviewed studies reached industrially relevant validation. Excluding non-English full texts and relying on three databases may underrepresent work from regions where manufacturing deployment is substantial, so the maturity pattern should be tested against additional languages and local evidence.
What the evidence does not yet show
- The review synthesises publications rather than generating new factory or worker outcomes.
- Thirty-eight reports sought at retrieval were not included, four non-English full texts were excluded and database coverage may miss relevant work.
- Technology-readiness and quality classifications involve reviewer judgement and do not replace independent system testing.
- Only six studies reached industrially relevant validation; no system reached the highest readiness levels.
- Worker acceptance, job quality, safety, displacement and long-term performance were rarely measured.
What to watch next
- Multi-site, preregistered industrial trials with worker participation and longitudinal outcomes.
- Common benchmarks that report safety, overrides, workload, downtime, maintenance and total cost.
- Explicit consent, privacy and governance rules for biometric and behavioural worker data.
- Independent replications of the six most operationally mature systems.
Evidence trail
Sources used for this report
Links checked 4 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Work & Skills
Who is Anthropic training to deploy Claude?
Anthropic says it will spend $100 million to train 10,000 nominated enterprise engineers by the end of 2027. The residency includes practical assessment and a 12-week workplace project, but the company has not reported completion, retention, safety or business-outcome evidence.
9 min · 1 source
Work & Skills
Does AI literacy keep people engaged at work?
A peer-reviewed Chinese study followed 1,108 employees for a year and found that higher baseline AI literacy predicted later work engagement. The design strengthens the timing evidence, but it does not show that AI training caused the difference.
8 min · 2 sources
Work & Skills
Who is workplace AI leaving behind? PwC finds an access gap—but not its cause
PwC surveyed 49,364 workers across 48 countries and regions. AI use rose, while reported access to learning fell; the cross-sectional responses show an association, not proof that AI caused the divide.
5 min · 2 sources
Reader discussion
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.