Can AI estimate survival after a Parkinson’s diagnosis?
A peer-reviewed Chinese registry study compared four survival models in 3,148 people with Parkinson’s disease. A transparent Cox model matched the machine-learning alternatives, but validation stayed within the same registry and no clinical-impact study was performed.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
Research topic
Whether clinical, genetic and treatment information can support interpretable estimates of all-cause mortality after a Parkinson’s diagnosis

At a glance
- 1The prospective Chinese Parkinson’s Disease Registry analysis included 3,148 patients enrolled at 19 tertiary hospitals from 2018 to 2020 and followed through 31 December 2024; 562 deaths were recorded and 146 people were lost to follow-up.
- 2Cox regression, random survival forest and XGBoost produced the same mean C-index of 0.716 in repeated cross-validation. The researchers selected Cox because it was similarly discriminative, easier to interpret and able to produce transparent absolute-risk estimates.
- 3A later-enrolment cohort of 913 people provided temporal validation, not independent external validation. The model’s C-index was 0.708, and it modestly underestimated observed mortality at two, three and four years.
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Living evidence record
Impact record IAI-1ARW1TR
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent support
Present
Record status
Monitoring
Last checked
3 October 2026
Source trail
2 direct sources across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
The question was prognosis, not diagnosis
This study did not ask whether an algorithm can detect Parkinson’s disease. It asked whether baseline information collected after diagnosis can estimate a person’s subsequent risk of death from any cause. Prognosis is difficult because Parkinson’s trajectories vary with age, disease duration, movement impairment, cognition, depression, falls, other illness, treatment and genetics.
The researchers used the Chinese Parkinson’s Disease Registry, a prospective multicentre cohort. The analytic sample contained 3,148 people enrolled at 19 tertiary hospitals between 2018 and 2020 and followed until death, loss to follow-up or 31 December 2024. Median follow-up was 5.33 years. There were 562 recorded deaths, 2,440 people alive at last contact and 146 lost to follow-up.
All-cause mortality is a clear endpoint, but it combines deaths related to Parkinson’s with cancer, cardiovascular disease, infection and other causes. A model can therefore describe associations in the registry without showing that any one Parkinson’s intervention will extend life. The intended use is risk stratification and discussion, not diagnosis or a recommendation to start, stop or intensify treatment.[1][2]
Four modelling approaches were compared on the same predictors
The final predictor set contained 11 variables available at baseline: age at onset, disease duration, GBA1 mutation status, type 2 diabetes, deep-brain-stimulation status, levodopa-equivalent daily dose, UPDRS Part III motor score, Hoehn and Yahr stage, history of falls, depression and cognitive dysfunction. Candidate sets were selected inside the resampling process rather than after looking at the full validation results.
Missingness was generally low, although depression was missing for 193 participants, or 6.13%. The team created 20 multiply imputed datasets. Within each outer training fold, imputation used training data only, while random-survival-forest and XGBoost tuning happened in inner folds. Five-fold cross-validation was repeated five times. That separation reduces a common form of information leakage in which validation data influence preprocessing or model selection.
The comparison covered a conventional Cox proportional-hazards model, a random survival forest, a survival tree and XGBoost adapted for survival analysis. Cox, the forest and XGBoost each reached a mean Harrell C-index of 0.716. Their two-, four- and six-year area-under-the-curve estimates were also close. The survival tree was weaker, with a mean C-index of 0.690.
A C-index of 0.716 means that, for roughly 72% of comparable patient pairs, the model ranked the person who died sooner as higher risk. It does not mean a 71.6% chance that an individual forecast is correct, and it says nothing by itself about whether the absolute percentage shown to a patient is well calibrated. Because the more complex methods did not deliver a meaningful discrimination gain, the authors chose Cox for its interpretability and reproducible absolute-risk calculation.[1]
The later cohort preserved ranking but exposed underestimation
For an additional test, the researchers trained the final Cox specification on 2,235 people enrolled in 2018 and 2019, including 428 deaths, then evaluated it on 913 people enrolled in 2020, including 134 deaths. This split was made before imputation and fitting. The later cohort produced a C-index of about 0.708 and time-dependent AUCs of 0.709, 0.692 and 0.725 at two, three and four years.
Discrimination held up reasonably well, but calibration was not perfect. Observed mortality was 3.14% at two years compared with 2.70% predicted, 6.72% versus 6.01% at three years and 10.67% versus 8.62% at four years. Observed-to-expected ratios of 1.17, 1.12 and 1.24 show modest underestimation, greatest at four years.
This is a temporal validation within the same national registry, not a genuinely independent external test. The same study network, clinical definitions and data-collection system connect the development and later cohorts. Validation in community hospitals, other Chinese regions and populations outside China could reveal different baseline mortality, care pathways and relationships between predictors and outcome.[1]
The web tool separates language-model assistance from risk calculation
The selected Cox model was implemented in a Streamlit web application that returns two-, four- and six-year survival and mortality estimates after the 11 predictors are entered. The numerical calculation is deterministic: pooled coefficients and baseline-survival values generate the result. A large language model does not calculate the mortality risk.
Separate DeepSeek-based modules can extract the 11 fields from a de-identified case summary and provide general education. Extracted values must be reviewed and confirmed before prediction. In a test of 50 constructed, de-identified clinical-style summaries—550 fields in total—the extraction module matched the adjudicated references on every field, including uncertainty handling.
That perfect result should not be mistaken for proven clinical safety. Fifty curated summaries cannot represent messy abbreviations, contradictory notes, unusual medication histories, multiple languages or copied-forward errors in real health records. The authors call the evaluation preliminary. A prospective study should measure silent extraction errors, correction time, usability, privacy controls and whether clinicians become over-reliant on a plausible pre-filled value.[1]
Several associations should not be read as treatment effects
In the pooled Cox model, age at onset above 50 was associated with higher mortality, as were longer disease duration, type 2 diabetes, higher levodopa dose, worse motor and disease-stage scores, falls, depression and cognitive dysfunction. Deep-brain-stimulation status was associated with lower risk. GBA1 carrier status showed a non-significant trend toward higher risk.
These coefficients describe prediction in an observational registry. They do not establish that deep-brain stimulation prevents death or that changing medication dose changes mortality. Treatment selection is influenced by age, disease severity, fitness, access and clinical judgement; those same factors can affect survival. The paper explicitly treats medication variables as baseline predictive markers rather than causal treatment effects.
One proportional-hazards diagnostic also warrants attention. The global test was not statistically significant, but Hoehn and Yahr stage showed evidence that its hazard relationship changed over time. That does not invalidate the entire tool, but it is another reason to test calibration across follow-up periods and settings rather than assuming one fitted relationship remains stable everywhere.[1]
The evidence remains one registry’s internally validated model
The registry’s size, 19-centre recruitment, 562 events, detailed neurological measures and careful training-fold-only imputation are meaningful strengths. Reporting followed TRIPOD+AI, analysis code is supplied as supplementary material, and the authors state that de-identified supporting data may be requested subject to ethics and data-use agreements. The dataset itself is not public because of privacy and Chinese data-protection requirements.
The limits are equally important. All sites were tertiary Grade-A hospitals, contributions varied by centre and some sites recorded too few events for robust centre-level validation. Major comorbidities such as ischaemic heart disease and malignancy were not incorporated, while the registry did not systematically capture the severity of broader health conditions. Most of the 146 losses to follow-up occurred in the first year and were censored at last contact rather than counted as deaths.
The Hunan Innovative Province Construction Project and the National Natural Science Foundation of China funded the work. The authors declare no competing interests. The available publication is a peer-reviewed accepted manuscript received on 7 July and accepted on 15 September 2026; the publisher says editing may still change presentation before the version of record.[1][2]
What would make the estimate clinically actionable
The next test should freeze the model and evaluate it in genuinely independent cohorts, including community care, other Chinese regions and countries with different mortality and treatment patterns. Researchers should pre-specify calibration targets, report error by age, sex, disease stage and care setting, and examine whether recalibration is needed before a percentage is shown to patients.
A clinical-impact study would then compare usual specialist care with care supported by the tool. It should measure whether forecasts improve follow-up planning and communication without increasing distress, fatalism, unnecessary intervention or inequity. The study should also test how clinicians explain uncertainty and what happens when input data are missing or the model conflicts with clinical judgement.
For now, the practical conclusion is measured. Detailed registry data can produce a moderately discriminative, interpretable survival model, and more complex machine learning did not outperform a conventional Cox approach. That is useful evidence against complexity for its own sake. It is not yet evidence that placing the calculator or its language-model intake layer into routine care improves decisions or outcomes.[1][2]
What this means for people
- Patients may eventually receive more structured discussions about prognosis, but the current model should not determine treatment or be presented as an individual life expectancy.
- Clinicians gain evidence that an interpretable statistical model can match more complex machine-learning approaches on this task, while retaining clearer assumptions and outputs.
- Hospitals considering the web tool would still need independent validation, privacy review, workflow testing, human confirmation and monitoring for calibration drift.
Global context
The cohort comes from 19 tertiary hospitals in China, so the study adds substantial Asian clinical evidence to a field often dominated by Western datasets. Its transportability is unproven: background mortality, comorbidity, genetics, treatment access and care pathways differ within China and across countries. External validation should include community and regional hospitals as well as populations in other parts of Asia, the Middle East, Europe, Africa and the Americas before the absolute-risk estimates travel.
What the evidence does not yet show
- The later-enrolment evaluation came from the same Chinese registry network; there was no truly independent external validation.
- All participating sites were tertiary hospitals, several contributed few patients or deaths, and generalisability to community care and other populations is unknown.
- The observational predictors include treatment variables and must not be interpreted as causal treatment effects.
- Important comorbidities and general-health severity were not systematically available, and 146 participants were lost to follow-up.
- The language-model intake evaluation used only 50 curated de-identified summaries and did not test real clinical workflow, privacy, usability or downstream harm.
What to watch next
- Independent geographic and centre-level validation with recalibration where baseline mortality differs.
- Prospective workflow testing against usual care, including communication quality, clinician correction burden and patient-relevant outcomes.
- Robustness when notes are incomplete, contradictory or multilingual and when the language-model layer encounters unfamiliar phrasing.
- The final version of record and any changes to model coefficients, calibration results, supplementary code or deployment claims.
Evidence trail
Sources used for this report
Links checked 3 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Health & Life Sciences
Can brain MRI reveal more than BMI?
A peer-reviewed study trained a deep-learning model on 45,702 MRI scans from six cohorts. Brain-derived features tracked BMI and separated several disease groups better than BMI alone—but the disease analysis stayed inside UK Biobank and cannot establish cause, diagnosis or clinical benefit.
9 min · 2 sources
Health & Life Sciences
Can this MRI model draw brain-tumour boundaries reliably?
A peer-reviewed multimodal segmentation model was developed on 2,422 public MRI volumes and externally tested on 125 cases. Accuracy remained useful but fell outside the development data; no prospective clinical workflow, radiologist comparison or patient-outcome test was performed.
8 min · 2 sources
Health & Life Sciences
Can AI support breast-ultrasound decisions across countries?
A peer-reviewed South Korean-led study validated an interpretable retrieval-augmented system across 8,311 images from 11 cohorts in seven countries and tested assistance with four readers. The retrospective evidence is encouraging, but it is not a prospective screening trial or proof of better patient outcomes.
8 min · 3 sources
Reader discussion
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.