Back to the news portal
Science & ResearchResearch paperResearchSource analysisNorth AmericaGlobal research

Can more data fix measurement errors in AI?

No, not in these simulations. Across five model families, noisy or misassigned input features reduced predictive performance and distorted feature-importance rankings; larger samples narrowed variation but did not remove the bias.

By The Impact of AI Editorial DeskReleased 11 October 2026 at 08:03 BST7 min read1 source

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

Share
Social links
LinkedInXBlueskyRedditEmail

At a glance

  • 1The researchers ran regression and classification simulations with sample sizes of 2,000, 3,000 and 4,000, six signal structures, two non-zero error levels and five model families.
  • 2Classical noise, categorical misclassification, Berkson error and group-summary assignment generally reduced predictive performance and changed the apparent importance of the affected feature.
  • 3Larger samples reduced variation between simulation runs but did not remove systematic bias, so collecting more records cannot compensate for a poorly measured predictor.
Key themesMeasurement errorMachine learningData qualityModel interpretationFeature importanceStatistical simulation

Research topic

How four common forms of error in predictive features affect accuracy and permutation feature importance across linear, kernel, tree-ensemble and neural-network models

The answer: more records made the estimates steadier, not the bad measurements truthful

Across the controlled experiments, substituting an error-burdened feature for its error-free counterpart usually lowered regression R-squared, increased mean squared error and reduced classification accuracy or F1. The direction held across multiple linear or logistic regression, support-vector machines, random forests, XGBoost and multilayer perceptrons. Increasing the simulated sample from 2,000 to 4,000 observations mostly narrowed the spread between runs; it did not eliminate the bias introduced by the faulty input.

That distinction matters for organisations accumulating large operational datasets. A large file can make an estimate look precise while the underlying measurement remains systematically wrong. If a hospital records a proxy instead of a patient's true exposure, a lender relies on a misclassified category, or an employer assigns a group average to an individual, model tuning cannot reconstruct information that was never measured correctly. The paper therefore supports data-quality audits before treating scale, architecture or hyperparameter search as a remedy.[1]

The Impact Brief · Free

Follow the evidence in science & research.

Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.

Choose your topics (optional)

One concise, source-linked briefing. Unsubscribe at any time.

What the simulations actually tested

The authors created synthetic regression and binary-classification problems with three features. Two features generated the outcome and one was deliberately non-predictive. They then replaced one predictive feature with an error-affected version while retaining an error-free benchmark. The study examined classical additive error, binary misclassification, Berkson error—where actual values vary around assigned values—and group-summary assignment, where five group means stand in for individual values.

Sample sizes were 2,000, 3,000 and 4,000. Six coefficient settings changed whether the error-prone feature was stronger than, equal to or weaker than the other predictive feature. Classical and Berkson error used standard-deviation levels of 0.5 and 0.75; binary misclassification used probabilities of 0.1 and 0.3. Each factorial cell was repeated ten times. Data were split 80:20 for training and testing, with stratification for classification, and hyperparameters were tuned separately for each model and measurement condition using 20 Optuna trials.[1]

Why feature importance can become misleading

The researchers calculated permutation feature importance with ten permutations for every fitted model. As classical error increased, the measured importance of the affected predictor generally fell, while the clean predictive feature stayed stable and the deliberately irrelevant feature remained near zero. In practical terms, a genuinely influential factor can appear less important merely because it is measured less reliably. Rankings can then reward the easiest variable to record rather than the variable most connected to the outcome.

This is not only a technical concern. Feature rankings are often shown to clinicians, auditors, regulators and managers as explanations for model behaviour. If measurement quality differs across demographic groups, sites or service settings, an explanation can also shift without any true change in the underlying relationship. The paper evaluates permutation importance rather than every explanation method, but its core warning travels: explanation tools describe the fitted data-generating pipeline, including its errors. They do not independently verify that the inputs represent reality.[1]

The model family did not provide a universal escape

Model architecture affected the size of the degradation, and interactions between model, error level and signal structure were statistically significant in many analyses. But no model family was consistently immune across all four error mechanisms and both task types. The largest losses tended to occur when the damaged feature carried more of the outcome signal. When the clean predictor was stronger, performance declined less because the model had another informative route.

That is more useful than a simple ranking of algorithms. It suggests the correct response depends on how error enters the workflow: laboratory imprecision differs from a miscoded category, and assigning neighbourhood averages to individuals differs from both. Practitioners need repeat measurements, validation subsets, sensitivity analysis or measurement-error models suited to the mechanism. Swapping a linear model for a neural network, or vice versa, may change robustness at the margin but does not make an error-prone feature accurate.[1]

What this changes for people who build and use models

Data teams should document how each consequential feature was obtained, what unit it represents, how often it is missing or recoded, and whether a value belongs to the individual or to a group. A model card that lists only predictive performance is incomplete when the input process can distort both decisions and explanations. Validation should include plausible error ranges and should examine whether performance and feature rankings change across sites, devices and subgroups.

For clinicians, teachers, workers, borrowers and other affected people, the practical issue is contestability. A decision should not become harder to challenge because a dashboard presents a precise feature-importance chart. Organisations need a route to correct an inaccurate source record and rerun the decision. The study does not quantify harm in any of these domains, but it shows why governance must reach upstream to measurement and record creation rather than beginning only after a model has been trained.[1]

Limits and what would change the assessment

The evidence is entirely simulated. The data contain three features, normal or Bernoulli distributions, specified additive structures and a deliberately simple non-predictive variable. Real datasets have correlated predictors, missingness, selection bias, repeated observations, nonlinear causal structures and error that may depend on the outcome or on protected characteristics. Ten replications per cell supported the paper's factorial analysis, but the reported post-hoc power estimate is based on observed effects and does not replace broader replication.

The study also evaluates permutation importance rather than SHAP, counterfactual explanations or causal attribution. Reproducible code was available on request rather than linked as a public repository in the article. Confidence would rise if the patterns were reproduced on realistic benchmark datasets with known measurement processes, under correlated and differential error, and across explanation methods. Prospective studies showing that targeted remeasurement or correction improves decisions would turn a strong simulation warning into direct operational evidence.[1]

What this means for people

  • People may receive confident but distorted decisions when an influential feature is recorded inaccurately.
  • A correction pathway should let an affected person challenge the source record as well as the model output.
  • Practitioners should treat feature-importance graphics as descriptions of measured data, not proof of causal importance.

Global context

Measurement error is a global infrastructure problem because models often move between hospitals, lenders, schools, factories and public services with different recording practices. The simulations do not establish the size of harm in any country, but they explain why a model can lose reliability when its input definitions or measurement devices change. Common reporting standards should therefore describe feature provenance and error, not just dataset size and headline accuracy.

What the evidence does not yet show

  • All data were simulated; no clinical, employment, lending or public-service decision was evaluated.
  • The three-feature designs are much simpler than operational datasets with correlation, missingness and selection effects.
  • Only non-differential feature error was studied, so error linked to the outcome or a protected group remains untested.
  • Permutation importance was the interpretation method; the results cannot automatically quantify distortion in every explanation technique.
  • Increasing sample size reduced simulation variability but the study did not test every possible error-correction strategy.

What to watch next

  • Replication with real datasets whose measurement process and repeatability are independently known.
  • Sensitivity analyses covering correlated, differential and subgroup-specific measurement error.
  • Comparisons of permutation importance, SHAP, counterfactual explanations and causal methods under the same errors.
  • Prospective evidence that remeasurement or correction improves consequential decisions and reduces unequal error.

Living evidence record

Impact record IAI-1MZCKKX

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent or research support

Present

Record status

Monitoring

Last checked

11 October 2026

Source trail

1 direct source across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Single-source reporting disclosure

This record analyses one direct source. It can establish what Scientific Reports published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.

Evidence trail

Sources used for this report

Links checked 11 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Science & Research

Did the arthritis AI travel between cohorts?

A peer-reviewed analysis of gut-microbiome data from 2,238 people reached a mean internal ROC-AUC of 0.834, but its genus-level model fell to 0.439 in 39 independently processed Shanghai samples. The study is a warning about cross-cohort transfer, not evidence for a rheumatoid-arthritis diagnostic test.

8 min · 1 source

Science & Research

Can a crash model help dispatchers if it already knows the outcome?

A Chicago-record model reported a 99.85% AUC, but its main 37-feature pipeline included seven injury-outcome fields. The paper describes post-crash classification—not prospective dispatch prediction—and does not report results for its reduced feature check.

7 min · 1 source

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.