Can METR’s AI time horizons be read literally?
Not as a simple ruler, this statistical reanalysis argues. Across 228 software tasks and 26 AI systems, more flexible models predicted held-out task families better than METR’s shared logistic curve and suggested that equal multipliers in human task time do not represent equal jumps in difficulty. The overall capability trend may still be real.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1The authors reanalysed 228 software tasks, 26 AI systems and 79 task families from METR Time Horizon 1.1.
- 2A shared monotone spline and explanatory item-response model beat the baseline across all displayed family-held-out scoring metrics.
- 3The audit does not reject broad capability progress; it challenges literal comparisons between equal multiplicative increases in reported horizon.
Research topic
Whether METR's AI task time horizon is estimated and interpreted consistently across software tasks and model generations

The direct answer: informative, but not a literal stopwatch
The preprint argues that METR's time-horizon graph should not be read as though human task duration were a uniform ruler of AI capability. A model's horizon is derived from the human completion time of tasks on which it is estimated to succeed at a chosen probability. That creates an appealing interpretation: a longer horizon sounds like an ability to carry out proportionally longer work. The audit finds that the mapping from human time to model difficulty is not that simple.
Using the same public task results, the authors fit more flexible models and predicted held-out task families better. Their curve is nearly flat across roughly two to 30 minutes of human time before becoming steeper. Moving from three to 30 minutes is therefore not equivalent to moving from 30 minutes to five hours, even though both are tenfold increases. The result challenges comparisons between multipliers; it does not show that the overall rise in model performance is fictitious.[1][2]
The Impact Brief · Free
Follow the evidence in science & research.
Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.
What data were reanalysed
The analysis uses public METR Time Horizon 1.1 data: 228 software-engineering tasks, 26 AI systems and 79 task families. A family groups related items, which matters because variants of the same problem may share difficulty features. Model success is observed over repeated attempts, while human completion time supplies the explanatory scale intended to make tasks comparable.
Human times are not exact constants. The paper says they were typically derived from about four attempts per task. For 29% of tasks, a valid human measurement was unavailable and an expert estimate was used. That adds uncertainty to the horizontal axis that a clean graph can hide. The meaning of an hour also depends on who performed the task, under what conditions and how estimation replaced observation.[1]
The baseline and two alternatives
The baseline is a shared-slope logistic model: each AI receives a position on a common curve linking log human time to success probability. The audit compares it with a shared monotone spline, which lets the curve bend while preserving the expectation that longer tasks are generally harder, and an explanatory item-response model. The latter gives tasks latent difficulty, models family structure and uses beta-binomial overdispersion.
If human time is a stable one-dimensional measure of difficulty, a simple shared curve should predict new families well. If family identity and a nonlinear relationship matter, a horizon can still summarise performance, but its numerical scale is a constructed benchmark quantity rather than a direct estimate of how long an autonomous system can work in any environment.[1]
How predictive performance was compared
The authors used five-fold cross-validation at task-family level. That is stronger than randomly holding out individual variants because it asks whether a model generalises to related groups it did not see during fitting. They evaluated 12 combinations of scoring rule and weighting: marginal log score, Brier score and smoothed elementary scores at 0.5 and 0.8 success thresholds, each under three weighting schemes.
Both proposed models beat the baseline on every displayed scoring combination. Consistency across metrics is more informative than one favourable number, although all comparisons still come from the same dataset and modelling choices. The spline and item-response approaches are not identical explanations; their shared advantage indicates that the simple curve leaves predictable structure unused.[1]
The alternative still recovers an overall ability trend
The item-response model yields an ability parameter for each system. After rescaling, those values have an R-squared of 0.996 with METR's estimates, according to the paper. The authors interpret this as evidence that the published plot may track a Rasch-like latent ability well even if the time label is not equally spaced in substantive difficulty. Ranking and broad progress can be robust while literal duration comparisons are not.
This prevents two overreactions. The audit does not warrant saying time-horizon research measures nothing. Nor does a strong correlation between two fitted summaries prove that either transfers to unrestricted work. Both are constructed from the same task suite, whose coverage and measurement choices bound what the latent ability means.[1]
Practical meaning, limits and what would change the assessment
A statement that horizons doubled or increased tenfold is incomplete unless readers know where on the fitted curve the change occurred, which task families support it and how uncertain the estimate is. Procurement and safety decisions should not use small differences as hard operational thresholds. If a decision depends on whether a system can handle work lasting a specific number of hours, it needs direct evaluation on representative tasks at that duration and in that environment.
The paper is an unreviewed statistical audit of software tasks up to about 30 human hours. It assumes the sampled tasks are useful probes and does not decide whether human time is the best measure for every kind of work. Confidence would rise with preregistered comparisons on new families, more measured human baselines, explicit modelling of time error, longer tasks and external tests beyond software.[1][2]
What this means for people
- A benchmark-hour estimate is not promised unattended work time.
- Researchers can preserve the progress trend while clarifying its nonlinear scale.
- Safety decisions need representative local tasks.
Global context
METR's measure is cited globally, but the audited items are software exercises rather than a geographic sample of work. Labour practices, tooling, language and organisational constraints change both human duration and successful completion. The statistical caution travels globally even when the numeric horizon does not.
What the evidence does not yet show
- Unreviewed preprint analysing one public software-task dataset.
- Human times typically use about four attempts; 29% of tasks use expert estimates.
- The suite covers software work and durations up to about 30 hours.
- Comparisons depend on sampled families and statistical specifications.
- Agreement between fitted ability estimates does not prove transfer.
What to watch next
- Family-held-out replication on newer suites.
- More human measurements and explicit time uncertainty.
- Longer tasks and nonsoftware domains.
- Horizon reporting with diagnostics and intervals.
- Links between benchmark horizons and organisational performance.
Living evidence record
Impact record IAI-1Q32N8U
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
9 October 2026
Source trail
2 direct sources across 2 source types.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
Evidence trail
Sources used for this report
Links checked 9 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Science & Research
Did the arthritis AI travel between cohorts?
A peer-reviewed analysis of gut-microbiome data from 2,238 people reached a mean internal ROC-AUC of 0.834, but its genus-level model fell to 0.439 in 39 independently processed Shanghai samples. The study is a warning about cross-cohort transfer, not evidence for a rheumatoid-arthritis diagnostic test.
8 min · 1 source
Science & Research
Can Tangermeme reveal what genomic AI has learned?
A peer-reviewed Nature Methods toolkit standardises prediction, perturbation, attribution and sequence-design operations around genomic deep-learning models. It can turn black-box outputs into testable hypotheses, but a model explanation is still not proof of a biological mechanism.
7 min · 3 sources
Science & Research
Can AI turn knee MRI into a personalised repair scaffold?
A physics-constrained generative pipeline produced printable digital scaffold designs from knee MRI-derived conditioning fields in 43 seconds per case. The study tested simulated geometry and mechanics—not manufactured implants, living tissue or patient outcomes.
6 min · 1 source
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.