Three field experiments find coding assistants raise completed developer tasks on average
A Management Science paper combines randomised trials involving 4,867 developers at Microsoft, Accenture and another large company and reports an average increase in completed tasks, with variation across experiments.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
Reads the full article in a natural voice. First play may take a moment to prepare.
Research topic
Longer studies should track defects, review time, skill development, collaboration and whether gains persist as tasks become more complex.
At a glance
- 1A Management Science paper combines randomised trials involving 4,867 developers at Microsoft, Accenture and another large company and reports an average increase in completed tasks, with variation across experiments.
- 2Field experiments are stronger than self-reported productivity claims, but completed tasks do not capture every dimension of code quality, maintenance, security or team learning.
- 3Longer studies should track defects, review time, skill development, collaboration and whether gains persist as tasks become more complex.
Living evidence record
Impact record IAI-12SBJQX
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent support
Present
Record status
Updated
Last checked
28 September 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what Management Science published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
What the source reports
A Management Science paper combines randomised trials involving 4,867 developers at Microsoft, Accenture and another large company and reports an average increase in completed tasks, with variation across experiments.[1]
Why it matters
Field experiments are stronger than self-reported productivity claims, but completed tasks do not capture every dimension of code quality, maintenance, security or team learning.[1]
Research question and evidence gap
Longer studies should track defects, review time, skill development, collaboration and whether gains persist as tasks become more complex. The firms are large and technologically mature; effects may differ for smaller organisations, other languages or less experienced teams.[1]
What the study can support
The evidence trail for this report begins with Management Science. The linked material is classified as Research paper, and the report keeps that provenance visible so readers can judge the claim at the correct level. The strongest conclusion directly supported by the record is this: A Management Science paper combines randomised trials involving 4,867 developers at Microsoft, Accenture and another large company and reports an average increase in completed tasks, with variation across experiments.
A research paper can expose methods, measurements and comparisons, but the label alone is not a guarantee that the result will replicate or transfer into routine use. The design, sample, baseline, uncertainty and real-world setting still determine how far the conclusion can travel. In this case, the practical significance is narrower and more useful than a general claim that AI is transforming the whole sector: Field experiments are stronger than self-reported productivity claims, but completed tasks do not capture every dimension of code quality, maintenance, security or team learning.[1]
Where the result may transfer
The human impact needs to be evaluated alongside technical capability. Developers may finish some work faster, while employers should avoid translating a noisy average into unrealistic individual quotas. That means tracking who receives a measurable benefit, who must change their work, what new oversight is required and whether a person has a realistic route to question or correct a harmful result.
The firms are large and technologically mature; effects may differ for smaller organisations, other languages or less experienced teams. Geography matters because infrastructure, language coverage, professional practice, regulation and public expectations can change the outcome. Evidence from one organisation or country is therefore a starting point for comparison, not a universal forecast.[1]
What replication needs to answer
The present boundary of the evidence is explicit: Results vary across the three experiments and do not cover all software work or long-term organisational effects. This does not make the development unimportant; it defines what cannot yet be claimed responsibly. Stronger confidence would require transparent methods, appropriate comparison groups or benchmarks, disclosed failures and results that other teams can examine.
The next test is equally concrete: Replication across company sizes and evidence on quality, security and junior-developer learning. The underlying research question is: Longer studies should track defects, review time, skill development, collaboration and whether gains persist as tasks become more complex. Until those points are answered, readers should treat the report as a verified account of the current evidence—not a prediction that every promised outcome will occur.[1]
What this means for people
- Developers may finish some work faster, while employers should avoid translating a noisy average into unrealistic individual quotas.
Global context
The firms are large and technologically mature; effects may differ for smaller organisations, other languages or less experienced teams.
What the evidence does not yet show
- Results vary across the three experiments and do not cover all software work or long-term organisational effects.
What to watch next
- Replication across company sizes and evidence on quality, security and junior-developer learning.
Evidence trail
Sources used for this report
Links checked 28 September 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Work & Skills
US payroll data show pressure concentrated in some entry-level AI-exposed jobs
Stanford Digital Economy Lab researchers report no economy-wide displacement but identify weaker outcomes for younger workers in selected occupations where generative AI can perform a larger share of tasks.
4 min · 1 source
Work & Skills
Hiring experiment finds AI skills can improve interview prospects
A study with 1,700 recruiters in the UK and US reports that AI skills increased interview invitations across office, software and design roles, sometimes offsetting age or education disadvantages.
4 min · 1 source
Work & Skills
OECD interviews show how early AI agents are being kept within human checkpoints
Practitioners in 25 organisations across 11 countries describe agents moving into real workflows, with autonomy limited around consequential actions. The interview study is a window into practice, not a global adoption rate.
4 min · 2 sources
Reader discussion
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.