Back to the news portal
Work & SkillsResearch paperResearchSource analysisNorth AmericaUnited States

Does an AI label bias code review—or does job rank matter more?

In a randomized vignette experiment, 447 Microsoft engineers did not penalize an AI-use disclosure on identical code, but rated the same code and author more favourably when the fictional author was labelled Principal rather than Level 1. One AI-normalized US workplace cannot settle how other teams respond.

By The Impact of AI Work & Skills DeskReleased 3 October 2026 at 05:00 BST9 min read2 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesWorkplace AISoftware engineeringCode reviewStatus biasDisclosureHuman judgement

Research topic

How AI-use disclosure and fictional author seniority affect engineers’ evaluation of identical code

The Impact of AI research cover asking whether an AI label or job rank biases code review, with identical conceptual code panels balanced beneath different metadata labels.
AI-generated editorial illustration. The code panels, metadata labels and balance are conceptual; they do not depict the study interface, a real employee, company system or measured chart.

At a glance

  • 1Four hundred and forty-seven engineers completed four code reviews each in a randomized 2×2 within-subjects experiment, creating 1,788 reviews of code held constant across AI-disclosure and author-seniority labels.
  • 2The study detected no AI-disclosure effect on the measured outcomes in this highly AI-accustomed sample, but a Principal label improved ratings of code effectiveness and author competence relative to a Level 1 label.
  • 3The unregistered vignette study came from one large US technology company; it did not observe live review decisions, promotions, pay, hiring or actual work by junior engineers.

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Living evidence record

Impact record IAI-168YDGX

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Present

Record status

Monitoring

Last checked

3 October 2026

Source trail

2 direct sources across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

The experiment separated the code from its labels

Code-review systems do more than show code. They can expose names, profile images, job levels and increasingly a statement that AI helped create a change. Those cues may shape judgement even when they say nothing about whether the code works. The new study isolated two of them: an AI-use disclosure and the fictional author’s seniority.

Researchers recruited 447 full-time software engineers at one large US technology company. Each participant reviewed four short snippets in a simulated interface, producing 1,788 reviews. For every participant, one junior and one principal author were labelled as having used AI and one at each level was labelled as not having used it. Authors, order and snippets were shuffled, while the code itself was independent of the label assignment.

That within-subjects 2×2 design is the study’s central strength. The comparison does not ask whether AI-generated code differs from human-written code; it asks whether reviewers judge the same material differently after reading metadata about how it was written or who supposedly wrote it. Four fictional, gender-ambiguous authors were used: two Level 1 engineers and two Principal engineers.[1][2]

The sample was already comfortable with AI

This was an AI-normalized workplace, not a general survey. Almost all participants, 97.8%, reported using AI to generate code at least occasionally, and 81.2% said they used AI for code review. Average comfort with disclosing AI use was close to the top of a seven-point scale. That context is not a footnote: it is the most plausible boundary around the null result.

The 447 participants spanned job levels: 10.1% were Level 1, 34.9% Level 2, 31.1% Senior and 23.9% Principal or above. Three-quarters were recorded as male, 18.6% as female and 5.8% had missing or prefer-not-to-say responses. Review was familiar work: 27.3% said they reviewed code multiple times a day and 23.9% daily.

Participants saw an AI sentence in a commit message rather than a brightly highlighted badge. The sentence said either that AI assisted code generation or that no AI was used. Profile layout stayed the same. This makes the manipulation reasonably close to a lightweight disclosure, but not necessarily to the structured co-author trailers or prominent warnings that production tools may show today.[2]

Reviewers scored code, competence and effort

The four snippets covered two programming languages. Three contained seeded bugs—one had four and two had two—while a fourth AI-generated tic-tac-toe snippet contained none. Reviewers rated quality, readability and security; those items formed a code-effectiveness composite. They also rated coding capability and whether the author seemed competent, skilful and masterful, forming an author-competence composite.

A third composite covered impressions of irresponsibility, laziness and giving up easily. Researchers separately analysed a hypothetical hiring recommendation and the number of inline comments. Comments were originally optional, but partway through collection at least one became mandatory because some reviewers left none even when bugs were present. The models controlled for snippet and this protocol change and included a random intercept for participant.

The study was not preregistered. The authors therefore treated the three composites as primary outcomes and item-level and robustness analyses as secondary. That transparency helps readers avoid treating every reported p-value as equally confirmatory, but the absence of preregistration still leaves more room for analytic decisions made after seeing the data.[2]

No AI-label penalty was detected in this setting

The AI disclosure had no detectable effect on any measured outcome. For the author-competence composite, the coefficient was −0.034 standardised units, with a 95% confidence interval from −0.130 to +0.061 and p=0.481. The models also found no measurable interaction between AI disclosure and seniority.

A null result does not prove that AI stigma no longer exists. The confidence interval rules out some effects but not every small one, and the company’s unusual level of AI adoption limits transfer to teams where use is newer, controversial or restricted. The disclosure was not highlighted, and there was no participant-level question asking whether each reviewer noticed it. Failure to notice would dilute an effect.

There is nevertheless evidence that the manipulation registered for at least some people. Spontaneous mentions of AI appeared in feedback on 28 of 238 AI-labelled reviews sampled for that check versus seven of 236 reviews without the label. The authors also report more than 99% power to detect an AI penalty as large as one in a prior study. The appropriate conclusion is therefore narrow: this experiment did not detect the earlier-sized penalty in this AI-accustomed population.[2]

The same code received a junior penalty

Seniority changed judgements even though the evaluated code did not. A Principal rather than Level 1 label increased the author-competence composite by 0.178 standardised units, p<0.001, and code effectiveness by 0.148, p=0.001. Code attributed to junior authors was rated lower for quality by 0.184 units, readability by 0.142 and security by 0.123.

The fictional junior authors were also seen as less masterful by 0.388 units, less skilful by 0.241 and more likely to give up easily by 0.170. The wider laziness composite pointed in the same direction but did not meet the authors’ significance threshold, p=0.075. Seniority did not measurably change hypothetical hiring recommendations or comment counts.

These are shifts in ratings, not evidence that the organisation actually underpays, rejects or promotes people because of a label. A field estimate in the paper suggests the vignette effect might explain about a third of a separately observed difference in pushback on junior-authored code, but the authors explicitly warn that cross-study scaling makes that an approximation, not a precise causal estimate.[2]

What the experiment cannot establish

Vignettes remove many features of real collaboration: ongoing relationships, ownership of a service, urgency, accountability, negotiation and repeated exposure to the same author. Those dynamics could amplify status bias or counteract it. The manipulation also compared the extremes of Level 1 and Principal; intermediate ranks may produce smaller or different effects.

The fictional profiles held gender ambiguous, so the design cannot test how seniority, gender, race or other protected characteristics interact. It should not be used to infer that a junior penalty has the same origin, magnitude or legal meaning as a gender penalty. Nor did it test AI systems that autonomously authored a change; AI was framed as a tool that assisted generation.

The manuscript lists authors from Microsoft, the Institute for Life at Work, the University of Washington and an independent researcher. It includes acknowledgements but no separate funding or competing-interest declaration was found in the reviewed manuscript. Several authors are Microsoft employees, and the participants came from the same large technology-company setting, facts that readers should weigh when judging independence and generalisability.[2]

The practical question is which metadata reviewers need

For engineering teams, the result argues against assuming that disclosure alone is the main fairness problem. If a job-level label changes perceived quality on identical code, interfaces may be importing hierarchy into a task that could be judged first on the change itself. Teams could test an initial metadata-minimised review, structured checklists or delayed identity disclosure for suitable repositories, while preserving ownership and escalation information where safety requires it.

Removing labels is not automatically fair or practical. Reviewers may need expertise, component ownership or accountability cues; anonymous review can also make collaboration harder. The evidence supports auditing interface choices, not a universal mandate. Organisations should compare review outcomes by level, inspect who receives subjective pushback and distinguish bug-finding from broad impressions of competence.

Confidence would rise with preregistered, multi-company field experiments using live pull requests, intermediate job levels and clearly noticed disclosure formats. Important outcomes include defects found, changes requested, review time, approval, rework, promotion signals and whether effects differ by gender, race, geography and AI familiarity. Until then, the paper shows a controlled junior-status penalty in one workplace—not a settled law of code review.[2]

What this means for people

  • Junior engineers may receive less favourable subjective evaluations even when the code is identical, making review process design a workplace-equity issue.
  • Reviewers could benefit from structured criteria that focus attention on observable defects and maintainability before status cues.
  • Organisations need local evidence: a null AI-label effect in one AI-normalized company does not guarantee that disclosure is neutral elsewhere.

Global context

The participants worked at one large US technology company, and nearly all already used AI for coding. Software cultures, job ladders, disclosure norms, labour protections and review practices differ across countries and organisations. Replication in small firms, public-sector teams, outsourced development, open-source projects and regions with lower AI adoption is necessary before applying the result globally.

What the evidence does not yet show

  • The within-subjects vignette experiment came from one large US technology company with unusually high AI use and disclosure comfort.
  • The study was not preregistered, and the comment requirement changed during data collection, although the model included that change as a covariate.
  • AI disclosure appeared in a short commit message and was not highlighted; there was no participant-level manipulation check showing that every reviewer noticed it.
  • The design compared Level 1 with Principal and did not test intermediate levels, live workplace consequences or interactions with gender and other protected characteristics.
  • Ratings, hypothetical hiring recommendations and comment counts do not establish effects on defects, pay, promotion, hiring or actual engineering performance.

What to watch next

  • Preregistered replications across companies, countries and teams with different levels of AI adoption.
  • Field experiments on live pull requests that measure defects found, approvals, rework, review time and career consequences.
  • Tests of prominent versus embedded AI disclosure and structured co-author metadata with confirmed reviewer attention.
  • Analyses of interactions among seniority, gender, race, geography, team familiarity and reviewer experience.

Evidence trail

Sources used for this report

Links checked 3 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Work & Skills

Should AI wait until people think first?

In an unreviewed study of 398 adults across five countries, a writing assistant withheld direct generation until users stated a position and supporting argument. The design shifted effort earlier and improved one conditional error-classification measure, but did not significantly improve overall judgment accuracy.

8 min · 2 sources

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.