Who is responsible when AI agents do the research?
Human authors remain responsible, a new peer-reviewed review argues. AI research agents should be treated as powerful tools—not authors or scientists—and their use should trigger explicit oversight, disclosure and post-publication accountability.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1The review argues that current AI research agents can cause and complete work but cannot accept moral responsibility, so they should not be listed as authors or described as scientists.
- 2Its proposed minimum is a three-part human duty: oversee and manage the agent, disclose its precise contribution and technical identity, and remain answerable for the work after publication.
- 3This is a peer-reviewed ethical and policy argument, not an experiment; it does not measure which controls reduce errors, misconduct or irreproducibility in real laboratories.
Research topic
How human authorship duties should change when partially autonomous AI agents make substantial contributions to scientific research

The direct answer: the human author remains accountable
A scientist cannot hand authorship responsibility to an AI agent simply because the system performed much of the work. That is the central conclusion of a peer-reviewed review by David Resnik, Mohammad Hosseini and Robert Pennock. The authors distinguish an agent's ability to cause things to happen from a person's capacity to understand duties, consent to publication and answer for misconduct. On their account, present systems may have causal agency, but they are not moral agents.
The practical consequence is deliberately demanding: people named as authors must stay substantially involved. They should supervise important decisions, reveal what the agent did and remain available to defend, correct or retract the work. If no human has contributed enough to qualify for authorship, the paper argues that the answer is more human scrutiny—not naming the software as an author.[1]
What kind of evidence this is—and is not
The article is a peer-reviewed review in Frontiers in Genetics' ethics, legal and social issues section. It develops a normative argument from established authorship principles, recent literature on AI-assisted science and anticipated agent capabilities. It does not recruit participants, compare laboratories, audit a dataset or test whether one governance regime produces fewer errors than another. There is therefore no sample size, effect estimate or experimental comparator to report.
That distinction is crucial. Peer review shows that the argument passed journal scrutiny; it does not turn the recommendations into measured proof. The review is best read as a proposed responsibility framework for institutions, journals and research teams. Its strength is conceptual clarity about who must answer when something goes wrong. Its main evidential limit is the absence of field evidence showing which controls work in practice.[1]
The authors separate causal agency from moral agency
AI research agents can plan workflows, search literature, formulate hypotheses, run code, analyze data, direct other tools and draft reports. Some systems can also connect to robots or laboratory equipment. Those capabilities make the agent causally important: its actions may directly shape a result. But the review argues that causal importance is not enough for authorship, because authors must also understand obligations, consent to publication and be capable of blame or correction.
The authors acknowledge that philosophers disagree about the threshold for moral agency and leave open the possibility that future systems could qualify. Their claim is about current systems. They say there is insufficient evidence that today's agents possess the understanding, motivation, consciousness or moral emotions needed to participate as responsible members of a scientific community. Calling them 'scientists' now would therefore blur a distinction on which accountability depends.[1]
Responsibility one: active oversight throughout the work
The first duty is to manage the agent as a fallible research assistant, not a sealed substitute for judgement. The review calls for human engagement at important stages including question formation, protocol design, data collection, analysis, interpretation and reporting. Discipline-specific controls would also cover data access, validation, equipment audits and workflow logs. The point is not to watch every token; it is to retain informed control over decisions that affect validity or safety.
That creates a real trade-off. Extensive review reduces some of the productivity gain that makes autonomous agents attractive. The authors accept the cost and argue that efficiency cannot outrank rigor, reproducibility, confidentiality or public safety. Their proposed standard also implies that a principal investigator needs enough time and expertise to challenge an agent's work. Merely approving a polished output at the end would not satisfy meaningful oversight.[1]
Responsibility two: disclose exactly what the agent changed
The second duty is specific disclosure. A paper should say which substantial tasks the agent performed—such as designing an experiment, selecting data, executing analyses, interpreting results or drafting text—and identify the system by name, model and version where possible. The review suggests putting core information in the methods and using an appendix or supplement for detail. A vague statement that AI was used would not tell readers what evidence pathway needs closer inspection.
For reproducibility, this proposal points beyond a one-line acknowledgement. Useful records may include prompts or goals, tool permissions, data sources, model updates, checkpoints, human interventions and outputs that were rejected. Some information cannot be shared because of privacy, security or intellectual-property constraints, but the reason for omission should itself be visible. Disclosure is not a transfer of blame; it is the audit trail that helps others evaluate the human authors' choices.[1]
Responsibility three: answer for the work after publication
Accountability continues after a paper appears. The review says human authors must answer questions, share methods and code where appropriate, respond to public or media concerns, investigate suspected errors and correct or retract the record when necessary. If an agent invents citations, copies text, fabricates data or produces misleading images, the people who deployed and approved it cannot treat autonomy as an excuse.
The authors recognize that this is severe. A supervisor is not automatically blamed for every independent wrongdoing by a student or employee. But an AI agent cannot be sanctioned, explain its intent or restore trust. Leaving nobody responsible would create a gap precisely where the public consequences of unreliable science can be greatest. The review calls for clearer rules distinguishing reckless deployment from ordinary negligence because existing misconduct procedures were not written for autonomous software.[1]
What journals and institutions could implement now
The framework can be translated into concrete editorial and laboratory requirements. Journals could require a structured agent-use statement tied to contribution roles; ethics and biosafety reviews could ask what tools can access or execute; institutions could retain versioned logs and name a responsible human for each consequential workflow. Independent replication or expert audit becomes especially important when an agent generated hypotheses and also selected the evidence used to support them.
The review also floats persistent identifiers for research agents, similar in tracking purpose—not moral status—to serial numbers on laboratory equipment. That could help researchers study recurring error patterns across versions. Yet model identity changes through updates, fine-tuning and connected tools, so any identifier would need to describe a configuration and time window rather than imply a stable, person-like entity.[1]
Limits, funding and what would change the assessment
The review focuses on responsibility and authorship, not every legal question raised by autonomous research. It is grounded largely in US research norms and examples, while authorship rules, misconduct procedures and legal liability vary internationally. It also does not quantify the cost of compliance or determine when an AI contribution becomes 'substantial.' Those thresholds will need discipline-specific and cross-border work.
The authors report support from US National Institutes of Health programmes and a National Science Foundation grant, no commercial or financial conflicts, and no generative AI use in writing the manuscript. Confidence in the framework would increase with prospective studies in real laboratories: compare oversight designs, measure undisclosed agent use, test reproducibility and track corrections or safety incidents. Evidence that agents can reliably understand duties, consent and accept meaningful accountability would challenge the authors' present conclusion about authorship.[1]
What this means for people
- Researchers retain personal and professional responsibility even when an AI agent performs substantial work.
- Readers, patients and the public gain a clearer route to challenge or correct AI-assisted findings when human accountability is explicit.
- Early-career research roles may change as routine work is automated, increasing the importance of preserving human training and critical judgement.
Global context
AI research agents are being developed and used across borders, while authorship, privacy, biosafety and misconduct rules remain nationally and institutionally fragmented. The review comes from US-based authors and draws heavily on US research-governance concepts. Its three-part framework is broadly portable, but implementation needs local law, discipline-specific risk controls and standards that remain auditable when models, data and connected tools are supplied internationally.
What the evidence does not yet show
- This is a peer-reviewed normative review, not an empirical study; it supplies no experimental sample, effect size or measured comparison of governance approaches.
- The recommended oversight, disclosure and accountability duties have not yet been prospectively validated across real laboratories.
- The argument applies to current AI systems and explicitly leaves open whether future systems could meet a different threshold for moral agency.
- Key operational terms, including substantial contribution and sufficient human involvement, require discipline-specific definition.
- Legal liability, misconduct standards and authorship policies vary across countries and institutions.
What to watch next
- Structured journal disclosure standards that record model, version, permissions, tasks and human checkpoints.
- Prospective audits comparing agent-assisted projects with different levels of human oversight.
- Institutional rules separating ordinary error, negligence and reckless failure to review agent output.
- Persistent configuration identifiers that allow recurring agent errors to be tracked without treating software as an author.
- International convergence—or divergence—on who bears legal and research-integrity responsibility.
Living evidence record
Impact record IAI-1EQ826T
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
8 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what Frontiers in Genetics published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 8 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Science & Research
Did AI just produce 372 families of new mathematics?
OpenAI has released 722 manuscripts grouped into 372 result families after posing about 4,000 problems to an internal model. The scale is extraordinary, but the repository says some unformalised results may contain errors and most have not yet received independent human review.
7 min · 3 sources
Science & Research
Can AI reveal hidden disease patterns in tissue?
A peer-reviewed method used ensembles of neural networks and statistical testing to find disease-associated spatial structures across rheumatoid arthritis, ulcerative colitis and dementia datasets. It is a research-discovery tool, not a diagnostic test or evidence of improved patient care.
8 min · 2 sources
Science & Research
Can an LLM spot what a clinical-trial paper leaves out?
A peer-reviewed benchmark found that a retrieval-plus-LLM pipeline reached F1 0.822 across 119 reporting questions on 40 held-out trial articles. Better evidence retrieval lifted F1 to 0.893, showing both the promise and the main bottleneck.
9 min · 2 sources
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.