Why do multi-agent AI systems keep failing?
An unreviewed analysis of 22,848 closed issues across 21 prominent open-source projects identified 944 genuine multi-agent problems. Handoffs, execution control and memory dominated—but repository reports cannot establish failure rates in production systems.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
Research topic
Which issues developers report in widely used open-source LLM multi-agent projects, what causes can be recovered from closed discussions, and which fixes practitioners document

At a glance
- 1The researchers collected 22,848 closed issues from 21 projects, used a local DeepSeek-R1 filter to flag 2,335 candidates, then manually retained 944 issues genuinely involving multi-agent behaviour.
- 2Orchestration and execution represented 208 of the 944 issues, or 22.0%; memory and state management followed at 14.7%. These are shares of the curated issue set, not system failure probabilities.
- 3Only 427 of the 944 issues had an identifiable solution in the discussion. Workflow optimisation was the largest solution category, accounting for 111 of those 427 solutions.
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Living evidence record
Impact record IAI-0UJJ215
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent support
Present
Record status
Monitoring
Last checked
3 October 2026
Source trail
3 direct sources across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 3 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
The question behind the issue tracker
Multi-agent AI products divide work among specialised model-driven components: one agent may plan, another call a tool, another retrieve information and another review the result. The arrangement promises broader capability, but it also creates more handoffs, shared state and stopping conditions. The researchers asked a practical software-engineering question rather than a benchmark question: what problems do developers actually report when building and using open-source LLM multi-agent systems, what appears to cause them, and what fixes are documented?
The evidence is a new, unreviewed preprint from researchers at Wuhan University, Central China Normal University and the University of Oulu. It mines repository discussions rather than running controlled model trials. That makes the study useful for finding recurring engineering patterns in real projects, but unsuitable for saying that a given framework fails a particular percentage of the time. A GitHub issue is a selected report, not a denominator of all successful and unsuccessful executions.[1][2][3]
How 22,848 issues became 944 cases
Projects had to describe multi-agent LLM functionality as a core component, have at least 3,000 stars, ten contributors and 100 closed issues, and show an update within the preceding year. The final 21 included AutoGen, CrewAI, MetaGPT, AgentScope, CAMEL, OpenAI Agents Python, Smolagents, Flowise, Ragflow and other frameworks, toolkits and workflow products. The team collected every closed issue available before 16 November 2025, producing 22,848 records. The popularity threshold intentionally focuses the sample on established public projects and excludes smaller or private systems.
A locally deployed DeepSeek-R1 model first classified the issue text and reduced the pool to 2,335 multi-agent candidates. The first author then manually reviewed those candidates; disputes were discussed with another author. The final dataset contained 944 issues judged to describe genuine multi-agent problems. In a 500-issue validation sample, two coders agreed strongly, with Cohen's kappa of 0.82; 102 of those 500 cases were retained as real multi-agent issues. The automated stage therefore helped triage the workload, but human interpretation still shaped the dataset.[2][3]
Handoffs, control flow and memory led the taxonomy
The authors used open coding and constant comparison to group the 944 cases into 17 issue categories. Orchestration and execution was largest at 208 cases, or 22.0%. It included failed handoffs, wrong execution order, loops that never ended, premature termination, resource contention and concurrency faults. Memory and state management was next at 14.7%, followed by feature requests at 13.2%, implementation issues at 10.6%, tool issues at 9.2% and integration issues at 6.7%. Security-labelled issues were only 0.3% of the curated set.
Those proportions describe what survived the study's selection pipeline. They do not show that security matters less than reliability, because maintainers may discuss vulnerabilities privately, users may not recognise them, and closed public issues favour problems that can be reproduced and resolved. Nor does the taxonomy separate failures caused by the foundation model from defects in schemas, libraries, prompts, state stores or orchestration code. Its value is architectural: extra agents add coordination surfaces where ordinary software defects and probabilistic model behaviour meet.[2]
What the reported causes and fixes show
The paper identified 541 causes across the issue discussions. Workflow problems accounted for 159, or 29.4%, followed by tool-integration problems and memory problems. Among orchestration and execution cases with an identified cause, 57.0% were attributed to workflow problems; among memory and state-management cases with an identified cause, 51.2% were attributed to memory problems. The examples include stale context being reused, handoff messages arriving too late, schemas disagreeing about whether an agent should send text or structured data, and tools not being registered where the next agent can reach them.
Solutions were visible for only 427 issues, 45.2% of the 944-case set. Workflow optimisation led with 111 solutions, or 26.0%, followed by configuration changes at 15.0%, tool-integration changes at 11.9%, infrastructure changes at 11.0% and improved error handling at 10.5%. These are practitioner-documented remedies, not experimentally compared treatments. A closed ticket can reflect a workaround, changed expectation or version update rather than a durable causal fix, and the researchers did not re-run every solution under a shared test harness.[2]
What teams can act on now
For software teams, the immediate implication is to treat agent orchestration as a distributed system, not as a longer prompt. Handoffs need explicit schemas, ownership and acknowledgement. Loops need budgets and deterministic termination conditions. Shared state needs versioning, isolation and observable updates. Tool calls need validation, timeouts, retries and clear permission boundaries. A workflow should preserve enough trace information to show which agent saw which state and why an execution stopped, without logging sensitive user content indiscriminately.
The issue set also argues against assuming that adding another agent automatically improves quality. Every added component can create a new interface, context copy and failure path. Teams should compare the multi-agent design against a simpler baseline and measure task completion, cost, latency, recoverability and human review burden. For consequential uses, a useful test is not merely whether agents produce an answer, but whether the system fails closed, exposes uncertainty and allows a person to reconstruct the chain without trusting the agents' own narrative of what happened.[2][3]
Limits, funding and what would change the assessment
The study includes only popular open-source GitHub projects and only closed issues created before 16 November 2025. Private deployments, unresolved tickets, support channels and smaller repositories are absent. The first author performed much of the coding and extraction, so consensus checks cannot remove all subjectivity. The model-assisted filter could miss quietly worded cases, while popularity and activity thresholds may overweight mature frameworks with active reporting cultures. The findings are descriptive and cannot compare vendors, predict an individual deployment or measure user harm.
The authors report partial support from China's National Natural Science Foundation under grants 92582203 and 62402348. The manuscript does not present a competing-interests statement. Confidence would rise with peer review, an independently reproduced coding sample, sensitivity analysis of the automated filter, exposure-based denominators such as executions or active installations, and controlled tests that compare proposed fixes. Longitudinal work should also ask whether orchestration, memory and tool failures decline across releases—or simply move into private operational logs.[1][2][3]
What this means for people
- Users can receive incomplete, repeated or inconsistent results when agent handoffs and shared state fail even if each model appears capable in isolation.
- Developers may face higher debugging and monitoring costs because the fault can sit between agents, tools, prompts and ordinary software dependencies.
- Organisations need human escalation, traceability and bounded execution before placing multi-agent workflows in consequential services.
Global context
The repositories and contributors are international, while the research team spans China and Finland. Public GitHub evidence travels widely but underrepresents private enterprise and public-sector deployments, lower-resource communities and projects hosted outside GitHub.
What the evidence does not yet show
- This is an unreviewed preprint submitted to a journal, not a peer-reviewed version of record.
- Closed GitHub issues are selected reports with no denominator of executions, users or installations, so category shares are not failure rates.
- The 21-project sample excludes smaller, private and less-active systems, and all issues were collected before 16 November 2025.
- Model-assisted filtering and researcher coding can introduce omissions and subjective classifications despite validation and consensus checks.
What to watch next
- Peer review and independent replication of the 944-case taxonomy and automated filtering pipeline.
- Controlled comparisons of explicit handoff schemas, termination rules, state isolation and tool validation.
- Operational studies with execution denominators, severity grading, user impact and longitudinal release data.
Evidence trail
Sources used for this report
Links checked 3 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Technology
Alibaba links chips, cloud, models and agents in a full-stack AI push
Alibaba Cloud used its Hangzhou conference to announce new Qwen models, proprietary processors, an agent-focused cloud and a mobile-agent platform, presenting them as one integrated technology stack.
4 min · 3 sources
Technology
Does self-hosting keep an AI coding agent away from sensitive source code?
IBM has made a customer-managed version of its Bob development agent generally available, including supported air-gapped deployments. Local inference can keep code inside an approved environment, but the announcement does not prove that every integration, model or generated change is secure.
6 min · 2 sources
Technology
What can OpenAI's always-on dots do—and when must they ask?
OpenAI's 29 September agent launch puts ongoing work and permission boundaries at the centre of its product pitch. The announcement establishes capabilities offered, not independently verified reliability.
4 min · 3 sources
Reader discussion
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.