Does human oversight improve perceived AI-audit outcomes?
A survey of 150 ESG-audit professionals linked automation and human oversight with higher self-rated efficiency, accountability and transparency. It measured perceptions—not audit speed, accuracy or detected greenwashing—and cannot establish cause.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1The study surveyed 150 auditors, compliance officers, sustainability managers and AI-governance analysts recruited by purposive and snowball methods.
- 2Automation and AI decision-making were positively associated with self-rated audit outcomes, while human oversight partially mediated every tested relationship.
- 3The study did not examine audit files, elapsed time, misstatements, greenwashing detection, regulatory findings or investor outcomes, so it cannot establish real-world effectiveness.
Research topic
Cross-sectional survey of perceived efficiency, accountability and transparency in AI-assisted ESG auditing, with human oversight modelled as a mediator
The answer: oversight tracked better perceptions, not verified audit performance
Professionals who reported more AI automation, AI-supported decision-making and human oversight also rated their ESG-audit processes as more efficient, accountable and transparent. Human oversight statistically explained part of the relationship between the two AI measures and all three reported outcomes. The pattern supports treating oversight as part of an AI-audit operating model rather than as a final sign-off added after automation.
It does not show that oversight caused better audits. Every central variable came from the same respondent in a one-time questionnaire, and the outcomes were perceptions. The study did not measure hours saved, errors found, emissions claims verified, greenwashing detected, restatements, enforcement outcomes or investor decisions. Its useful contribution is a testable governance hypothesis, not proof that an AI-assisted ESG audit works better.[1]
The Impact Brief · Free
Follow the evidence in finance & business.
Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.
Who answered the survey and what they were asked
The researchers collected 150 online responses over one month from ESG auditors, compliance officers, sustainability managers and AI-governance analysts in private, public and non-governmental organisations. Recruitment used professional networks, audit forums and LinkedIn groups focused on ESG. Participants were chosen purposively and then through snowball referrals because experience with AI-driven ESG auditing is specialised.
That method can reach knowledgeable respondents, but it does not produce a representative sample. People active in professional online groups or referred by peers may be more engaged with AI and governance than the wider profession. The paper's authors are based in Greece, but respondent geography is not reported clearly enough to treat the sample as nationally representative or globally balanced.
Respondents used five-point agreement scales to rate AI automation, AI decision-making, human oversight, audit efficiency, accountability and transparency. The analysis used partial least-squares structural-equation modelling. The achieved sample exceeded a simple ten-times rule but was slightly below the paper's 155-to-160 target from an inverse-square-root calculation for detecting a small effect around 0.20.[1]
What the structural model found
Reported AI automation was positively associated with perceived efficiency, with a standardised path coefficient of 0.351, accountability at 0.288 and transparency at 0.341; all had p-values below 0.001. AI decision-making had smaller positive direct associations with efficiency at 0.203, accountability at 0.248 and transparency at 0.297. The first two had p-values of 0.001 and the third below 0.001.
Both AI variables were also associated with human oversight: 0.494 for automation and 0.440 for decision-making. Oversight in turn was associated with efficiency at 0.426, accountability at 0.436 and transparency at 0.397, all with p-values below 0.001. Significant indirect paths remained alongside direct paths, so the authors described partial mediation.
The indirect association between automation and efficiency through oversight was 0.210; corresponding figures were 0.216 for accountability and 0.196 for transparency. For AI decision-making, they were 0.187, 0.192 and 0.175. These coefficients describe questionnaire constructs in this sample. They are not percentages of improvement and should not become claims that oversight reduced errors or raised productivity by a particular amount.[1]
Why perception and performance must stay separate
A respondent enthusiastic about an organisation's AI programme may rate automation, oversight and outcomes positively at the same time. Organisational culture, senior sponsorship or personal attitude could raise every score. The authors ran a Harman single-factor check for common-method bias, but a statistical diagnostic cannot remove the basic same-source, same-time design problem.
Efficiency is especially easy to overread. The survey captured agreement with statements about process, not timestamps, staff hours, cost per engagement or time to assurance. Accountability and transparency were also reported constructs, not independently reviewed evidence of traceable decisions, successful challenges, disclosed limitations or comprehensible models.
The paper distinguishes process transparency from technical explainability. A well-documented workflow may still rely on an opaque model; an interpretable model can sit inside a process with poor responsibility and escalation. Organisations need both operational records and appropriate model evidence rather than using a favourable survey score as proof of either.[1]
What firms, auditors and regulators can take from it
The defensible implication is to specify oversight before automating an ESG-audit task. A firm can name the person accountable for each model-assisted judgment, define which evidence must be checked, preserve source-to-output traceability and document when a human changed or rejected a recommendation. High-risk assertions, such as emissions baselines or materiality decisions, need stronger review than routine document sorting.
Audit committees should ask for operational measures alongside staff sentiment: elapsed time, rework, exception rates, confirmed misstatements, false alarms, unresolved data gaps and the proportion of model outputs receiving substantive review. Regulators and standard-setters need evidence that oversight is capable of changing a decision, not merely that a human clicked approve.
For workers, meaningful oversight requires time, access to underlying evidence and authority to challenge the system. If workload targets make review nominal, a human-in-the-loop label adds little protection. Training should cover both ESG subject matter and the limits of the AI system, while escalation routes should protect staff who identify unreliable evidence or misleading claims.[1]
Evidence limits, ethics and disclosures
The cross-sectional design prevents causal inference and cannot establish the sequence implied by the mediation model. Purposive and snowball recruitment means coefficients describe this sample and its networks rather than population parameters. The marginally smaller-than-target sample, self-report measurement and unclear geographic distribution add uncertainty. The model-fit SRMR was 0.063, while NFI of 0.795 was below commonly preferred levels.
The authors acknowledged that measuring predictors and outcomes at one time may inflate relationships. They recommended longitudinal work, sector comparisons and explainable-AI variables. Data are available from the corresponding author on request because of confidentiality commitments rather than through an open repository.
The study received no external funding, and the authors declared no conflicts. The university ethics committee waived formal review under its rules because the project had no funding, collected no identifying or sensitive personal information and involved no vulnerable populations. The authors reported obtaining informed consent.[1]
What would change the assessment
Confidence would rise with a preregistered study following audit teams over time and combining surveys with system logs, reviewed files and independently adjudicated outcomes. A stronger comparison could examine similar engagements before and after a defined oversight intervention, or compare teams assigned different protocols while controlling for task difficulty and organisational support.
Useful denominators include total engagements, model-assisted decisions, human overrides, confirmed errors, unresolved exceptions and staff hours. Reporting should separate organisations, sectors and jurisdictions and test whether relationships hold outside networks already interested in AI. Evidence would weaken the interpretation if objective outcomes fail to improve, if oversight becomes ceremonial, or if automation shifts work into hidden review and correction.[1]
What this means for people
- Audit staff need protected time and authority for review; a nominal approval step is not meaningful oversight.
- Companies should measure rework, exceptions and confirmed errors rather than infer effectiveness from perceptions.
- Investors and the public still need independent assurance that sustainability claims are accurate.
Global context
The author team is based in Greece, but the paper does not clearly report respondents' geographic distribution. ESG rules, assurance standards, professional roles and AI governance vary by jurisdiction. The associations should not be treated as globally representative until replicated with transparent country and sector denominators.
What the evidence does not yet show
- Cross-sectional questionnaire data cannot establish that AI or oversight caused better outcomes.
- All central variables were self-reported by the same respondents, creating common-method and perception-bias risks.
- Purposive and snowball recruitment produced a non-representative sample; respondent geography was unclear.
- The sample was slightly below one of the paper's targets, and NFI was below commonly preferred levels.
- No objective audit speed, accuracy, misstatement, greenwashing, enforcement or investor outcome was measured.
What to watch next
- Longitudinal or controlled comparisons of defined human-oversight protocols.
- Objective audit-file, system-log, error, override and staff-time measures.
- Replication across sectors, jurisdictions and representative professional samples.
- Evidence that reviewers have time, information and authority to change model-assisted decisions.
Living evidence record
Impact record IAI-0AUNZCL
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
10 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what Journal of Risk and Financial Management published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 10 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Finance & Business
What would SAP gain from TechWolf?
SAP has signed an agreement to acquire TechWolf, whose software maps tasks, skills and labour-market signals for workforce planning. The price is undisclosed, closing remains conditional and no public evidence yet shows that deeper integration will improve hiring, reskilling or job design.
7 min · 2 sources
Finance & Business
Is AI financing becoming circular?
A Bank for International Settlements analysis maps 1,246 AI firms and 972 investment relationships, finding that commercial links accompanied 16.1% of AI-to-AI deals by count but 46.4% by disclosed value. The pattern is consequential, yet it does not show that the transactions were improper, unprofitable or certain to spread financial stress.
10 min · 5 sources
Finance & Business
Who is financing AI's power backbone?
General Catalyst, Koch Equity Development and co-investors have agreed to put $2 billion of preferred equity into Axiom at a $37.5 billion enterprise value. The capital is expensive, the separation is unfinished and the disclosed terms do not prove AI demand will deliver the promised returns.
8 min · 2 sources
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.