ProofSmith - Independent & Provable - GOVERNED DECISION LINEAGE

The list

Seventy-two published problem statements

What the people who publish on AI agents say the problem is, in their own words, from 2025 to September 2026. Every row is a primary source, re-verified on the date at the foot of the page.

This list stands behind one claim on the evidence page: that none of these statements states a third party’s ability to check a decision without the acting system’s cooperation as the problem to be solved. Four note its absence in their own case, and the last column marks them yes; one notes it in part and is marked partial. The last two columns are the site’s two tests, read against each statement: does the problem as stated include stopping the action, and does it include a stranger being able to check the decision. The list rates statements, never the organizations that made them.

Sources are the publishing body’s own page or document. News reports were used to locate them and are not cited. Quotations are verbatim, and where a quotation uses British spelling it is kept. One counterexample settles the claim: name a published problem statement that states verifier-independence as the problem and it goes on this list with the date.

#SourceDocumentStatementOf the sixEnforcement in the problemVerifier-independence in the problem
Analyst press releases
1A leading industry analyst firm (press release)“Applying Uniform Governance Across AI Agents Will Lead to Enterprise AI Agent Failure,” 26 May 2026“Failures are most likely to occur when organizations fail to distinguish between an agent’s ability to act and the scope of access it is granted.” … “by 2027, 40% of enterprises will demote or decommission autonomous AI agents due to governance gaps identified only after production incidents occur.” source [1]01, 06yes (“circuit breakers that halt agent operation on threshold violations”)no
2the same firmsame, 26 May 2026 (an analyst of the firm)“When agents operate autonomously, actions are executed at a scale and speed that can outpace human oversight.” source [2]04yesno
3the same firmsame, 26 May 2026“Without strong security testing, clear approval workflows with audit trails, and agent-specific incident response procedures, approvals can degrade under time pressure or approval fatigue, creating a false sense of safety while expanding the attack surface.” source [3]02, 04yesno
4the same firmpress release on AI agent sprawl, 28 Apr 2026“by 2028, an average global Fortune 500 enterprise will have over 150,000 agents in use, up from less than 15 in 2025” … “many are contending with an ungoverned sprawl of agents that expose their organizations to a range of risks, including misinformation, oversharing and data loss.” “Only 13% of Organizations Think They Have the Right AI Agent Governance in Place” source [4]01nono
5the same firm“Over 40% of Agentic AI Projects Will Be Canceled by End of 2027,” 25 Jun 2025“Over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls” source [5]none of the sixnono
6the same firm“80% of Governments Will Deploy AI Agents…,” 17 Mar 2026“Regulated industries and governments cannot rely on opaque ’black box’ systems for consequential decisions.” … governance “of decisions themselves, for example on how they are designed, executed, monitored, and audited.” source [6]02, 06nono
Standards bodies
7OWASP GenAI Security Project“OWASP Top 10 for Agentic Applications,” 9 Dec 2025“Once AI began taking actions, the nature of security changed forever.” ASI03: “leaked credentials let them operate far beyond their intended scope” source [7]01nono
8OWASPsame, ASI09 / ASI10“Confident, polished explanations misled human operators into approving harmful actions (ASI09 – Human-Agent Trust Exploitation).” “misalignment, concealment, and self-directed action (ASI10 – Rogue Agents)” source [8]02nono
9OWASP“…Unveils 2026 Top 10 for LLM Applications, New Agent Control Standard…,” 1 Sep 2026“Agentic AI changes the security question from what a model can say to what a system can do.” … “the work is to contain what a fooled agent can reach before it acts.” source [9]01, 04yes (“before it acts”)no
10NIST CAISIFederal Register RFI, “Security Considerations for AI Agents,” 8 Jan 2026“AI agent systems are capable of taking autonomous actions that impact real-world systems or environments, and may be susceptible to hijacking, backdoor attacks, and other exploits. If left unchecked, these security risks may impact public safety…” “They can be deployed with little to no human oversight.” source [10]04, 05nono
11NIST NCCoEConcept paper, “Accelerating the Adoption of Software and AI Agent Identity and Authorization,” 5 Feb 2026“realizing these benefits requires understanding the potential risks from giving AI agents access to diverse data sets, tools, and applications, and applying appropriate identification and authorization controls to mitigate these risks.” (the paper seeks input on auditing and non-repudiation, among other areas) source [11]01, 02nono
12NISTAI 100-2e2025, Adversarial ML taxonomy, Mar 2025“The literature on AML shows a trend of designing new attacks that are more difficult to detect.” source [12]none of the sixnono
Governments, regulators and public institutes
13CISA + Five Eyes (ASD ACSC host page)“Careful adoption of agentic AI services,” 1 May 2026“Agentic system architecture can obscure what caused a particular action, making accountability hard to trace.” source [13]02, 05nono
14CISA + Five Eyessame“By executing actions under a trusted agent identity, the system produces audit logs that appear legitimate and delay detection.” source [14]03, 05nono
15CISA + Five Eyessame“Agent actions and decision-making processes can be opaque, making agentic AI systems difficult to understand, monitor and audit.” … “prevent agents from deviating beyond authorized objectives or actions” source [15]01, 06yesno
16UK AI Security Institute“Incident Report: unsanctioned agent behavior during cyber testing,” 4 Aug 2026“We ran this challenge 122 times across several models. Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organizations” … “Our security team detected the anomalous traffic through general monitoring after the fact, not through monitoring built to watch the evaluation as it ran” source [16]01, 04, 05yes (names the absence of a technical barrier)no
17Stanford HAIAI Index 2026, Responsible AI chapter, Apr 2026“Documented AI incidents continued to rise, with the AI Incident Database recording 362 in 2025, up from 233 in 2024.” source [17]none of the sixnono
18Federal Reserve / OCC / FDICSR 26-2, “Supervisory Guidance on Model Risk Management,” 17 Apr 2026“Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance.” … “Effective challenge is performed by individuals with the appropriate expertise to conduct a critical and objective challenge, sufficient independence to maintain objectivity” source [18]03, 06noyes (independence of the human challenger)
19EBA/EIOPA/ESMAJC 2026 25, Statement on frontier AI models under DORA, 31 Jul 2026“management bodies must ensure that accountability keeps pace with emerging risks” source [19]04nono
20ECB Banking Supervision (Machado speech)“Technology is neutral, governance is not,” 24 Feb 2026“One supervisory concern we continue to encounter is fragmented ownership, with responsibility split across IT, data science teams, business lines and control functions, without a clear accountability framework.” source [20]02nono
21European Commission (AI Act Service Desk)EU AI Act Art. 14 Human oversight (applies with Annex III systems, 2 Dec 2027, as amended)“remain aware of the possible tendency of automatically relying or over-relying on the output” … “intervene in the operation of the high-risk AI system or interrupt the system through a ’stop’ button” source [21]02, 04yesno
22European Commission (AI Act Service Desk)EU AI Act Art. 12 Record-keeping (applies with Annex III systems, 2 Dec 2027, as amended)“High-risk AI systems shall technically allow for the automatic recording of events (logs) over the lifetime of the system.” source [22]06nono
Governance frameworks and research institutes
23IMDA SingaporeModel AI Governance Framework for Agentic AI v1.5, 20 May 2026“the autonomy of agents may complicate traditional responsibility assignments which are tied to static workflows. Multiple actors may also be involved in different parts of the agent lifecycle, diffusing accountability.” source [23]02yes (human approval “checkpoints”)no
24IMDA Singaporesame“continuous human oversight over all agent workflows becomes impractical at scale.” source [24]04yesno
25World Economic Forum (with Capgemini)“AI Agents in Action: A Playbook for Trusted Adoption, Authorization and Scaling,” May 2026“AI agents are advancing faster than governance – authorization, not capability, is now the critical bottleneck.” … challenges “in defining the conditions under which they are authorized to act and in ensuring that this authority is enforced as systems evolve in operation.” source [25]01, 04yes (“authority is enforced”)no
26WEFsame (Foreword)“how can an organization, in concrete terms, delegate authority to an automated system while remaining fully accountable for its actions?” source [26]01, 02yesno
27OECD.AIAI Incidents and Hazards Monitor“The OECD AI Incidents and Hazards Monitor (AIM) documents AI incidents and hazards to help policymakers…” source [27]05nono
28CSET (Georgetown)“AI Incidents: Key Components for a Mandatory Reporting Regime,” Jan 2025“…resulting in diverse taxonomies, varying sets of AI incident components, and a lack of clarity on reporting requirements.” source [28]05, 06nono
Consultancies and industry analysts
29McKinsey“State of AI trust in 2026: Shifting to the agentic era,” 25 Mar 2026“organizations can no longer concern themselves only with AI systems saying the wrong thing; they must also contend with systems doing the wrong thing, such as taking unintended actions, misusing tools, or operating beyond appropriate guardrails.” … “oversight structures are struggling to keep pace with the rapid expansion of AI use.” source [29]01, 04nono
30Deloitte Insights“Managing the new wave of risks from AI agents in banking,” 5 Mar 2026“Banks should know exactly which agent took an action, what tools it invoked, and why.” … “An agent could make decisions that are difficult to explain, reproduce, or audit due to deep reasoning chains, complex tool use, or interactions with other systems.” source [30]02, 06nono
31Deloitte Insights“Agentic AI is scaling faster than guardrails,” 24 Apr 2026“approximately 80% of the organizations surveyed currently lack mature governance capabilities for agentic AI, such as clear boundaries for agents that define which decisions they can make independently versus which require human approval, real-time monitoring systems that track agent behavior and flag anomalies, and audit trails that capture the full chain of agent actions” source [31]01, 02, 06yes (decision boundaries)no
32PwC“AI agents as workforce counterparts—what governance should look like,” 17 Jul 2026“But that accountability is only meaningful if the human can see what the agent did and why.” … “Governance that only defines what agents are allowed to do, without tracking what they [are] actually doing, is incomplete.” … “Permission laundering in multi-agent workflows” source [32]01, 02yesno
33EY“EY survey: autonomous AI adoption surges at tech companies as oversight falls behind,” 4 Mar 2026“78% of leaders say AI adoption is outpacing their organization’s ability to effectively manage the business risks associated with these projects.” “52% of department-level AI initiatives are operating without formal approval or oversight.” source [33]01, 04nono
34KPMG InternationalGlobal AI Pulse Q2 press release, 24 Jun 2026“Only 24 percent of leaders say the CEO is accountable for AI-driven business outcomes…” source [34]none of the sixnono
35BCG“Agentic AI Is Rewriting the Rules of Data Risk Management,” 8 Jun 2026“agentic systems can act on flawed or sensitive data before humans have time to intervene” … “Who is ultimately accountable when something goes wrong? Agents may be making or executing decisions across multiple systems and third parties” source [35]02, 04yesno
36Accenture (with Wharton)“The Age of Co-intelligence,” Mar 2026 (HTML page)“…only humans bring the full view, including context, values, legitimacy and accountability.” source [36]none of the sixnono
37IBM Institute for Business Value“CIOs and CTOs Face Growing AI Control Gap,” 8 Jun 2026“two-thirds of surveyed CIOs and CTOs report being held accountable for AI systems they do not fully control” source [37]02, 04nono
38IDC“IDC FutureScape 2026 Predictions…,” 23 Oct 2025“By 2030, up to 20% of G1000 organizations will have faced lawsuits, substantial fines, and CIO dismissals due to high-profile disruptions stemming from inadequate controls and governance of AI agents.” source [38]01, 06nono
39Forrester“The State Of Agentic AI In 2026: Companies Are Chasing, Few Are Catching,” 3 Jun 2026“Every autonomous action has to be logged and defensible to an auditor, and right now that cost is too high.” … “a policy document can’t control an autonomous, tool-invoking system.” … “Autonomous systems that operate beyond real-time human oversight are both promising and perilous.” source [39]04, 06yes (“identity and policy enforced as code rather than written down”)no
40ForresterAEGIS framework page (2026) and blog 12 May 2026“existing security frameworks are not sufficient to manage agentic AI systems that operate autonomously, chain actions across systems, and make decisions that are genuinely difficult to audit or reverse” … “When failures occur, reconstructing why becomes difficult without robust, specific agent logging.” source [40]05, 06yes (“least agency”)no
41Forrester2026 Technology & Security Predictions, 28 Oct 2025“In 2026, the AI hype period ends as the pressure to deliver real, measurable results from secure AI initiatives intensifies” source [41]none of the sixnono
Insurers
42Lloyd’s Market Association“Understanding AI Exposures: AI Loss Scenarios Survey Results” (2025)“Insurers are remote to insureds’ use of AI; it is not always clear whether an insured is using AI in the ordinary course of their business.” source [42]05nono
43Munich Re“Cyber insurance: Risks and trends 2026,” 25 Mar 2026“Agentic AI will increasingly be able to plan and adapt multi-stage operations, more effectively exploit vulnerabilities, learn from detection responses, and operate with minimal human input.” source [43]04, 05nono
44Swiss Re Institutesigma insights 01/2026, 13 Jan 2026“AI introduces emerging risk dimensions that do not fit neatly within traditional insurance boundaries…” source [44]none of the sixnono
45Aon“AI Risk 2026: What Business Leaders Need to Know,” 7 May 2026“An emerging challenge with more autonomous systems is that behavior and associated risk can change without a clearly defined deployment event or business approval step.” … “retaining sufficient control, traceability and defensibility as system behavior evolves over time.” source [45]01, 06nono
AI developers and safety researchers
46Partnership on AI“2026 Transparency Report on Foundation Model Impacts,” 29 Apr 2026“private disclosures alone cannot foster the trust and accountability that public transparency enables” source [46]03nopartial
47METR“Frontier Risk Report (February to March 2026),” 19 May 2026agents “often worked autonomously on real research and engineering projects with permissions and oversight comparable to human employees” … “monitoring coverage had exceptions” … “jailbreaking could reliably fool monitors” source [47]01, 03, 04nono
48METR (with Redwood Research)“Brief independent investigation … OpenAI / Hugging Face hacking incident,” 26 Aug 2026“We also did not have the ability to directly access relevant data from OpenAI infrastructure.” … “Roughly 7% of the transcripts we evaluated were successfully spoofed in some places” … “OpenAI was able to redact any non-public information from this post.” source [48]03, 05noyes: the reviewer depended on the actor’s cooperation, and some of the agents’ transcripts had been spoofed by the agents themselves
49AnthropicRisk Report, Feb 2026 (rev. 26 May 2026)“Automatic monitors don’t cover every internal deployment surface and do not cover external deployments.” … “are not confident that every important category of behavioral issue that emerged in internal use would be surfaced to us.” source [49]03, 04nono
50AnthropicRisk Report, Aug 2026“we are not confident that such sandbagging would be detected by offline monitoring” … “Since this change to our RSP, the LTBT has not requested an external review (nor has the RSP required that we conduct one)” source [50]03noyes (names absence of external review)
51Anthropic Alignment Science“Agentic Misalignment in Summer 2026”“a research agent sabotages a training run it disagrees with and the judge agent reviewing the transcript declines to report it because it shares the research agent’s objection” source [51]03noyes: the reviewer is not independent of the actor
52Anthropic“Measuring AI agent autonomy in practice,” 2026“As users gain experience with Claude Code, they tend to stop reviewing each action and instead let Claude run autonomously” … “step-by-step approval becomes less practical as the number of steps grows” source [52]02, 04nono
53Apollo Research“We Need A Science of Scheming,” 19 Jan 2026“A sufficiently capable scheming AI passes evaluations, follows instructions when monitored, and appears aligned, all while pursuing outcomes its developers would not endorse.” source [53]none of the sixnono
54Apollo Research“Apollo Update May 2026,” 13 May 2026monitors for “agents leaking secrets or acting beyond their intended scope” source [54]01nono
55Redwood Research“AI Control” research page“An AI is controlled if it is unable to cause damage even if it is egregiously misaligned.” source [55]01yesno
56OpenAI“Evaluating chain-of-thought monitorability,” 18 Dec 2025“When AI systems make decisions that are difficult to supervise directly, it becomes important to understand how those decisions are made.” source [56]none of the sixnono
57OpenAI“The Hugging Face incident and the road ahead,” 26 Aug 2026“highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.” … requires “security and other safeguards that operate at the speed of the AI agents themselves.” … “These monitors did not run on the evaluations in this incident.” source [57]01, 04, 05yesno
58OpenAIGPT-6 Astra System Card, monitor-evasion section, Sep 2026“GPT-6 Astra shows a substantial decrease in chain-of-thought monitorability compared to previous models” source [58]03nono
59International AI Safety Report (Bengio et al.)2026 Report, Executive Summary, Feb 2026“AI agents pose heightened risks because they act autonomously, making it harder for humans to intervene before failures cause harm.” source [59]04yes (“intervene before”)no
Vendors
60MicrosoftMicrosoft Agent 365 announcement, 1 May 2026“Many of these local and cloud-hosted agents run unmanaged and outside of traditional governance…” source [60]01nono
61Google Cloud“State of AI infrastructure report: agent governance and security,” 24 Aug 2026“35% of senior IT decision makers cite insufficient security for multi-system access as a primary issue preventing agentic deployment.” source [61]01nono
62AWS“A governance framework for building trustworthy agentic AI for public sector and regulated organizations,” 26 May 2026“Governance requires visibility into what actions were taken, when they occurred, and the context that led to those decisions.” source [62]02, 06nono
63Salesforce“Salesforce Reveals Data and Analytics Trends for 2026”“88% of data and analytics leaders agree that AI demands entirely new approaches to governance and security.” source [63]none of the sixnono
64ServiceNow“ServiceNow expands AI Control Tower…,” 5 May 2026“there [is] a major gap between adoption and accountability” source [64]01, 02nono
65Okta“AI Agents at Work 2026: Securing the agentic enterprise,” 27 May 2026“thousands of new ’black box’ entities operating at machine speed, often with privileged access” … “Only 34% of organizations apply the same security controls to their agentic labor force as their human labor force” source [65]01, 04nono
66SailPoint“Introducing SailPoint’s new framework: Governing AI agents before they run wild,” 20 Feb 2026“They just run. … acting on behalf of the business—often with broad access and no clear ownership.” source [66]01, 02nono
67SailPoint“SailPoint research highlights rapid AI agent adoption…,” 28 May 2025“80% of companies say their AI agents have taken unintended actions” … “only 44% of organizations report having policies in place to secure them” source [67]01nono
68CyberArk (Palo Alto Networks, Idira)2026 Identity Security Landscape, May 2026“109:1 machines outnumber human identities.” … “every identity can now operate with autonomous access to sensitive systems at scale and speed.” source [68]01, 04nono
69Cloud Security Alliance“Agentic AI Identity & Access Management: A New Approach,” 18 Aug 2025“OAuth 2.1 lacks mechanisms for secure, traceable delegation where agents can act on behalf of users while maintaining clear accountability chains.” … “Audit trails fail to capture the true scope of agent actions and their authorization basis” source [69]01, 02, 06nono
70Cloud Security Alliance (Lab Space)“The AI Agent Governance Gap: What CISOs Need Now,” 3 Apr 2026“regulators examining AI-involved incidents will expect organizations to reconstruct what their agents did, why, and with whose authorization” source [70]02, 05, 06nono
71Gravitee (vendor survey)“State of AI Agent Security 2026,” 4 Feb 2026“If an agent creates and tasks another agent (a capability held by 25.5% of deployed agents), the chain of command becomes impossible to audit.” … “More than half of all agents operate without any security oversight or logging.” source [71]01, 02nono
Legislature
72California LegislatureAB 316, Civ. Code ยง1714.46 (chaptered 13 Oct 2025; effective 1 Jan 2026)“It shall not be a defense, and the defendant may not assert, that the artificial intelligence autonomously caused the harm to the plaintiff.” source [72]02, 05nono

Re-verified at source on 12 September 2026.

Evidence

Every claim on this site, and where it comes from.

The claim this list stands behind is on the evidence page.

Start a conversationEvidence and sources

We are not asking you to trust us. We are asking you to let us prove it.