When you deploy an agentic workflow, you have not automated a task. You have delegated authority.
Nobody signed for it. That is the problem.
Automation is a tool doing what you told it. Delegation is an actor deciding on your behalf, inside your systems, under your credentials, with your name on the outcome. Every legal system in the world already knows what to do with delegation: the principal answers for the agent. The industry chose that word itself, and it was the correct word.
What follows is the foundational framework. The realities that make governance unavoidable, the responsibilities those realities create, the actions required, and the evidence that proves you took them. It closes on the one function almost every AI governance program is missing, including some very good ones.
The realities
The dual attack surface
AI exposure is no longer employees pasting confidential data into consumer chatbots. That is one face. The second face is agentic: multi-agent workflows and background integrations holding unreviewed system credentials, executing code, and acting across APIs without a human in the loop.
The scale of the second face is not intuitive. Non-human identities — agents, service accounts, keys, tokens — now outnumber human identities in the enterprise by roughly 100 to 1. The authorization surface has left the org chart entirely, and most of those authorizations were never written down, never reviewed, and carry no revocation date.
That is delegation at a scale no organization has ever attempted with human delegates, granted without any of the ceremony that granting authority to a person would involve.
Observability and enforcement are not governance
Monitoring proves you watched. Firewalls and data-loss prevention prove you blocked. Neither proves you consciously governed.
This is the most common and most expensive category error in the field. An organization with excellent telemetry and no governance artifacts has a complete record of events and no record of decisions. When the question is whether reasonable care was exercised, a log of what happened does not answer a question about what was decided, by whom, and on what basis.
A log is what happened. A record is what was decided, by whom, and on what basis.
The asymmetry of access and competence
Deploying a generative model or an autonomous agent takes minutes. Building the institutional comprehension to supervise one — model drift, prompt injection, hallucination, toxic data combination — takes quarters, and in my experience across security and network leadership, most organizations have not started.
The gap widens rather than closes, and the reason is arithmetic. Capability deploys at software speed. Comprehension moves at organizational speed. Meanwhile the threat side compounds: the interval from CVE disclosure to weaponized exploit has fallen from roughly 2.3 years in 2018 to about 10 hours, and criminal breakout time now runs near 29 minutes. Any control that requires a meeting to invoke is already too slow.
Exposure is no longer voluntary
A fourth reality belongs beside the other three, and it undoes the most common executive assumption in the category.
Your organization does not choose whether it uses AI. AI is now embedded inside ordinary business software purchased years ago for other reasons. The summarization feature in the CRM. The drafting assistant in the document suite. The classification model inside the ticketing system. “We don’t use AI” is not true at any organization operating today, and the employee using it often does not know they are using it either.
This is why inventory by job title fails. Exposure, not role, is the axis.
The template defense has expired
Regulators and cyber-insurance underwriters no longer accept a boilerplate acceptable use policy as evidence of due care. They want demonstrable diligence backed by operational evidence.
The obligations are not pending. EU AI Act Article 4 (AI literacy) has been in force since February 2, 2025. Article 50 (transparency) since August 2, 2026. Colorado SB 26-189 takes effect January 1, 2027. High-risk obligations under Annex III follow on December 2, 2027, and Annex I embedded systems on August 2, 2028.
Whether the frontier labs are moving too fast or not fast enough is an interesting argument. It is also entirely frontier-side, and it changes nothing about what a deploying organization already owes.
What they make you responsible for
Delegation liability is fiduciary
Deploying an agentic workflow is the legal delegation of authority, not a software procurement. Executives and boards carry direct liability for actions taken by agents operating on their infrastructure, under credentials their organization issued.
Directors have understood the analogous duty for decades in other domains: a board that fails to establish any system of oversight over a known area of risk has failed at the duty itself, whatever the outcome. The legal principle is not new. What is new is that the delegate is not a person, has no judgment about whether an instruction is reasonable, and can act thousands of times before anyone notices.
Ignorance is not a defense
“We did not know our staff or our software agents were routing data through third-party models” is an admission of a standard-of-care failure. It is not a mitigating fact.
This is worth stating plainly, because the instinct runs the other way. In most operational failures, not knowing reads as bad luck. In a duty-of-care regime, not knowing is the thing you were supposed to prevent.
Observation must replace attestation
Employee surveys and vendor promises are attestation. Somebody told you.
Defensible accountability requires observation: empirical, log-based verification of which models, tools and endpoints are actually being touched. Reading a summary of what happened checks agreement. Only observation checks truth.
Hold that principle for a few minutes, because it has a limit, and the limit is the most important section in this framework.
Action one: literacy on the record
Technological leverage has always been gated behind demonstrated competence. Heavy machinery, aviation, commercial driving, firearms. The pattern is old and it is not controversial.
But note where the duty sits, because this is where most organizations get it exactly backwards. A driving licence and a firearms permit gate an individual operating in public, and the obligation belongs to the person. Article 4 does not work that way. The literacy duty sits with the deploying organization. A company cannot discharge it by telling employees to go and get trained on their own time. The company has to hold the record: who was assessed, at what depth, against which version of the material, and when.
Depth is matched to operational authority.
| Seat | What the seat must be able to do |
|---|---|
| Board and C-suite | Systemic oversight, delegation risk, disclosure obligations, algorithmic liability |
| CISO and security leaders | Agentic threat taxonomy, model supply chain, deployment gates, identity isolation for non-human actors |
| CTO, CIO and IT | Tooling supply chain, agent identity, evaluation and deployment gates |
| Legal and risk | Cross-regulatory mapping, vendor retention clauses, defensible evidence standards |
| General workforce | Shadow-AI recognition, data sanitization, the non-negotiable boundaries of human accountability |
Every seat needs the same five capacities: recognition, boundary, verification, escalation, accountability. Only the depth differs. That is precisely why a single org-wide module does not satisfy Article 4, and why literacy you cannot evidence is, to a regulator or an underwriter, indistinguishable from no literacy at all.
Action two: a continuous lifecycle
Ad-hoc workshops and annual surveys are replaced by a programmatic arc.
| Function | What it does |
|---|---|
| Discover | Continuously audit perimeter, proxy and DNS traffic against known AI endpoint registries to surface human and agentic usage |
| Assess | Evaluate observed tools against risk tiers, toxic combinations, and regulatory frameworks (HIPAA, SOC 2, NIST AI RMF, ISO/IEC 42001) |
| Establish | Enforce clause-anchored acceptable use policies and unambiguous decision rights |
| Prove | Produce dated, third-party-verifiable audit artifacts and board reports a regulator or underwriter can independently confirm |
| Sustain | Re-scan on a monthly cadence to catch model drift, new integrations and emerging requirements before an audit does |
Governance is a continuing program, never a point-in-time audit. A single dated report proves you looked once. The systems move underneath the people: a model version updates, a vendor enables an AI feature inside software you already owned, an MCP server is registered to an agent that had a narrower remit last quarter. Nobody changed their behaviour, and the boundary moved anyway.
The missing function: Disclose
Here is the gap, and it sits inside Discover.
Network telemetry finds what crosses the wire. It is very good at that, and it should be the backbone of discovery. But it cannot find:
- The employee using a personal account on a personal device on a personal network.
- The contractor whose traffic never touches your infrastructure.
- The analyst who does not know the summarization feature she has used for six years now calls a model.
- The agent authorized last quarter for a narrower purpose, still inside policy on paper, now doing something nobody would approve if asked today.
Shadow AI is a disclosure problem before it is a detection problem.
An organization that disciplines the first employee who admits pasting customer data into a chatbot has purchased silence, and will govern a fiction from that day forward.
Aviation solved this exact problem fifty years ago, and the solution is documented, public and directly transferable.
In 1976 NASA established the Aviation Safety Reporting System. NASA sits between the reporter and the regulator as a neutral third party. Reports are de-identified before analysis — “all information that might assist in or establish the ID of persons filing ASRS reports... will be deleted.” The FAA commits that it “will not use any reports submitted to NASA under the ASRS... in any enforcement action.”
The conditions are the part that makes it work. Protection applies where the violation was inadvertent and not deliberate, where no accident occurred, where no criminal offence is involved, and where the report was filed promptly. It is not amnesty. It is a bounded, principled trade: tell us what happened, and we will treat it as information rather than evidence.
The result is the largest voluntary safety dataset in aviation, assembled almost entirely from events no monitoring system would ever have recorded, because they harmed nobody and only the crew knew.
Disclose belongs in the lifecycle as a peer of Discover, not an appendix to it.
| Function | Source of truth | What it cannot see |
|---|---|---|
| Discover | Network, proxy, DNS, endpoint telemetry | Anything that never crosses your wire |
| Disclose | A protected internal reporting channel | Anything nobody is willing to say |
Build the channel on the NASA pattern. Route it away from the employee’s own manager. De-identify before analysis. Protect the inadvertent explicitly and in writing. Exclude the deliberate just as explicitly. Publish back what you learn, because a channel that consumes disclosures and returns nothing stops receiving them.
Note what this does to the attestation principle above. Observation beats attestation for everything that crosses the wire. For everything that does not, a protected disclosure is the only evidence that will ever exist. The two are not in tension. They cover different halves of the same surface.
The acceptable use architecture
A modern AI acceptable use policy governs programmatic data ingestion, algorithmic liability and autonomous delegation. It is not a warning about chatbots.
| # | Clause | What it establishes |
|---|---|---|
| 1 | Scope and entity definition | Applies to personnel, contractors, vendors, automated scripts and autonomous agents alike. Establishes that agentic tools act under delegated corporate authority and that the deploying manager retains accountability for the agent's actions |
| 2 | Data classification and ingestion boundaries | Boundaries mapped to enterprise tiers — Public, Internal, Confidential, Regulated. Prohibits proprietary code, customer records or trade secrets entering any tool that trains on inputs or retains session data |
| 3 | Tool classification and approved registry | Three operational tiers: Sanctioned (enterprise-licensed, contractual zero data retention), Sandbox Only (isolated, synthetic data), Prohibited (public consumer models). Unsanctioned endpoints, browser extensions and unmonitored API calls are out of policy |
| 4 | Human-in-the-loop and authority caps | Documented human review for high-stakes workflows: personnel decisions, credit determinations, production commits, legal interpretation. Hard caps forbidding agents from executing financial transactions, modifying security perimeters, or binding the company without step-up authentication |
| 5 | Work-product accountability | All model output is unverified drafting. Full legal, technical and factual responsibility sits with the human who publishes or deploys it, including for hallucination, infringement and defamation |
| 6 | Synthetic identity and attribution | Prohibits unauthorized voice cloning, non-consensual biometric simulation and synthetic personas. Mandates unprompted disclosure when an external party is interacting with an agent, which is also Article 50's requirement, now in force |
| 7 | Observability, logging and enforcement | States plainly that proxies and endpoint monitors log prompt metadata, token volume and endpoint access for audit. Defines consequences, from credential revocation through to termination and contractual liability |
Clause 1 is where the delegation principle becomes enforceable rather than rhetorical. Until a document says in writing that an agent acts under delegated corporate authority and names the manager who retains accountability for it, the authority was granted by nobody and is owned by nobody.
Clause 4 is the one that fails most often in practice, and it fails quietly.
An authority cap that exists in the policy document but not in the agent’s permission scope is a sentence, not a control.
Decision rights
Accountability that is shared is accountability that is absent. Every operational row below has exactly one Accountable owner.
| Governance domain | Board / Exec | Legal & Risk | CISO / InfoSec | AI Gov / CAIO | IT Ops | BU Lead / Deployer |
|---|---|---|---|---|---|---|
| AUP drafting and updates | I | A | C | R | C | C |
| Tool vetting and vendor retention terms | I | C | A | C | R | C |
| Data ingestion boundaries and classification | I | A | R | C | C | I |
| Agentic access control and non-human IAM | I | I | A | C | R | C |
| Human-in-the-loop verification | I | I | I | C | I | A / R |
| Work-product accuracy and IP liability | I | C | I | I | I | A / R |
| Shadow AI network discovery and telemetry | I | I | A | C | R | I |
| Protected disclosure channel and intake | I | C | C | A | I | I |
| Regulatory filing, board audits, evidence defence | I | C | C | A | I | C |
| AI incident response and model deprovisioning | I | C | A | C | R | I |
| Role-based qualification and literacy | I | C | C | A | I | R |
A Accountable, single final decision-maker · R Responsible, executes · C Consulted, two-way · I Informed, one-way
Three operating principles govern the matrix.
One accountable owner per row. This matrix is the instrument that makes delegation legible. An agent’s authority ultimately traces to a named human in one of these columns, and if it does not trace to exactly one, it traces to nobody. Diffusing accountability across Security, IT and Legal guarantees that nothing is enforced when a breach or a hallucination occurs.
Separate infrastructure from intent. IT Operations provisions endpoints and proxies; InfoSec sets the perimeter rules. Business unit leaders own the operational and factual output of their agents outright, which forecloses blaming the model or the vendor for a flawed decision.
Non-human identity is identity. Agents interacting with enterprise APIs are governed under the same IAM discipline as human privilege, owned by InfoSec. The disclosure row is deliberately owned by AI Governance rather than Security, because a reporting channel owned by the enforcement function is a reporting channel nobody uses.
Vendor due diligence
Data rights and training exclusions
- Contractual zero data retention. In writing: prompts, context windows, retrieval embeddings and outputs are never stored, logged, or used to tune shared or foundational models.
- Ephemeral session scoping. Server-side cache and session state purged on response termination, with no secondary replication into telemetry or QA stores.
- Tenant isolation. Dedicated encryption keys (CMEK) and partition isolation enforced cryptographically or logically at the compute layer.
- Telemetry scrubbing. Where operational metrics are logged, prompts and raw payloads are stripped programmatically before entering observability pipelines.
Security architecture
- Prompt injection defences, direct and indirect: hardened input validation, system-prompt isolation, architectural boundaries against payload execution.
- DLP and PII redaction upstream of the inference endpoint.
- Model supply chain disclosure: upstream base models, licensing chain, training data provenance, third-party API dependencies.
- Encryption: TLS 1.3 in transit, AES-256 at rest, including vector indexes and inference buffers.
Agentic execution boundaries
- Least privilege tied to human-approved service accounts.
- Sandboxed execution in ephemeral, network-isolated containers with no route to corporate subnets.
- Deterministic circuit breakers on recursion depth, token consumption per task, and concurrent outbound calls.
- Step-up authorization, structurally enforced, before an agent finalizes a financial commitment, commits production code, or alters IAM privilege.
Compliance and auditability
- Third-party attestations: current SOC 2 Type II with AI processing in scope, ISO/IEC 42001, ISO 27001.
- IP indemnification, unconditional, against third-party infringement claims arising from generated output.
- Audit logging and reproducibility: model version, temperature, applied system prompts and tool invocations, retained at least 365 days, with integrity evidenced by content fingerprint rather than asserted.
- Regulatory classification: where the vendor’s system sits under the EU AI Act — prohibited practice, high-risk under Annex I or III, subject to Article 50 transparency, or minimal risk — plus conformity documentation. General-purpose AI obligations sit in their own chapter and are a separate question from risk tier; ask both.
Instant disqualification
| Red flag | Exposure |
|---|---|
| Clickwrap terms reserving rights to improve models on customer data | Irrevocable loss of IP and trade secrets |
| No option for customer-managed encryption keys | Exposure to vendor-side credential compromise |
| Opaque third-party model routing without notice | Uncontrolled supply chain risk and regulatory non-compliance |
| No granular role-based access control at workspace level | Lateral privilege escalation across departments |
| No documented incident response for hallucination or leak | Indefensible liability under corporate due-care standards |
What you owe, and where to start
Six things, and none of them are optional once AI is in the building, which it is, whether anyone chose it or not.
- Assessed literacy, by seat, on the record. Depth matched to authority, scored the same way every time, issued as a dated credential.
- A boundary artifact. One page of what the agents may not do, with a named signer and a date it last changed. Not a training module. An artifact.
- Decision rights with one accountable owner per row. Written down, and the same document a regulator would be shown.
- A continuous cadence. Because the systems move underneath the people.
- A protected disclosure channel. Open to the inadvertent, closed to the deliberate, routed away from the reporting line.
- Artifacts someone outside your company can verify. Dated, third-party-verifiable, and produced before anyone asks for them.
That is a programme, not an afternoon. So start with the smallest piece of it that produces something real.
This week, do three things. Write the one-page list of what your agents are not allowed to do, and put a name and a date on it. Ask your CISO what percentage of actual AI exposure the current discovery tooling can see, and listen carefully to the part of the answer about personal devices. Then find out whether an employee in your organization could report an AI mistake without it going through their own manager.
None of the three requires a budget. All three produce a record where there was none.
An auditor, an underwriter, or a board will not ask whether you took AI seriously. They will ask where the record is. Ignorance of the obligation has never been a defence against it, and the obligation is already in force.
Delegation without a signature is the exposure. Put a name on it.
Twelve questions, no account required, and it produces the first page of the record.
- ASRS Immunity Policies — NASA Aviation Safety Reporting System: de-identification, enforcement waiver, exclusions
- ASRS Program Briefing — NASA
- The Evolution of Crew Resource Management Training in Commercial Aviation — Helmreich, Merritt & Wilhelm, FAA
- EU AI Act — Article 4 (AI literacy, in force February 2, 2025), Article 50 (transparency, in force August 2, 2026), Annex III high-risk (December 2, 2027), Annex I embedded (August 2, 2028)
- Colorado SB 26-189 — effective January 1, 2027
- SanctumShield: Agent Governance — the two faces of shadow AI, four risk surfaces, the ~100:1 non-human identity ratio (Astrix Security, via the CIS Controls v8.1 MCP Companion Guide, 2026)
- SanctumShield: Why Now — regulatory timeline, exposure vectors, and threat velocity figures (CrowdStrike 2026 Global Threat Report)
As of September 17, 2026.
