§ Perspective · Founder POV · September 14, 2026

Who actually
answered your prompt?

The attacks in Anthropic’s September 2026 threat report are ordinary. The economics are not — and the chapter drawing the least attention is the one most companies should read first.

By Lindsay Hiebert · Founder · CISSP

The question

A question for the CISOs reading this first. Can you name every router, proxy, reseller and coding harness sitting between an employee’s prompt and whatever actually answers it?

Anthropic’s fourth threat intelligence report landed on September 10. One hundred fifty-four pages, seven harm areas, activity disrupted between December 2025 and August 2026. I read all of it. The part that should change what a mid-market security program does next is not in the chapters getting the coverage.

The report shows the threat, the path shows the risk. What the report shows: agent swarms, breaches in hours, relayed sessions. What most organizations miss: the model path, governance, and AI literacy. The governance question — who actually answered the prompt, and can you show the paper on every hop between the employee prompt and the answering provider?
The report shows the attacks. The missed story is the path.

The attacks are ordinary. The economics are not.

The report is explicit that none of the intrusions it documents depended on a technique defenders have never seen. Stolen credentials, unpatched edge devices, exposed services, phishing. The tradecraft is the tradecraft.

What changed is the economics. Anthropic reports that AI has collapsed the labor and tooling gap that used to separate a well-resourced state programme from an opportunist, and that the same work now runs inside agent harnesses at machine speed and in parallel — breaches completed in two to three hours, dozens of victims handled at once by a single operator. The report’s own conclusion is the one worth carrying into a board meeting: sophistication has stopped being a reliable signal of who is behind an operation. Intent separates actor classes now, not capability.

That has a direct consequence for anyone who has ever been told their organization is too small to interest a serious adversary. The report’s observation on opportunistic actors is that security through obscurity is no longer viable and everything internet-connected is a target. The cost of coming after you has fallen to the point where being unremarkable is not a defense.

Attackers run agent swarms

The operational detail is where this gets uncomfortable, because it will be familiar to anyone who has built with agents this year.

Anthropic describes one cluster running a lead agent that decomposed reconnaissance and post-exploitation across parallel subagents, kept target lists and standing instructions in persistent memory so a campaign could resume mid-flight, and ran a collection fleet unattended. The report documents the same group running an autonomous loop against security-appliance firmware. In the influence-operations chapter it describes doctrine written into markdown files and reused across hundreds of sessions, with banned-word lists carried inside the agents themselves.

Read that list again as an architecture rather than as a threat. A lead agent. Decomposed subtasks. Persistent memory across sessions. Written standing instructions. Version-controlled doctrine. Explicit prohibitions the agent carries with it.

That is a governed agent programme. It is running on the other side.

The mirror is worth naming plainly, because it is the whole argument of this piece: attackers write down what their agents are permitted to do, and carry those instructions forward deliberately from one session to the next. Most defenders cannot produce that document for their own agents.

Detection stopped imposing cost

One case makes the point sharply. Anthropic reports that a cluster it tracks as GTG-20006 — an attribution it assesses as consistent with public reporting on Midnight Blizzard — pointed monitoring agents at its own malware, detecting when a security product flagged a sample and rebuilding it automatically.

The report’s framing of what that does to the defender’s economics is the sentence security leaders should sit with: adversaries can now bypass detections faster than defenders can deploy them.

Detection is not thereby useless. It is the floor. But a floor is not a record. A detection stack tells you what you caught. It does not tell a regulator, an underwriter or opposing counsel what you authorized, who reviewed it, and when that authorization expires. Those are different questions, they are asked by different people, and only one of them is answered by a dashboard.

In fairness to the counterweight, the report does not read as an argument that defense is failing. Anthropic documents its own countermeasures in the same volume — organization-level attribution, strengthened extraction classifiers, identity verification — and reports that the influence operations it disrupted mostly failed to reach authentic audiences, with their widest reach coming through state media distribution rather than through the AI. The picture is not that the attackers have won. It is that the cost of attacking fell faster than the cost of defending, and the gap between the two is now bridged by evidence rather than by tooling.

The chapter nobody read

The section drawing the least attention is the one most companies should read first, and it is not an intellectual-property story. It is a third-party risk story.

Anthropic reports that several AI labs relayed their own users’ prompts into Claude. In one case the report describes a provider serving Claude’s responses to people who believed they were using a different model — roughly three hundred thousand relayed requests over ten days. The relayed sessions are described as carrying names, email addresses, corporate data and live credentials, much of it moving through model routers common in the United States and Europe. In one instance the report describes a relayed session exposing a government database credential.

Sit with the shape of that, not the scale of it. An employee opens a tool their company approved. They type something containing a customer name, a contract term, a key. The prompt is relayed to a provider nobody in the organization has ever assessed, under terms nobody has read, and the answer comes back branded as the tool they chose. Every control that governs that exchange was written for a vendor relationship that does not describe what actually happened.

And the line that should end the debate about whether this is a governance problem or a procurement one: the report states that the safeguards preventing misuse do not transfer when a model is distilled. The capability moves. The controls do not.

The model path, drawn

The model path: an employee prompt passes through an IDE plug-in or coding harness, a model router, a reseller or proxy, and a serving provider before reaching the answering provider. The two ends are marked on the approved-model list with the control status note that the endpoint is visible but the path still needs evidence; each of the four hops between them is marked an unassessed hop, labelled with what it can see — full prompt, repository context, local files, chosen destination, account identity, response, retention terms — and whether the organization has paper on it: usually none, rarely, sometimes, or assumed but often unverified. The four are bracketed as the unassessed subprocessor path.
The approved model list covers the two ends. Everything between them is a subprocessor.

Most organizations maintain an approved model list. Very few maintain a model path. Those are not the same document, and the difference is every hop the diagram shows: the IDE plug-in, the coding harness, the router that chooses a provider on price or latency, the reseller with its own terms, and the provider that actually serves the response.

Each hop can see the prompt. Each hop is a place where data rests, is logged, or is relayed onward. The governance question is not whether the hop is malicious. It is whether you have paper on it: a named owner, an assessment, a processor term, and a date. Where the answer is no, you have an unassessed subprocessor in the middle of your workflow — and, as the report’s distillation chapter shows, the model you believe is answering may not be the model answering.

What this is actually asking of you

All of it lands on governance rather than on tooling. If your agents hold credentials and standing permissions that were never written down, reviewed or revoked, you have the attacker’s architecture without the attacker’s discipline.

Four things worth doing this month.

Inventory the model path, not just the approved model list. The list names what you permitted. The path names what actually carries the data — and it is the one an assessor will ask to see.

Treat routers and resellers as processors and paper them accordingly. They receive personal and confidential data on your behalf, which is the definition that matters. Outsourcing the work never outsources the obligation, and the report’s relayed-session cases are what that principle looks like when it is ignored.

Rotate and scope non-human credentials on an evidenced schedule. The credentials in those relayed sessions were live. The control that matters is not whether a key exists but whether anyone can show when it was last scoped and rotated.

Write down what each agent may do, who approved it, and when that approval expires. The adversary in the swarm case did exactly this. It is not an advanced practice. It is the baseline, and most organizations do not have it.

Do all four and you still are not finished, because none of it is self-evidencing. What a regulator, an underwriter or opposing counsel asks for is a record: an AI acceptable use policy that names who may use what, an executive risk report with severity rationale and named owners, a board memo showing the decision reached the people accountable for it, and a dated, third-party-verifiable way for an outsider to confirm the first two are genuine without being handed their contents.

And it is not a one-time exercise. The report covers activity through August 2026. The next one will cover the months after it. An assessment that was accurate in September is a historical document by the following quarter — which is why governance is a programme with a cadence, and why an artifact with a date on it is worth more than an assessment without one.

Literacy is the floor

Underneath all of it sits recognition, because a control nobody understands is a control nobody applies.

The users in the relayed-session cases had no way of knowing which provider answered them. A person who cannot recognize that possibility cannot meaningfully acknowledge a policy about it, cannot flag it when it happens, and cannot be expected to ask the question that would surface it. This is not a training nicety. EU AI Act Article 4 has required AI literacy of providers and deployers since February 2, 2025, and Article 50’s transparency obligations have applied since August 2, 2026.

The version that works is role-depth rather than one organization-wide module: what a board member needs to recognize is not what a developer needs to recognize, and neither is what the person pasting a contract into a browser tab needs to recognize. The SanctumShield AI Literacy Academy is built for that shape, with a record of who completed what.

Three questions before you go

The report rewards being read rather than summarised. Three questions from it, with the evidence and the page reference on each answer — then ten more if you want them.

Cyber operations
1 / 3

What best describes the report’s shift from AI as an assistant to AI as an orchestrator in cyber operations?

Choose an answer to see the evidence

Close

The uncomfortable finding in this report is not that adversaries have new capabilities. It is that they have adopted, deliberately and at scale, the operating discipline most defenders have not: written instructions, persistent memory, explicit scope, and a record that survives the session.

Sophistication is no longer the signal. What separates one organization from another now is whether anyone can produce the record.

So the question is worth asking twice, because it is answerable today and it is the one you will be asked later. Who actually answered your prompt — and can you show the paper on everything in between?

Run the free Shadow AI Risk Calculator. Twelve questions, no account. It starts the inventory this report is asking every organization to have.

Source

Anthropic, Detecting and countering misuse of AI: September 2026, published September 10, 2026 — the report. Page references are the report’s own printed page numbers.

  • Scope: seven harm areas, activity December 2025 – August 2026p. 3
  • Collapsed labor and tooling gap; sophistication no longer a reliable signalp. 5
  • Cost inversion: detections bypassed faster than they are deployedp. 9
  • Security through obscurity no longer viable; everything internet-connected is a targetp. 12
  • Agent swarms, persistent campaign memory, unattended collection fleet, appliance looppp. 24–26
  • Familiar attacks, changed economics: two-to-three-hour breaches, parallel victimspp. 38–39
  • Intent rather than sophistication distinguishes actor classesp. 38
  • Monitoring agents detecting and rebuilding flagged malware (GTG-20006)pp. 6, 9
  • Influence-operations doctrine in markdown, reused across hundreds of sessionspp. 42–43
  • Influence operations mostly failed to reach authentic audiencespp. 43–44
  • Distillation defined; access via proxy transfer stations, fake accounts, stolen keyspp. 143–144
  • Safeguards do not transfer when a model is distilledp. 146
  • Relayed user sessions containing names, corporate data and credentialsp. 146
  • Responses served to users who believed they were using a different model; ~300k requests in ten daysp. 148
  • Relayed session exposing a government database credentialp. 150
  • Countermeasures: organization-level attribution, strengthened classifiers, identity verificationpp. 153–154
Test yourself on the report

Can you pass the AI misuse quiz?

Ten questions drawn from all seven harm areas in the same report. Five minutes, no account, and every answer carries the report’s own page reference so you can check it yourself. See if you can score 80% or better — and whether you can spot what most teams miss.

Take the quiz →
Want the whole report first?

This piece is an argument about one chapter. If you want the map instead — all seven harm areas summarised chapter by chapter, with the page number behind every claim — start with the briefing.

The report in ten minutes →Read all 154 pages →
Free Shadow AI Risk Audit

See what your current stack is missing — in 12 questions.

The SanctumShield free Shadow AI Risk Calculator runs in your browser. No account, no email, no credit card. Twelve questions, instant risk score, three primary findings tailored to what you submit.

Perspective · outside the 27-week sequence · see the full series →

Who Actually Answered Your Prompt? — SanctumShield