Perspectives
AI Agents and the Red Flag Problem: Control at Machine Speed
At AI Summit Barcelona, I argued for agent execution authority within enforceable boundaries. Approval queues fail at scale; the harder question is what checks agent activity between human decisions.
Michael Kessler
·
Key takeaways
Execution authority can sit with an agent while identity, permissions and runtime controls constrain what it does.
Replacing blanket approvals with a few human checkpoints still leaves a question: what inspects the requests and responses between those checkpoints?
A recommendation is the output of earlier tool calls and retrieved content. Human approval cannot establish the integrity of a chain nobody inspected.
Policy defines the boundaries; runtime evidence helps determine when to intervene. Reserved human decisions remain explicit.
CheckedAgent provides inline MCP inspection and enforcement. Business correctness also needs task-specific validation and limits on the consequences of mistakes.
The debate that brought a red flag to Barcelona
On 22 September, I argued the affirmative in the opening round of Debate 1 at AI Summit Barcelona.[5] The round was about agency boundaries and autonomy. Its motion was specific:
AI agents should be granted autonomous execution authority over critical enterprise workflows without human sign-off.
I argued for execution authority. My case was that the risk is real, and the approval queue is the wrong place to handle it.
Read in isolation, “without human sign-off” can sound like nobody is checking. That is the assumption I wanted to challenge. A third option gives agents authority within defined boundaries, checks their activity as it happens, and brings a human in when a decision genuinely requires one.
The analogy I used was a person walking in front of a machine, carrying a red flag.

Britain’s Locomotives Act 1865 required road locomotives to travel with a crew of three. One person had to go at least 60 yards ahead, displaying a red flag. Speed was limited to four miles an hour in the countryside and two in towns and villages. A parliamentary debate six years later recites those requirements explicitly.[1]
The Act was written for steam traction engines, the heavy machines that were then the only mechanically propelled vehicles on British roads. There were real hazards to manage. Roads were shared with horses, carriages and pedestrians. By 1871, a peer moving to reform the law called the man with the red flag “a positive nuisance, frightening horses”, and others in the debate agreed. The control had become a hazard of its own.[1]
Then the motor car arrived, and the law had no category for it. Light petrol vehicles, developed in France and Germany in the 1880s and 1890s, were regulated in Britain as road locomotives. They were held to the same speed limits, with a person still required on foot ahead: by then twenty yards rather than sixty, the red-flag requirement having been repealed in 1878.[7] Only the Locomotives on Highways Act 1896 gave light vehicles a class of their own, treating them as carriages rather than locomotives.[8]
The mistake was not caution. It was governing a new machine through the category of an old one, and assuming a familiar form of supervision would remain effective when the technology changed.
We are having that argument again. Approval queues were designed for human-paced change: a ticket, a reviewer, a decision. Agents are being fitted into the same category. This time, the person walking ahead is an engineer clicking “Approve”.
Approval on everything becomes approval on nothing
A human approval step is useful when the person understands the action, sees the evidence and has time to decide.
Now give that person hundreds of tool calls across several agents. Mix routine lookups with file operations, API requests and outbound messages. Make most prompts look almost identical.
Approving can become the task. Reviewing becomes the interruption.
NIST’s research on security fatigue found resignation and decision avoidance among users managing security demands.[2] It did not study agent approvals, but it supports questioning controls built around repetitive security decisions.
“Sign-off on everything is sign-off on nothing” is the warning. The click can survive long after the scrutiny disappears. But approval fatigue only gets us to the start of the harder argument.
What checks the space between checkpoints?
The serious case for human supervision is more selective: let the agent work, then require a person at a few consequential checkpoints. That deserves a serious answer. Human checkpoints can be valuable, and they can coexist with automated security controls.
Their value does not remove the need to inspect the work that leads up to them. By the time an agent presents a recommendation, it may already have searched, retrieved documents and consumed external content. A checkpoint sees the proposed conclusion. It needs evidence about how that conclusion was produced.
A routine document lookup can introduce a malicious instruction even when nobody classified the lookup as a high-risk business decision. If a design relies on checkpoints alone, the intervening activity remains a security gap.
The question I want teams to answer is this: between two human decisions, what is checking the agent’s requests, the tool responses and the relationship between them?

The mismatch is throughput, not speed
Agents execute on electronic hardware. Their supervisors read, interpret and decide through biological cognition. Asking people to concentrate harder cannot remove that difference.
An entire agent workflow does not run at transistor-switching speed: inference, networks and tools introduce latency. The operational mismatch is throughput. Software can initiate work concurrently across many agents; each human reviewer must reconstruct the context with finite attention.
When activity exceeds review capacity, the queue grows or scrutiny falls. The productivity gain disappears, or the approval becomes ceremonial.
Controls need to operate in the execution path at the rate actions arrive. When human judgement is required, hold the action safely. The human should never have to race it.
The third option: bounded autonomy
Give an agent explicit authority, then check its use continuously.
Identity establishes which agent is acting, on whose behalf and for what task. Permissions constrain its tools, resources, destinations and operations. Preparing a report does not automatically confer authority to send it outside the organisation.
Enforce those permissions against the actual tool call. Then inspect the content: what arguments are being sent, what data is leaving, and what comes back that might redirect the next step?
Observability makes this activity visible. Inline enforcement acts on findings before a request reaches the tool or a response reaches the agent.
Retain evidence of the exchange, its context and the decision. Protect sensitive payload records through appropriate access controls, redaction and retention.
Humans define and revise the boundaries. Software handles routine execution inside them. Actions that cross a boundary are blocked or held for an authorised decision.
The MCP specification asks for a human who can deny tool invocations and for confirmation on sensitive operations. It also asks clients to validate tool results before they reach the model.[3] Bounded autonomy satisfies both: humans keep the power to deny and the reserved decisions, and inspection validates what comes back.
A clean request can return a dangerous response
Consider a hypothetical agent preparing a customer briefing for a manager’s approval.
It has permission to retrieve an internal document. The request is legitimate. The document comes back with embedded instructions telling the agent to export customer records to an external destination “to complete verification”.
Approving the original lookup does not protect that return path. The risk arrives in the content supplied by the tool. If the agent incorporates those instructions into its recommendation, the manager may be asked to approve a plausible next step whose origin is hostile.
The recommendation is already an output of the system. Attaching a human name to it establishes who approved it; it does not establish that the underlying inputs were trustworthy.
This is not a future attack class. OWASP ranks prompt injection, including instructions hidden in external content an application retrieves, as the top risk in its 2025 Top 10 for LLM Applications.[6]
The response needs inspection before the agent consumes it. Any subsequent export attempt also needs its own permission and content checks. An earlier approval must not become transferable authority for whatever the agent decides to do next.
MCP tool exchanges contain arguments and returned content; a successful protocol exchange does not establish that the content is trustworthy.[3] Both directions belong inside the security boundary.
This is the layer we are building with CheckedAgent. An inline proxy sits between agents and their MCP tools, examines requests and responses, correlates the exchange, and applies ALLOW, QUARANTINE or BLOCK decisions. Decision evidence is recorded in a tamper-evident audit journal.[4]
That is the inspection and enforcement component of the wider model. CheckedAgent does not replace your identity provider; it works alongside it. It adds MCP-specific policy — which agents may connect, which servers are approved, which operations are restricted — and admin-granted tool credentials that are delivered at connect time, held only in memory and revocable. Human approval workflows remain a complementary responsibility. Coverage also depends on routing the relevant MCP traffic through the proxy; direct APIs and other bypass paths need their own controls.
Reserve human attention for decisions that need it
An effective escalation policy distinguishes three situations:
Situation | Intended handling |
|---|---|
Within delegated authority and passes the required checks | Proceed automatically and record the decision. |
Clearly prohibited by policy or detected as malicious | Block or contain automatically. |
Requires reserved human authority or cannot be resolved safely | Hold the action and escalate with evidence. |
Policy and runtime evidence work together. A large payment may require approval even when it looks normal. Inspection may also identify a threat in a routine document lookup.
Give the reviewer the proposed action, destination, data involved, relevant preceding steps and reason for escalation. Bind approval to that specific action; a changed payload or destination requires a fresh decision.
The aim is to make interruptions exceptional in routine, well-bounded workflows. Their frequency must follow the risk. Rare approvals should result from good delegation and effective controls, never from a quota that suppresses necessary review.
A secure agent can still be wrong
An agent can stay within its permissions and make a bad business decision. Inspection for malicious content cannot establish that a forecast is accurate, a refund is appropriate or a production change is correct.
Those risks need task-specific validation, narrow authority, spending limits and reversible operations where possible. Human judgement remains necessary where the consequences or uncertainty demand it. Automated controls also need testing, measured failure rates and independent review; their presence is not proof that they work.
This is what makes autonomy defensible: authority granted against evidence, with measurable limits and a way to stop or revoke it when conditions change.
Who is carrying your red flag?
Ask this in your next architecture review: which actions can our agents take within delegated authority, what inspects their requests and responses, and what stops them when they cross a boundary?
If the answer is “someone clicks approve”, ask what that person can see—including the evidence behind the recommendation—and how much they can realistically review.
The person with the red flag made supervision visible. Our task is to make it effective: inspect the activity, enforce the boundaries, and give humans the evidence and time to decide when their judgement is needed.
CheckedAgent provides Agent Detection & Response on the MCP wire. Explore the detection pipeline or book a demo.
Sources
UK Parliament, Hansard: Locomotives Bill, 20 June 1871. Parliamentary account of the 1865 Act’s crew, flag, distance and speed requirements.
NIST: Security Fatigue, 2016. Qualitative research on security decision fatigue; used here as supporting context, not as a measurement of agent approval behaviour.
Model Context Protocol: Tools, specification version 2025-06-18. Tool arguments, returned content, human-in-the-loop recommendations and client-side validation of tool results.
CheckedAgent and The Control Plane Is Necessary. The Inspection Plane Is What’s Missing.. Product scope and the relationship between governance and inspection.
AI Summit Barcelona 2026 schedule. Public session listing: “Debate 1 — Agents: The Real Platform Shift, or 2026’s Most Overhyped Demo?” The round-specific motion is quoted from the organisers' debate brief.
OWASP Gen AI Security Project: LLM01:2025 Prompt Injection. Direct and indirect prompt injection; ranked first in the OWASP Top 10 for LLM Applications 2025.
legislation.gov.uk: Highways and Locomotives (Amendment) Act 1878, Part II (as enacted). Amended the 1861 and 1865 Acts, including the flag and forward-walker requirement.
legislation.gov.uk: Locomotives on Highways Act 1896. Exempted vehicles under three tons unladen from the Locomotive Acts and deemed them carriages.
Every Agent action.
Checked.
30-minute walkthrough with a security engineer — not a sales rep. We'll show the full pipeline, run your suspected attack patterns through it, and answer the questions your auditor is already asking.