“The agent can only issue refunds up to fifty dollars.” It is a reassuring sentence, and the behaviour it describes may well be what the ledger shows. The question that decides whether it is a control is what makes it so, and there are three possible answers, each of which Module 0 places at a numbered control point.
Control point 1, the system instructions. The limit is written into the text the model is given at the start of every request. Compliance is then probabilistic, because the model reads that text alongside everything else in its context and weighs it against the rest. Anything that changes the context can change the behaviour: a long conversation, an unusual phrasing, a retrieved document containing contrary text, or a model version upgrade. Nothing outside the model evaluates the limit, so nothing rejects an action that exceeds it and nothing necessarily records that it was exceeded. A limit at this control point is therefore advisory, and enforcement strength is the term for what happens when the agent attempts to exceed a limit. An advisory limit is read by the model and may be disregarded, because nothing outside the model evaluates it. A preventive limit causes the action to be refused, because a component between the model and the target system rejects the call. A detective limit permits the action and records it, so that somebody can find it afterwards. The classification describes the mechanism rather than the quality of the wording.
Control point 4, the harness. The code that assembles the tool call reads the amount and refuses to issue the call when it is above the threshold. This is a preventive control, it is deterministic, and it can be tested by reading the code path. It holds for every call that passes through that code and for no others, which is why it fails when a second path reaches the same API, and a second path is what an agent-to-agent handoff or a new channel creates.
Control point 5, the target system. The refund service itself rejects amounts above the threshold for this identity. This is the strongest of the three, because it holds irrespective of which caller arrives and therefore survives a caller whose behaviour has been influenced by whoever controls the model’s input. It is also the only one of the three that a change to the agent cannot weaken.
Establishing which of the three is in place, and saying so plainly, is the substance of the work, because the three carry different residual risk and the same sentence describes all of them. The words used in an interview are a reliable indicator of which is meant: “we told it not to” describes the system instructions, “we do not let it” describes the harness, and “it cannot” describes the target system. Only the third is worth anything as a control, and only where the configuration has been read.
This is a specific case of a principle that predates AI. A control enforced by the component it is meant to constrain is not a control. A payment limit enforced by the payment requester would not be accepted, and an access restriction enforced by the party being restricted would not be accepted. The system instructions are a request to the model, delivered through the same channel that carries every other input to the model, and they compete with those inputs for influence rather than governing them.
The reason that competition exists is worth stating precisely, because it is the part the phrase “system prompt” conceals. The system instructions are not a privileged layer that the model consults separately. They are text in the context window, sharing that window with user messages, retrieved documents, tool outputs, and conversation history. Their influence can therefore be diluted by volume, contradicted by retrieved content, or removed entirely when the context is truncated to fit. Treating them as configuration rather than as input is a category error, and it produces an assurance position that reads as preventive and behaves as advisory.
Testing by asking the agent establishes less than it appears to, because the thing being sampled is a probabilistic process rather than a rule. An agent that refuses one request has demonstrated one refusal on one phrasing, which is a single sample from a probabilistic process, and reporting it as evidence that a control holds is an evidence-quality defect rather than a shortcut. The configuration is what establishes the control: the API specification, or the validation rules in the harness, read directly.
The transaction record then corroborates it from the other direction. Where the limit is enforced at control point 4 or 5, no transaction above it exists, so the population of transactions is itself the test. Where transactions above the limit do exist, the question is settled without any further inspection. A single row above the stated limit ends the discussion, because it is an instance of the system acting rather than an account of what it was designed to do.
This brief is one of ten behind Module 1, The tool surface, a free guided walkthrough of an AI agent assessment at a credit union. The same engagement can be run unassisted, with the check questions above put to you against evidence rather than against a description.
NexNith advises boards and audit committees on exactly this work. How we work