What is Control Layers?
Governance & ControlThe ordered checkpoints an AI system passes through between an input and a consequence: prevention, semantic validation, authorisation, output validation, audit, and recovery. Each layer runs outside the model, and each one can stop the request.
Why It Matters
A single guardrail fails in a single way. Filter the input and a well-formed malicious request still gets through. Validate the output and the damage is already done. The reason to layer controls is not thoroughness for its own sake: it is that each layer catches a different class of failure, and the cost of an error rises as the request moves from a sentence toward an action.
The layers also decide where a control can be honest. A check that runs before execution can refuse; a check that runs afterwards can only report. Systems that put everything in the second category produce excellent postmortems.
The Six Layers
Input prevention. Screening what arrives before the model sees it: injection patterns, personal data that should never reach the endpoint, and payloads that do not match the expected schema.
Semantic validation. Whether the request means what it appears to mean, checked against the domain: the intent it expresses, the entities it names, and whether it belongs to a workflow the system is allowed to run.
Action authorisation. The decision point at the tool boundary: allow, deny, or escalate, evaluated against policy, credentials, and limits. This is the layer that separates an assistant from an agent, because it is where a proposal becomes a consequence.
Output validation. What leaves the system: whether claims are supported by what was retrieved, whether personal data has leaked into a response, and whether the output complies with the rules that apply to this surface.
Audit and evidence. The record, written outside the agent so it cannot be revised by the process it describes, holding the policy version, the verdict, and the reviewer.
Recovery and learning. The way back and the way forward: reversal for actions already taken, a kill switch for the system as a whole, and any learning gated behind evaluation rather than shipped on the strength of an anecdote.
Where It Breaks
Layers that live in the prompt. Instructions describing how the model should behave are a request, not a layer. Every one of the six has to be enforced by a service the model cannot reason around.
Authorisation without identity. A layer that decides what may happen but not on whose behalf collapses when two users share a credential, which is the common shape of an agent deployment.
Evidence that the agent writes. A record produced by the process it describes is testimony rather than evidence. It has to be written outside, and tamper-evident.
Recovery nobody has tested. A reversal path that has never been exercised is an intention. The same is true of a kill switch that has only been described in a runbook.
Layers added after the incident. Controls are cheapest to design into an architecture and most expensive to retrofit, which is why the layer map belongs in the design review rather than the remediation plan.
How Flytebit Handles It
We design the layer map before the build, with each layer enforced by a service rather than an instruction, and the enforcement points are the ones our governance and risk engagement installs. The architecture side is our AI architecture design work, and the enforcement discipline is described in Governing Agentic AI.
More info
- OWASP Top 10 for Agentic Applications Most of the ten risks are governance problems rather than model problems, which is what the layers exist to address.
- EU AI Act: consolidated text (Article 14, human oversight) Oversight as a designed capability of the system, not a person watching a dashboard.