What is Failure Design?
Delivery & EngineeringDeciding in advance what an AI system does when it is wrong, uncertain, or unavailable: who it tells, what it undoes, and what it stops doing. Designed before launch rather than discovered during an incident.
Why It Matters
Every system fails. The engineering question is whether the failure was designed, and in AI systems the default is that it was not: the model is asked to answer, so it answers, and the failure arrives as a confident sentence rather than an error code.
Failure design is the practice of deciding the failure behaviour while the architecture is still open. It is also where the difference between a pilot and a production system shows up first, because a pilot has never had to fail in front of anyone.
What It Covers
The refusal path. What happens when the system should not answer: an unsupported question, a request outside its authority, a source it cannot reach. Declining is a feature, and it has to be built rather than hoped for.
Degradation. How the system behaves when a dependency is slow or down: queue, hand off to a person, or stop. The wrong default is to keep answering from memory.
The reversal. What can be undone, within what window, and how the reversal is confirmed against the system of record rather than assumed.
The escalation. Who is told, with what context, and how quickly. An escalation that arrives without the evidence is a second problem rather than a solution.
The stop. The kill switch, and what the system does with work already in flight when it fires.
The record. What the failure looked like from inside, kept so the next one is diagnosable rather than mysterious.
Where It Breaks
No refusal path. The most common gap. A system that answers everything has no way to decline, so it invents, and the invention is what somebody acts on.
Silent degradation. Falling back to a smaller model or a cached answer without telling anyone produces a system that looks healthy and behaves differently.
The reversal nobody tested. A rollback path described in a runbook and never exercised is a plan, not a control. The first test of it should not be the incident it was written for.
Failure handled by the model. Asking the model to decide when it is unsure produces a system that is confident about its own confidence. The judgement belongs in the layer outside.
Monitoring that measures uptime. A service can be up and wrong. Availability metrics describe the infrastructure; the failure modes above are about the behaviour.
How Flytebit Handles It
We write the failure behaviour into the design review: the refusal path, the degradation rule, the reversal window, the escalation route, and the stop condition, each enforced outside the model. Then we test them, because a failure path that has not been exercised is a claim rather than a control. The gate between prototype and production is our implementation oversight work, and the enforcement points are designed in our architecture engagement.