What is Omission Error?

Evaluation
Definition

A failure where the system leaves something out rather than stating something wrong: a missing finding, an uncited claim, a dropped medication. In clinical AI it is the dominant error class and the hardest to catch, because the output reads as complete.

Why It Matters

Most AI quality work targets things the system got wrong. An omission error is what the system did not mention, which is a different problem: there is no incorrect sentence to flag, no contradiction to catch, and no failed check to read.

That difference matters in clinical work. On NOHARM, a clinical safety benchmark, the best-performing models still produced recommendations with potential for severe harm in roughly 1 in 14 consultations, and failures of omission drove most of the serious ones. A capable model is not automatically a safe one, and the gap shows up as absence rather than error.

Why It Evades Review

The output reads as finished. Ambient documentation that omits a finding is fluent, internally consistent, and formatted correctly, so a reviewer checking it for problems finds none. The only way to see the gap is to compare the artifact against the source it was built from.

Reviewer behavior compounds it. Once drafts are usually good, reading becomes confirming, and a reviewer’s attention goes to phrasing rather than coverage. Time pressure pushes the same way: the fast review looks at what is there.

How It Is Measured

Completeness needs its own check, separate from accuracy. Accuracy asks whether what the artifact says is supported. Completeness asks whether everything the encounter required is present, which means comparing against the source rather than reading the draft.

Two mechanisms make that practical. First, an automated pass that extracts the findings, medications, and actions from the source and flags what the artifact does not mention, so the reviewer starts from a list of gaps. Second, clinician-graded sets scored for completeness as well as correctness, because a harness that only measures what the output says will report a clean score on an artifact that left out the finding.

How Flytebit Handles It

Every clinical artifact gets an omission check against its sources before a person sees it, and the gaps arrive with the draft rather than after sign-off. Evaluation sets built with the client’s clinicians score completeness as a first-class dimension alongside accuracy, and human review is positioned as the check on what the automated pass could not resolve. The industry application is on our Healthcare & Life Sciences page, and the evaluation method is in Evaluating Agentic AI Systems.

More info

On flytebit.com

Reviewed by Jayaveer Bhupalam, Founder & CTO Last updated September 28, 2026