What is pass@k and pass^k?

Evaluation
Definition

Two ways to score repeated attempts at the same task. pass@k is the chance that at least one of k attempts succeeds, which measures capability. pass^k is the chance that all k succeed, which measures reliability. They move in opposite directions as k grows.

Why It Matters

A single pass or fail tells you almost nothing about a system that behaves differently on every run. Ask the same question five times and you may get four good answers and one that invents a policy. Reporting “80 percent correct” hides the shape of that failure, and the shape is what decides whether the system is safe to deploy.

These two metrics separate the halves. pass@k asks whether the capability exists at all: try k times, did one attempt work? pass^k asks whether the capability is dependable: try k times, did every attempt work? A system can score high on one and low on the other, and which number matters depends entirely on who absorbs a bad answer.

The Two Metrics

pass@k. The probability that at least one attempt out of k succeeds. It rises toward 1 as k grows, so a larger k flatters it. This is the right measure when a person reviews the output and picks the good one: a coding assistant, a drafting tool, anything where the human is the filter.

pass^k. The probability that all k attempts succeed. It falls toward 0 as k grows, so a larger k punishes it. This is the right measure when one bad answer reaches a user with nobody in between: an automated pipeline, a customer-facing agent, a scheduled job.

The arithmetic is unforgiving and worth internalising. At 80 percent per run, three attempts give 0.8³, which is 51.2 percent. The same system reads as “80 percent reliable” or “a coin flip” depending only on which metric you print.

Where It Breaks

Printing the flattering one. pass@k is easier to make look good, because it improves as you sample more. A team that reports it while shipping an unattended pipeline has measured capability and deployed reliability.

A k chosen for convenience. One or two attempts is not a sample. The metrics only mean something at a k that reflects real usage, and real usage is rarely three.

Per-run accuracy as a proxy. A model that is right 80 percent of the time on a single attempt tells you nothing about whether the tenth consecutive step in a workflow will hold. Long chains compound: ten steps at 95 percent each is 59.8 percent end to end.

Averaging across task types. A number blended across easy and hard cases hides which one is failing. The metric belongs per task class, next to the count of attempts behind it.

Ignoring the cost of k. Sampling five times costs five times as much. If the metric that makes the system look good requires it, the economics belong in the same document as the score.

How Flytebit Handles It

We choose the metric from the failure mode rather than the other way round: pass^k where a bad answer reaches a user unattended, pass@k where a person reviews before it lands. Every published figure carries its k and its sample size, and evaluation runs against your own cases rather than a vendor benchmark. The reasoning is set out in Evaluating Agentic AI Systems, and the runtime side is covered by agentic observability.

More info

On flytebit.com

Reviewed by Jayaveer Bhupalam, Founder & CTO Last updated September 30, 2026