← BackReference (opens in a new tab)

Ninety-two percent correct

An aggregate accuracy score does not settle acceptability. · PM Drills · 22 · 30 sec–2 min

Ninety-two percent correct · 30 sec–2 min

Situation

An AI feature is correct 92% of the time.

The team asks whether that is good enough to launch. You have not seen the errors or the evaluation dataset.

Think before scrolling

What determines the answer?

Consider task type, error severity, reversibility, review, baseline, and how errors are distributed.

How a strong PM thinks

Inspect the failures and the operating context.

Eight percent weak brainstorming suggestions may be manageable. Eight percent incorrect account actions may be unacceptable. Ask what “correct” means, whether cases represent real use, and whether a small critical category fails disproportionately.

Compare with a simpler baseline and evaluate human review or constrained scope. Launch readiness depends on consequences and controls, not accuracy in isolation.

Takeaway

A quality number needs a task and a cost of failure.

Read the error distribution before deciding what users may rely on.