"We were told to require 90% accuracy." Here is what to require instead.
Buyers now arrive with a target list — 90% correct answers, 95% of answers with valid supporting evidence, 70% resolved without escalation, sub-five-second responses. The list is a reasonable starting shape. Every number on it is unenforceable as written, because a percentage is a fraction and none of these come with a bottom half.
A vendor can satisfy every one of those targets and still be unfit for your library. Ninety percent correct on questions the vendor chose, graded by the vendor, over the prose portion of your documents, is not the same claim as ninety percent correct on your desk's real intake, graded by your expert, including the tables. The first is easy. The second is the job.
Convert each target into a question before you put it in an RFP:
| The target you were given |
What to ask instead |
| "Correct answer rate ≥90%" |
"Ninety percent of which questions, chosen by whom, graded by whom — and what is it on tables alone?" |
| "≥95% of answers have valid supporting evidence" |
"How often does the cited source not actually support the answer? Who checked that, and how many did they check?" |
| "≥70% answered without escalation" |
"What is the decline rate on questions that were answerable? A system that refuses too often gets abandoned; one that never refuses is worse." |
| "Zero permission violations" |
"Fine, and non-negotiable. Now: what happens when a document's approval status changes after indexing?" |
| "Median response time <5 sec" |
"Show me p95, not median. And I would rather wait twelve seconds for a cited answer than get an uncited one instantly." |
Our own view, stated plainly: do not set a single accuracy threshold at all. Require the breakdown instead. A vendor who can show you retrieval and answer quality per document type, graded against your experts' labels, with their scorer's own false positive rate attached, has told you more than any threshold can. A vendor who clears a threshold you set has told you they can clear a threshold you set.
There is no meaningful industry benchmark for this. Any number you are handed — including by an AI — is a shape, not a standard.