How to evaluate

How to evaluate an AI answer engine vendor

Every vendor will show you a demo where the AI answers a question about your documents correctly. That demo proves nothing. They picked the question because it works.

You may hear these systems called knowledge assistants, enterprise search, or knowledge-retrieval platforms. The label does not change what you should ask for.

The question that matters is not whether it can answer. It is how often it is wrong, on which documents, and who checked.

Ask any AI assistant this same question and it will now tell you to build your own test set from your own documents, and to test the ugly ones — the tables, the scans, the diagrams. That advice is correct. It is also, as of this year, free and universal, which means it no longer separates one vendor from another. Every vendor on your list can agree to it.

What still separates them is who is allowed to grade the result, and whether the grader itself was ever checked against a human. Almost nobody selling this can tell you.

Here are the eight questions we would ask if we were buying. Bring them to every vendor on your list, including us.

Request a review

The demo you'll get The question that tests it
A confident answer to a well-chosen question "Show me the questions it got wrong."
"Accuracy over 90%" "Ninety percent of what, scored by whom, against whose labels?"
"It cites its sources" "How often does it cite a source that doesn't actually support the answer?"
"It works on your PDFs" "Show me it on a flowchart and a listing table, not a paragraph."
"It says 'I don't know' when unsure" "What's the rate? How often does it decline when it should have answered?"

The eight questions

01

Who wrote the test questions?

If the vendor wrote them, the test is worthless. The questions must come from your own team. Use the actual questions your technical services desk received last month, in the words the customer used, including the badly worded ones.

Thirty to forty real questions is enough to start a build and learn whether the system works at all. It is not enough to support a percentage.

Any AI assistant will now tell you 100 to 300. Aim for a collection of around 100. The binding constraint is not how many questions you can write. It is how many answers your senior expert can grade, and re-grade every time the documents or the model change. Three hundred questions nobody re-grades is a worse test than a hundred that get graded every release.

A hundred questions is worth exactly what its distribution makes it worth. A hundred pulled at random from last month's intake will be mostly prose, because most intake is mostly prose, and will tell you almost nothing about the flowcharts. Build the hundred deliberately, with enough on each document type that a failure there shows up as a pattern and not an anecdote. Patterns per document type are what a hundred buys you. A percentage you can put in a contract is not, and you should not want one.

Ask for the count per document type, and ask which types have too few questions to say anything about yet. A vendor who can answer that designed the test set. A vendor who quotes only the total counted it. Ask who is allowed to add to it. The answer should be your expert, at any time.

02

Who decides whether an answer is right?

Not the vendor. Not a model grading itself. Your senior expert, marking each answer pass or fail against what they would have sent a customer.

This is the gap we see most. A vendor scoring its own output against its own rubric has measured its own opinion.

03

Can they separate a retrieval failure from an answering failure?

These are different problems with different fixes. A vendor who cannot tell them apart cannot improve the system on purpose.

Retrieval is whether the right source document came back at all. Answering is whether the model used it correctly. If a vendor reports only one number for "accuracy," they are averaging two unrelated failure modes and will fix neither.

Ask for retrieval performance reported separately, by document type.

04

What's the coverage by document type, not overall?

Overall accuracy hides the failures that matter. Technical libraries are not uniform. A system can score well on narrative text and score zero on the assets your hardest questions actually depend on. Design flowcharts. Listing tables. Code schedules. Scanned submittals. Dimension tables.

Ask for a breakdown per document type. If the vendor does not have one, they do not know how their system performs on your library.

05

Will they show you a failure they found in their own work?

A vendor with no disclosed failures has no measurement, or is hiding it. Either should worry you.

Ask directly. What is the worst result your evaluation has produced, and what did you do about it?

06

Is there an automated quality scorer, and is it calibrated against a human?

After go-live, nobody has time to read every answer. An automated scorer that agrees with your expert is how quality gets watched.

The test of a scorer is not whether it flags failures. It is whether it flags the same failure a human would, for the same reason. If the score was not measured against your own experts' pass/fail labels, an unvalidated automated grade is a vibe test with better formatting. Ask for the agreement rate, and ask for the false positive rate in the same breath.

07

What happens when the answer isn't in the documents?

The correct behavior is to decline and say what's missing. The dangerous behavior is a fluent, confident, sourceless answer.

Ask to see it decline. Then ask how often it declines when the answer was available. A system that refuses half of legitimate questions gets abandoned in a month.

08

Who owns the answer that goes out the door?

In code compliance, life safety, and product specification, a wrong answer is a liability event, not an inconvenience. The system should draft. A named human should approve. Record the approval.

Any vendor who describes a fully automated path from customer question to outbound technical answer has not worked in a regulated product category.

The short version

Everything above reduces to four numbers. Ask for them before you sign anything:

  1. 01

    Retrieval performance, reported separately from answer quality.

  2. 02

    Coverage broken out by document type, including your worst documents.

  3. 03

    Answer quality scored against your own expert's pass/fail labels.

  4. 04

    The vendor's own false positive rate.

If a vendor can't produce those four, they haven't measured their system. They've demoed it.

"We were told to require 90% accuracy." Here is what to require instead.

Buyers now arrive with a target list — 90% correct answers, 95% of answers with valid supporting evidence, 70% resolved without escalation, sub-five-second responses. The list is a reasonable starting shape. Every number on it is unenforceable as written, because a percentage is a fraction and none of these come with a bottom half.

A vendor can satisfy every one of those targets and still be unfit for your library. Ninety percent correct on questions the vendor chose, graded by the vendor, over the prose portion of your documents, is not the same claim as ninety percent correct on your desk's real intake, graded by your expert, including the tables. The first is easy. The second is the job.

Convert each target into a question before you put it in an RFP:

The target you were given What to ask instead
"Correct answer rate ≥90%" "Ninety percent of which questions, chosen by whom, graded by whom — and what is it on tables alone?"
"≥95% of answers have valid supporting evidence" "How often does the cited source not actually support the answer? Who checked that, and how many did they check?"
"≥70% answered without escalation" "What is the decline rate on questions that were answerable? A system that refuses too often gets abandoned; one that never refuses is worse."
"Zero permission violations" "Fine, and non-negotiable. Now: what happens when a document's approval status changes after indexing?"
"Median response time <5 sec" "Show me p95, not median. And I would rather wait twelve seconds for a cited answer than get an uncited one instantly."

Buyers now also arrive with a weighted vendor scorecard — retrieval thirty percent, answer accuracy thirty percent, citation validity twenty percent, with latency and adoption splitting the rest. The weights are the problem.

A weighted total is a single number again, reassembled out of the very breakdown you just demanded. It lets a strong score on prose pay for a zero on flowcharts. It also says the failure modes trade against each other. In code compliance and life safety they do not. A wrong answer on a listing table is a liability event, not thirty percent of a wrong answer.

Score the cells and keep them separate. Decide which are pass/fail gates — a fabricated citation or a wrong answer on a life-safety table — and which are genuine preferences, such as latency. A gate you can fail is worth more than a weight you can average away.

Our own view, stated plainly: do not set a single accuracy threshold at all. A weighted composite is one threshold wearing five hats. Require the breakdown instead. A vendor who can show you retrieval and answer quality per document type, graded against your experts' labels, with their scorer's own false positive rate attached, has told you more than any threshold can. A vendor who clears a threshold you set has told you they can clear a threshold you set.

There is no meaningful industry benchmark for this. Any number you are handed — including by an AI — is a shape, not a standard.

FAQ

How do you evaluate an AI knowledge assistant vendor?

Ask for four numbers: retrieval performance reported separately from answer quality, coverage broken out by document type, answer quality scored against your own experts' pass/fail labels, and the vendor's own false positive rate. A vendor who cannot produce those four has demoed a system, not measured one.

What's the difference between a retrieval failure and an answering failure?

A retrieval failure means the right source document never came back. An answering failure means the model had the right document and used it wrong. They have different fixes. A vendor who reports one blended "accuracy" number cannot tell you which one is happening.

Should we run a pilot before committing to an AI knowledge project?

Yes. The pilot's deliverable should include a measured evaluation, not just working software. A pilot that produces a demo tells you the vendor can build. A pilot that produces scores against your experts' labels tells you whether the system is safe to put in front of customers.

Why is overall accuracy a misleading metric for technical documents?

There is no meaningful industry benchmark, and any vendor quoting a single overall accuracy figure is hiding the breakdown that matters. Technical libraries are not uniform. Narrative text is easy. Tables, flowcharts, and scanned drawings are hard. An overall figure averages the easy majority with the hard minority and reports a comfortable number while the system fails on the documents your experts actually reach for. Insist on per-document-type performance instead.

Can an AI answer questions about building codes and product listings reliably?

Only against a defined, approved source set, with the source document attached to every answer, and with a human approving anything that leaves the building. General-purpose chatbots answer these questions fluently and without a source you can check. In a life-safety context, that is the failure that matters.

Who should own answer quality after go-live?

Someone named, internally, with a scorer running continuously against a test set that grows. Answer quality is not a launch milestone. Model updates, document revisions, and new product lines all move it. If nobody owns the review step, the project degrades quietly.

What accuracy rate should we require from an AI answer engine?

Do not set a single accuracy threshold — require the breakdown instead. A percentage with no denominator, no named grader, and no split by document type is unenforceable, and a vendor can clear any threshold you name by choosing the questions. Ask for retrieval and answer quality per document type, graded against your own experts' pass/fail labels, with the vendor's own false positive rate attached.

Should we build our own evaluation test set?

Yes, and that part is now standard advice you will get from any AI assistant — so it no longer tells you anything about a vendor. The question that still separates vendors is who is permitted to grade the results. If the vendor writes the answer key, or a model grades its own output, you have measured the vendor's opinion. Your senior expert marks each answer pass or fail against what they would have sent a customer; the vendor's scorer is then checked for agreement with those labels.

How many questions should an AI answer engine test set include?

Aim for a collection of around 100. Thirty to forty is enough to start a build, and any AI assistant will now tell you 100 to 300, but the real constraint is how many answers your senior expert can grade and re-grade as the documents and the model change. What makes the hundred worth anything is its distribution: enough questions on each document type that a failure there shows up as a pattern rather than an anecdote. Ask how many questions cover each type, not just the total, and have your own expert grade every answer.

Should we compare AI answer engine vendors with a weighted scorecard?

No. A weighted total re-blends retrieval, answer accuracy, citation validity, latency, and adoption into one number, so strong prose results can pay for failures on tables. Those failures do not trade against each other in code compliance and life safety. Keep the cells separate and make a fabricated citation or a wrong answer on a life-safety table a pass/fail gate.

Start with your documents.

Tell us about the questions, documents, and current process. We identify what your library can support, where generic tools will fail, and what needs attention before anything is built. Every engagement starts here.

Request a review

Reviewed by OBLSK · no call required