How to evaluate

How to evaluate an AI answer engine vendor

Every vendor will show you a demo where the AI answers a question about your documents correctly. That demo proves nothing. They picked the question because it works.

You may hear these systems called knowledge assistants, enterprise search, or RAG platforms. The label does not change what you should ask for.

The question that matters is not whether it can answer. It is how often it is wrong, on which documents, and who checked. Almost nobody selling this can tell you, because they never built the tests.

Here are the eight questions we would ask if we were buying. Bring them to every vendor on your list, including us.

The demo you'll get The question that tests it
A confident answer to a well-chosen question "Show me the questions it got wrong."
"Accuracy over 90%" "Ninety percent of what, scored by whom, against whose labels?"
"It cites its sources" "How often does it cite a source that doesn't actually support the answer?"
"It works on your PDFs" "Show me it on a flowchart and a listing table, not a paragraph."
"It says 'I don't know' when unsure" "What's the rate? How often does it decline when it should have answered?"

The eight questions

01

Who wrote the test questions?

If the vendor wrote them, the test is worthless. The questions must come from your own team. Use the actual questions your technical services desk received last month, in the words the customer used, including the badly worded ones.

Ask for a written test set of at least 30 to 40 real questions before any build starts. Ask who is allowed to add to it. The answer should be your expert, at any time.

02

Who decides whether an answer is right?

Not the vendor. Not a model grading itself. Your senior expert, marking each answer pass or fail against what they would have sent a customer.

This is the gap we see most. A vendor scoring its own output against its own rubric has measured its own opinion.

03

Can they separate a retrieval failure from an answering failure?

These are different problems with different fixes. A vendor who cannot tell them apart cannot improve the system on purpose.

Retrieval is whether the right source document came back at all. Answering is whether the model used it correctly. If a vendor reports only one number for "accuracy," they are averaging two unrelated failure modes and will fix neither.

Ask for retrieval performance reported separately, by document type.

04

What's the coverage by document type, not overall?

Overall accuracy hides the failures that matter. Technical libraries are not uniform. A system can score well on narrative text and score zero on the assets your hardest questions actually depend on. Design flowcharts. Listing tables. Code schedules. Scanned submittals. Dimension tables.

Ask for a breakdown per document type. If the vendor does not have one, they do not know how their system performs on your library.

05

Will they show you a failure they found in their own work?

A vendor with no disclosed failures has no measurement, or is hiding it. Either should worry you.

Ask directly. What is the worst result your evaluation has produced, and what did you do about it?

06

Is there an automated quality scorer, and is it calibrated against a human?

After go-live, nobody has time to read every answer. An automated scorer that agrees with your expert is how quality gets watched.

The test of a scorer is not whether it flags failures. It is whether it flags the same failure a human would, for the same reason. If the score was not measured against your own experts' pass/fail labels, an unvalidated automated grade is a vibe test with better formatting. Ask for the agreement rate, and ask for the false positive rate in the same breath.

07

What happens when the answer isn't in the documents?

The correct behavior is to decline and say what's missing. The dangerous behavior is a fluent, confident, sourceless answer.

Ask to see it decline. Then ask how often it declines when the answer was available. A system that refuses half of legitimate questions gets abandoned in a month.

08

Who owns the answer that goes out the door?

In code compliance, life safety, and product specification, a wrong answer is a liability event, not an inconvenience. The system should draft. A named human should approve. Record the approval.

Any vendor who describes a fully automated path from customer question to outbound technical answer has not worked in a regulated product category.

The short version

Ask for four numbers before you sign anything:

  1. 01

    Retrieval performance, reported separately from answer quality.

  2. 02

    Coverage broken out by document type, including your worst documents.

  3. 03

    Answer quality scored against your own expert's pass/fail labels.

  4. 04

    The vendor's own false positive rate.

If a vendor can't produce those four, they haven't measured their system. They've demoed it.

FAQ

How do you evaluate an AI knowledge assistant vendor?

Ask for four numbers: retrieval performance reported separately from answer quality, coverage broken out by document type, answer quality scored against your own experts' pass/fail labels, and the vendor's own false positive rate. A vendor who cannot produce those four has demoed a system, not measured one.

What accuracy should an AI document assistant have?

There is no meaningful industry benchmark. Any vendor quoting a single overall accuracy figure is hiding the breakdown that matters. Insist on per-document-type performance instead. A system can score well overall while failing completely on flowcharts, listing tables, and code schedules. That is where the hardest questions live.

What's the difference between a retrieval failure and an answering failure?

A retrieval failure means the right source document never came back. An answering failure means the model had the right document and used it wrong. They have different fixes. A vendor who reports one blended "accuracy" number cannot tell you which one is happening.

Should we run a pilot before committing to an AI knowledge project?

Yes. The pilot's deliverable should include a measured evaluation, not just working software. A pilot that produces a demo tells you the vendor can build. A pilot that produces scores against your experts' labels tells you whether the system is safe to put in front of customers.

Why is overall accuracy a misleading metric for technical documents?

Technical libraries are not uniform. Narrative text is easy. Tables, flowcharts, and scanned drawings are hard. An overall figure averages the easy majority with the hard minority and reports a comfortable number while the system fails on the documents your experts actually reach for.

Can an AI answer questions about building codes and product listings reliably?

Only against a defined, approved source set, with the source document attached to every answer, and with a human approving anything that leaves the building. General-purpose chatbots answer these questions fluently and without a source you can check. In a life-safety context, that is the failure that matters.

Who should own answer quality after go-live?

Someone named, internally, with a scorer running continuously against a test set that grows. Answer quality is not a launch milestone. Model updates, document revisions, and new product lines all move it. If nobody owns the review step, the project degrades quietly.

Start with your documents.

Tell us about the questions, documents, and current process. We will review the opportunity before recommending what to build. If it is worth building, the next step is a six-week Technical Answer Pilot at $38,000. You get working software answering your team's real questions against your real documents, plus a measured evaluation scored against your experts' labels. You keep both regardless of what you decide next.

Request a review

Reviewed by OBLSK · no call required