01
Who wrote the test questions?
If the vendor wrote them, the test is worthless. The questions must come from your own team. Use the actual questions your technical services desk received last month, in the words the customer used, including the badly worded ones.
Thirty to forty real questions is enough to start a build and learn whether the system works at all. It is not enough to support a percentage.
Any AI assistant will now tell you 100 to 300. Aim for a collection of around 100. The binding constraint is not how many questions you can write. It is how many answers your senior expert can grade, and re-grade every time the documents or the model change. Three hundred questions nobody re-grades is a worse test than a hundred that get graded every release.
A hundred questions is worth exactly what its distribution makes it worth. A hundred pulled at random from last month's intake will be mostly prose, because most intake is mostly prose, and will tell you almost nothing about the flowcharts. Build the hundred deliberately, with enough on each document type that a failure there shows up as a pattern and not an anecdote. Patterns per document type are what a hundred buys you. A percentage you can put in a contract is not, and you should not want one.
Ask for the count per document type, and ask which types have too few questions to say anything about yet. A vendor who can answer that designed the test set. A vendor who quotes only the total counted it. Ask who is allowed to add to it. The answer should be your expert, at any time.