Compare answer quality and token use on the same questions
Access arranged as part of a pilot
Analect Proof is the measurement workflow in the Analect platform. It takes a document and its decomposition, asks the same questions of both against a model you choose, and totals the tokens each route consumed. It reports three things separately: whether the answers agreed with each other, whether they were correct against ground truth you supply, and whether they were complete.
Why proof comes first
Most claims about AI retrieval are made in a deck and tested nowhere. Proof exists so that the first conversation with your finance team is about numbers from your own documents rather than numbers from ours — including the questions where the fragment route did no better.
Three steps to a number
Pick the data and the baseline
A document from your library or a fresh upload, and the baseline to compare against: the model reading it whole, your current retrieval workflow, or both.
Pick the model and supply ground truth
Any model from the catalog, and the known-correct answers to score against where you have them. Without ground truth the test can only report agreement, not correctness.
Run the test
The same questions go to each route. The report gives agreement, correctness where ground truth exists, completeness, and the token totals per category. It exports.
Four measures, never conflated
Agreement, correctness and completeness answer different questions, and a tool that reports one of them as all three is not measuring anything useful.
Baseline agreement
Did the fragment answer say the same thing as the same model's whole-document answer? This is what our published test measured, on nine direct factual questions, and it matched on eight of them.
Factual correctness
Was the answer true, judged against known ground truth? An error shared by both routes counts as agreement but not as correctness. We have not published this separately.
Completeness
Did the answer contain everything a complete answer needed? An answer can agree with the baseline, be factually correct, and still omit a material fact that both routes dropped.
Citation support
Can every statement in an answer be traced to the assertion it came from? Fragments carry a link to their source sentence by construction, which is a property of the design rather than a benchmark result.
What we found on our own material
On nine direct factual questions across two documents, fragment answers agreed with the same model's whole-document answers on eight, using 70 to 77 percent fewer retrieved context tokens. Agreement is not correctness: we have not published factual correctness or completeness as separate measures, and the ninth question is not yet published with its reason.
The zeroes are part of the report
Proof reports the zeroes as well as the wins. Whole-document workloads cost more on fragments, and the report says so. That is the point: a result you can take to a CFO has to be a result that could have gone the other way. Where a question came back worse on fragments, the report shows the question, both answers and the reason.
What a scored question contains
One row per question, per route. The report is only worth anything if a row can come back negative, so every field below is reported whichever way it goes.
| Field | What it contains |
|---|---|
| The question | As asked, chosen before anyone saw the fragments. |
| Answer, per route | What each route returned: the whole-document baseline, your existing retrieval workflow where you have one, and the fragment route. |
| Agreement | Did the fragment answer say the same thing as the baseline? Agreement only. It says nothing about either answer being right. |
| Correctness | Scored against ground truth you supply. Without ground truth this field is empty, and the report says so rather than substituting agreement for it. |
| Completeness | Did the answer contain everything a complete answer needed? An answer can agree, be correct, and still omit a material fact both routes dropped. |
| Citations | Which fragments were used, and the source sentence each came from, so a reviewer can check rather than trust. |
| Tokens, by category | Retrieved context, total input, output. Reported separately, because a percentage without its category is meaningless. |
| Verdict | Including “no better” and “worse”. These rows stay in the report. |
This is the shape of the report, not a result. The figures in yours come from your own documents.
What we will not claim in advance
These are answered by running this on your own documents, not by a specification sheet. They are the things most likely to decide whether the approach works for you, which is exactly why guessing at them would be unhelpful.
- Which baselines to compare against — the whole-document read, your current retrieval workflow, or both
- Whether you have ground truth for the questions, and if not, who will mark the answers
- How many questions and documents constitute a sample you would actually believe
Run the test on your documents.
A pilot runs the same questions against your documents and their fragments, marks the answers and totals the tokens, so the evidence is yours rather than ours.