Use less context. Measure the answer.
Analect turns document content into addressable knowledge fragments, helping AI retrieve the information a question needs. Test answer quality, token use and storage requirements on your own documents before deciding where to deploy it.
Where it fits Designed for targeted, repeated questions across a document collection. Whole-document tasks and frequently changing material need a separate cost and quality assessment.
The cause is structural
A language model has no index into a document. To answer any question it reads everything, so every question pays for every token, and the same document is paid for again on the next question. Retrieval helps by sending less of the document, but a chunk is still prose. Compression shaves the prose. Neither changes what a unit of knowledge is.
Most enterprise material is unstructured, and most of it is never queried at all. The standard answer — chunks, embeddings and indexes layered on top of the source estate — works, and it makes the estate larger rather than smaller.
Change the unit of knowledge
Analect reads a document once and decomposes it, one way, into knowledge fragments: small structured statements of what the document actually asserts, each tied back to the sentence it came from. Once the knowledge is addressable, a question can retrieve the dozen fragments that hold its answer rather than the eight thousand tokens around them.
Two honest lines, because they frame everything else. Decomposition does not make a document smaller — at whole-document level it makes the material larger. And if your workload is to summarize entire documents, fragments cost more, not less. The value is in targeted, repeated questioning of a stable corpus.
What actually happens when you ask
“What is the recommended monitoring interval after the initial dose?”
- 01The question arrivesNo document is loaded. The question is matched against the fragment index.
- 02Fragments are retrievedThe handful of structured statements that bear on dosing and monitoring, and nothing else.
- 03The model answersFrom roughly 355 tokens of fragments rather than the 8,000 tokens the whole guideline costs to read.
- 04The answer cites its assertionsEach statement used links back to the sentence it came from, so a reviewer can check the answer rather than trust it.
Token figures from the clinical guideline in our published test. Your documents and questions will produce different numbers.
What we measured, and what we did not
Our published test describes the documents, questions and settings used, together with the results and the limitations. A pilot establishes whether the findings carry over to your workloads.
Three measures, kept apart
Baseline agreement
Did the fragment answer say the same thing as the same model's whole-document answer? This is what our published test measured, on nine direct factual questions, and it matched on eight of them.
Factual correctness
Was the answer true, judged against known ground truth? An error shared by both routes counts as agreement but not as correctness. We have not published this separately.
Completeness
Did the answer contain everything a complete answer needed? An answer can agree with the baseline, be factually correct, and still omit a material fact that both routes dropped.
Citation support
Can every statement in an answer be traced to the assertion it came from? Fragments carry a link to their source sentence by construction, which is a property of the design rather than a benchmark result.
The token result
The storage result
Limitations
- The comparison is against the same model reading the whole document. It is not a comparison against your current retrieval workflow, which is the comparison that usually matters more and is often a harder one to beat.
- The sample is small: nine direct factual questions across two documents. It is enough to show the effect exists and nowhere near enough to predict its size on your estate.
- The token figures cover retrieved context. They do not net off the cost of decomposing the documents in the first place, or of re-processing them when they change.
- The storage figure is the fragment layer against the source file. It excludes source retention, metadata, indexes, replicas and backups. If you keep the sources — and you should — your total storage goes up, not down.
- Whole-document work goes the other way. At document level a fragment set is larger than the source, so summarize-everything workloads cost more on fragments, not less.
These figures are from a small ETT sample: nine direct factual questions across two documents, compared against the same model reading each document whole. They show the effect exists. They do not predict its size on your estate, and they exclude the cost of decomposing documents in the first place.
Everything a benchmark has to record
This is the record we keep for every test. It is published without a form, because evidence that supports a public headline should be inspectable without talking to sales. Rows we have not yet recorded in a publishable form are listed below the table rather than left out of it.
| Field | Published | Why it matters |
|---|---|---|
| Document types | Clinical guideline and one further document | Results do not transfer evenly across document types. Tables, forms and scanned material behave differently from prose. |
| Questions | 9 questions, direct factual questions against the source material | Question count and question type determine what the result covers. |
| Baseline agreement | 8 of 9 questions agreed with the same model's whole-document answer | The measure we actually took. It is agreement with a baseline, not a measure of correctness. |
| Token categories | Retrieved context tokens. Input and output totals, cached and billed categories not reported separately | A percentage means nothing until you know whether it covers retrieved context, total input, all billed tokens or money. |
| Token result | 70–77% fewer context tokens on direct questions. One factual question answered from 355 tokens of fragments against 8,000 tokens to read the document whole | The headline figure, with its unit stated. |
| Retained storage scope | Fragment tables measured at roughly 300 KB against a 3.5 MB source document, about 91% smaller. This is the fragment layer only | It excludes source retention, metadata, indexes, replicas and backups. Total retained infrastructure is a different and larger number. |
| Comparison baseline | The same model reading the whole document | Not compared against the customer's existing retrieval workflow, which is usually the more relevant comparison and often a stronger one. |
Not yet published
These are recorded for every pilot and reported to that client. They are not yet published for our own test, and until they are, treat the figures above as an indication that the effect exists rather than a measure of its size.
- Test date. Model behavior and token prices both move. A result without a date cannot be interpreted.
- Corpus. Which documents, how many, and where they came from.
- Task selection. How the questions were chosen, and by whom. Questions selected after seeing the fragments would not be a fair test.
- Model and version. A result is a property of a specific model version, not of models in general.
- Prompt and retrieval settings. How many fragments were retrieved, and how. Retrieval settings move both cost and accuracy.
- Ground truth and scoring. Who marked the answers, against what, and whether they knew which route produced each one.
- The disagreement. The ninth question, and whether the difference was a missing fact, a wrong answer, an incomplete answer or a problem with the baseline. This is the single most useful row in the table.
- Factual correctness. Scored against known ground truth rather than against the baseline. Not yet measured separately.
- Completeness. Whether answers contained everything a complete answer needed. Not yet measured separately.
- Latency. Retrieval adds a step. Whether the round trip is faster or slower end to end is a separate question from token count.
- Ingestion and refresh effort. Decomposition has an up-front cost and a cost every time a document changes. Both have to be paid back out of the per-question saving.
- Price date. Any monetary figure is only valid at a stated set of prices.
Four workflows, and how a buyer decides with each
Every answer shows its working
Every fragment links back to the sentence it came from, so an answer assembled from fragments arrives with its evidence attached: which statements were used, from which documents, at which lines. A chunk retrieved by similarity can cite a passage; a fragment cites the assertion itself.
Only the fragments a question needs are sent to a model, under that provider's terms, which we name in the design. That reduces how much of your text leaves your estate. It is not by itself anonymisation or a security boundary, and we do not present it as one — where each processing step happens is confirmed for your deployment and validated by our technical lead.
Your source documents stay where they are
A smaller derived representation is not a reason to delete the originals, and nothing in Analect suggests you should. Fragments are an index into your knowledge, not a replacement for the record.
Each fragment links back to the sentence it came from, which only works while that source exists. Source updates are handled by re-processing the changed document; deleting a source removes the fragments derived from it.
Four situations where this does not pay
We would rather tell you before a pilot than during one. If your workload is on this list, the honest answer is that Analect is not the right tool for it.
You need the whole document every time
Summarisation, full-document review and drafting from a complete source all need the entire text. Fragments are larger than the source at that level, so this costs more, not less.
Your documents change constantly
Decomposition has an up-front cost and is paid again whenever a document changes. A corpus rewritten weekly may never repay it.
You ask each document one question, once
The saving comes from asking a stable corpus many questions. One question against one document does not amortise the processing.
Your corpus is small
If reading everything is already cheap, the arithmetic does not need changing. Fix the problem you have.
Against the alternatives, by criteria
Retrieval systems do not all behave the same way, and a comparison that says they do is not helping anyone evaluate anything. These are the criteria worth measuring on.
Prompt compression
Reduces the tokens in each request by dropping low-value text. It is lossy and applies per request, so nothing is reused by the next question. Compare on: what is lost, and whether the saving survives repeated questioning.
Retrieval and graph RAG
Sends part of a document rather than all of it, and can be very effective. The retained estate usually grows, because embeddings and indexes sit on top of the sources. Compare on: retrieval relevance, total retained storage, and refresh cost when documents change.
Gateways and routers
Reduce the price per token rather than the number of tokens. They are not an alternative to either of the above and combine with them. Compare on: model coverage, routing quality and observability.
Parsing and AI ETL
Prepare documents for a downstream pipeline. Upstream of this question rather than an alternative to it. Compare on: format coverage, extraction accuracy and table handling.
Deterministic decomposition, measured on the two bills that reach a CFO, with the measurement shipped as part of the product rather than asserted in a brochure. The right comparison for you is against whatever you run today, not against a model reading whole documents — and Proof will run that comparison.
Where the pressures stack
Sectors with large document estates, repetitive questioning, tight budgets and binding sovereignty rules.
Government and public services
Doing more with less is the operating condition, and the data cannot leave.
Legal and professional services
Precedent and matter files where every answer must cite its source.
Scope, test, review, decide
Pilot
A fixed-fee pilot on your documents. We decompose a slice of the estate, run the proof suite, and report accuracy, tokens and storage, zeroes included.
Pass
If the numbers do not justify going further, you pass, and the pilot report is yours to keep.
Play
If they do, we productionize: ingestion at scale, routing across the model catalog, governance and the storage layer, in your environment.
- Fixed fee, agreed before the work starts.
- The comparison baseline includes your current retrieval workflow, not only a model reading whole documents.
- The report includes the questions where fragments did no better, and says why.
- If the numbers do not justify going further, the report is yours and there is nothing further to decide.
Measure it on your own documents.
A fixed-fee pilot on an agreed slice of your estate returns your numbers: answer quality against a baseline you choose, token use with its categories stated, and storage under both retention scenarios.
Procurement and technical questions
What is Analect?
Analect is an enterprise knowledge platform from ETT. It decomposes unstructured documents into addressable knowledge fragments so that AI systems answer from the fragments a question needs rather than re-reading whole documents, cutting token usage and storage while keeping every answer traceable to its source.
How much does Analect reduce AI token costs?
On our own sample, answering direct factual questions from fragments used 70 to 77 percent fewer retrieved context tokens than reading the source document whole. That figure covers retrieved context, not total billed tokens, and it does not net off the cost of decomposing the documents in the first place. The sample was nine questions across two documents. Your figures come from a pilot on your own material.
Does Analect give the same answers as reading the whole document?
On eight of nine benchmark questions the fragment answer agreed with the same model's whole-document answer. Agreement with a baseline is not the same as being correct: if both routes share an error, that scores as agreement. We have not published factual correctness or completeness as separate measures, and we distinguish all three rather than reporting agreement as accuracy.
How much storage does Analect save?
On our sample the fragment tables for one document measured roughly 300 KB against a 3.5 MB source, about 91 percent smaller. That is the fragment layer measured against the source file — it excludes source retention, metadata, indexes, replicas and backups. If you keep your source documents, and you should, your total retained storage goes up rather than down. Store models both scenarios and distinguishes bytes reduced from bill reduced.
Do we have to delete our source documents?
No, and nothing about Analect suggests you should. Fragments are an index into your knowledge, not a replacement for the record. Each fragment links back to the sentence it came from, which only works while the source exists. Changed documents are re-processed; deleting a source removes the fragments derived from it.
Is Analect a RAG tool?
It solves a related problem differently. Retrieval sends less of a document; Analect changes what a retrievable unit is. Fragments are structured statements rather than prose chunks, produced by a deterministic decomposition rather than by running a language model over the corpus. Whether that is better for you is a question for a measured comparison against your existing retrieval workflow, which is what a pilot runs.
Does Analect work with our existing models?
Yes. Analect is model-agnostic. Forecast prices workloads across a catalog of public models, carrying the date the catalog was priced, and Proof lets you test the models you actually use against your own documents.
Where does our data go?
Decomposition runs inside your environment, and when a question is answered only the fragments that question needs are sent to whichever model you have chosen — under that provider's terms, which we name in the design. Where each processing step happens is confirmed for your deployment and validated by our technical lead rather than assumed. Fragmentation reduces how much text leaves your estate; it is not by itself anonymisation or a security boundary, and we do not present it as one.
What does a pilot involve?
A fixed-fee engagement on an agreed slice of your estate. We decompose the documents, run the question tests against a baseline we agree with you — including your current retrieval workflow where you have one — weigh the storage, and hand you a report that includes the results that came back flat. If the numbers do not justify going further, the report is yours and you walk away.
Which workloads does Analect not help with?
Four of them. Summarize-everything work, because at whole-document level a fragment set is larger than the source. Corpora that change constantly, because decomposition is paid again on every change. Documents you question once, because there is nothing to amortise against. And small corpora, where reading everything is already cheap. We would rather tell you that before a pilot than during one.