Understand what each AI workload costs
Access arranged as part of a pilot
Analect Forecast is the AI cost modeling workflow in the Analect platform. Describe your workloads in plain language or import your provider bill, price each one across the model catalog, and see the reasoning behind every recommendation before you commit spend.
The problem finance actually has
AI spend arrives as one line on a provider invoice. Nobody can say which workload drove it, whether the model was the right one for the job, what caching would have saved, or what next quarter looks like. That is not a pricing problem. It is an operating-model problem, and it is not solved by finding a cheaper model.
What Forecast does
Describe the work
Workloads captured in plain questions or full detail: volumes, context, tool calls, budgets.
Start from your bill
Import the provider CSV. See what you actually paid against what the work should cost, and your effective discount.
Recommendation
Which models in the catalog can do each job, what every scenario costs, and what your budget ceiling would buy. Every forecast carries the date its prices came from.
Report
One row per workload with the reasoning behind each pick and caching as its own lever, plus the uncertainty on each figure. Prints to PDF.
Ingestion and routes
What the corpus costs to make answerable before anyone asks it a question, what re-processing costs when documents change, and the fragment route priced against the document route including both.
The Forecast loop
Measure
Import the bill or describe the workloads. Establish what is actually being asked, how often, at what size.
Model
Price every workload across the catalog. Run scenarios. Set the ceiling and see what it buys.
Decide
Route each workload to the best-fit model. Choose fragment or document per job. Switch caching on where it pays.
Govern
Publish the report. Own the budget. Re-forecast as prices, models and volumes change.
Step four feeds step one. Every re-forecast starts from the bill you just paid.
From an invoice to an operating cost
Per workload, not per invoice
Cost attaches to the job that caused it, with an owner and a reason.
A ceiling that means something
The budget planner shows what your number buys, so trade-offs are explicit before the spend.
A forecast that states its date
Model prices move constantly. Every forecast carries the date of the catalog it was priced against, because a figure without one cannot be checked.
Forecast turns AI from a black-box invoice into a managed operating cost. It is also the tool that prices the fragment route against the document route — including the up-front decomposition and the cost of refreshing it — which is how you find out whether the saving survives contact with your actual corpus.
What a forecast row contains
One row per workload. Finance signs off the report, so every row has to carry the reasoning as well as the number.
| Field | What it contains |
|---|---|
| Workload | The job being priced, described in business terms rather than as an API call. |
| Volume and shape | Requests per period, context size, output size, tool calls per task. |
| Model and price | The model chosen, its price, and the date the catalog was priced. A cost without a price date cannot be checked. |
| Route | Fragment route against document route, priced separately so the comparison is visible rather than asserted. |
| One-off costs | Decomposing the corpus before anyone asks it a question. Paid once, and again whenever a document changes. |
| Ongoing costs | Refresh as documents change, plus storage for whatever you retain. |
| Caching | Modeled as its own lever, because it moves the bill independently of model choice. |
| Range, not a point | A low and high figure with the assumptions that produce each. A single number would be false precision. |
This is the shape of the report, not a result. The figures in yours come from your own documents.
What you can change, and what moves when you do
Every figure rests on these. They are editable rather than baked in, because a model whose assumptions you cannot see or change is not evidence.
| Assumption | Effect |
|---|---|
| Request volume | Scales the recurring cost roughly linearly. The dominant term for most workloads. |
| Context size per request | Where the fragment route earns its keep, or does not. Small contexts have little to save. |
| Model choice | Changes price per token by an order of magnitude. Often a bigger lever than anything else on this list. |
| Cache hit rate | Applies only to repeated prefixes. Optimistic assumptions here are the most common way a forecast turns out wrong. |
| Corpus churn | How often documents change, and therefore how often decomposition is paid again. High churn can erase the saving entirely. |
What we will not claim in advance
These are answered by running this on your own documents, not by a specification sheet. They are the things most likely to decide whether the approach works for you, which is exactly why guessing at them would be unhelpful.
- Which of your provider bill exports parse cleanly, and which need mapping
- Your actual workload mix, taken from the bill rather than estimated
- Whether your contract pricing differs from public list pricing, and by how much
See your workloads priced.
A pilot prices your real workloads across the catalog and puts the fragment route next to the document route, so the saving is a number before it is a decision.