A governed document pipeline is an inference bill in a costume. Optim sits underneath Gothink AIR and routes every page to the smallest model that still clears your accuracy floor, meters cloud and token usage in real time, and forecasts the next quarter as a distribution.
You do not buy Optim. It ships in the runtime at every tier, and it is why your cost per page beats the vendor quoting you next.
Prices per token keep falling and bills keep climbing, because usage grows faster than prices drop — and a pipeline that reads a whole estate is the largest inference bill most firms will ever run.
You cannot control a real-time cost with a monthly report — and you cannot cut a bill you can’t predict or attribute.
A clean native-PDF statement and a photographed handwritten amendment are not the same problem, and paying frontier prices for both is how document AI budgets die in year two. Optim scores each page and sends it down the cheapest lane that still holds the line your risk function drew.
Swipe to follow the routing →
A forty-page pack is rarely uniform — page 3 a clean table, page 27 a faxed signature block — so routing is decided page by page, and one bad page does not price the whole document at the top rate.
Text layer, image quality, layout density, table structure, language, and the field set the pack expects from that page — with no content leaving your boundary to decide.
If the cheaper route cannot meet a field’s confidence threshold, the page escalates to a larger model; if that still cannot clear it, the value abstains into the steward queue. Optim is never permitted to move a value past a gate.
A page that fails the cheap lane is re-run, not accepted — at a fraction of a cent, versus a confidently wrong field in a client statement.
Estates are repetitive — the same master agreement underlies four hundred amendments — so recurring context is cached rather than re-sent, which is most of the saving on a large corpus.
Which models occupy the small, mid and large lanes is configuration set at Fitting; in your own VPC, your endpoints, contracts and rates — Optim routes between them, it does not resell them.
Routing is one lever of four. They compound in a deliberate order — cheapest and least costly to quality first. In a representative model, $63.9k → ~$22.4k of monthly inference before any commitment discount.
| Lever | What it does | Effect |
|---|---|---|
| Semantic + context caching | Similar prompts served from cache; repeated context billed at a fraction of fresh input. On a document estate this is the big one — the same master agreement underlies four hundred amendments. | first lever |
| Model routing & right-sizing | Straightforward pages go to a smaller model; escalates only when the confidence gate demands it. | ~80% / answer |
| Batch & off-peak scheduling | Latency-tolerant traffic moves to asynchronous endpoints. A nightly backfill of ten years of archives does not need the interactive path. | ~50% cheaper |
| Provisioned throughput | The forecast sets your 24/7 baseline; the floor is pre-bought and bursts fall back to pay-per-token. | forecast floor |
Caching is first because it is the only lever that gets better with popularity; every other lever trades something — a smaller model trades headroom, batching trades latency, a commitment trades flexibility — while cache hits trade nothing.
Routing is second because it is measurable per page: the router scores routable traffic against your confidence gates before moving any of it, and the split is a number your team can read.
Not a monthly bill you reconcile afterwards — a live meter inside your boundary that knows which family, queue and tenant caused the spend, while the run is still going.
| Metered | Attributed to | Refresh | Why it matters |
|---|---|---|---|
| Input & output tokens | Document family, pack, run, queue | Live, per run | Tells you which estate is expensive before finance does |
| Model lane mix | Small / mid / large share per family | Live, per run | A drift toward the large lane is the earliest signal your estate quality changed |
| GPU hours & idle time | Node pool, region, tenant | Live | Idle capacity scales to zero between bursts — in your VPC that is your bill, not our margin |
| Cache hit rate | Document family | Live | A falling hit rate explains a rising bill without anyone guessing |
| Cost per page | Family, pack, business unit | Live, rolling | The single number to compare against whatever the last vendor put in a slide |
| Cost per exception | Steward queue, approver | Per run | Shows what a threshold actually costs, so setting one becomes an informed decision |
Document estates are bursty in ways averages hide — quarter-end, renewal season, a portfolio transfer, a regulatory deadline — so an average forecast is comfortable and wrong exactly when it matters. Optim forecasts a distribution and provisions against the tail that hurts.
A distribution tells you when the expensive version of next quarter becomes likely — the only moment intervention is still cheap.
Every forecast is a range with stated confidence — capacity provisioned against the upper band, cost planned against the middle, so a busy quarter is a plan rather than an incident.
Short horizon drives autoscaling during a run; long horizon tells you which volume band you will be in at renewal — before the conversation, not during it.
Learned from your own run history — day-of-month effects, quarter-end spikes, family growth and lane-mix drift. It is a model of your estate, not an industry benchmark.
Budget thresholds fire when the forecast crosses them, not the spend — a projection that breaches next month is something you can still act on.
Model an unonboarded document family against your measured cost per page before committing — useful for deciding which estate to bring on next, and for saying no to one.
Optim predicts from your history; a genuinely new document family or a step change in volume sits outside the distribution until it has run a few times. When the model is unsure, the bands widen — and that widening is shown rather than smoothed away.
A cost panel beside the accuracy and throughput panels, broken down by document family and queue. Same screen, same login, same roles.
Every run closes with a measured cost per page and the lane mix that produced it — including the free evaluation, so you can model your full estate before anyone quotes you.
Usage, attribution and forecast series are readable through the same authenticated API as records and lineage, so this data lands in your own FinOps tooling.
Forecast breaches and lane-mix drift fire as webhooks into whatever you already use. Optim does not ask to own your alerting.
It was previously sold on its own as a cloud and AI spend platform, and it has not been switched off or abandoned — it has been repositioned. The forecasting, attribution and routing engine now runs underneath Gothink AIR, where the unit economics work.
If you hold a current Optim commitment, the same support, release cadence and named contact hold for the term you contracted; renewal is an honest conversation, not a migration deck.
Two crowded markets and two buyers cannot be served well by one small team, and we would rather say so than sell you a second product we cannot properly support. If cloud and AI spend governance is your actual problem, the FinOps category has good dedicated vendors we will point you to.
Existing Optim customer wondering what this means for your renewal? You will get a straight answer.
Talk to us about an existing agreementGothink AIR reads your document estate into validated records with a page reference behind every field, deployed where your regulator requires, with the weights staying yours.