Calculate cost per accepted result

For a document workflow, the useful denominator might be accepted documents. For a support assistant, it could be resolved cases that meet the agreed quality criteria. Define that unit with the process owner before comparing models. Counting successful API responses can make an unreliable workflow look efficient.

Use an illustrative invoice review as a starting point. Include extraction, retrieval, model calls, storage, retries and reviewer time. Record how many invoices were accepted, rejected or sent back for correction. Keep the calculation visible enough that the team can dispute its assumptions.

A simple working measure is total workflow cost divided by accepted outcomes. It is a model for discussion, not a universal accounting rule. Separate one-time engineering costs from recurring operation and state how review time is estimated.

Observe the whole path

Google’s SRE guidance identifies latency, traffic, errors and saturation as core monitoring signals. For an AI workflow, those service signals remain useful, and we propose adding task quality and the human review burden. An application can be available while its answers become less useful.

Attach a stable run identifier to the steps of a request. Attribute the cost to model use, retrieval, external tools and retries where possible. Inspect slow or expensive outliers; an average can hide a small class of work that consumes most of the budget.

Keep the record proportionate. You may need timings, versions and counts without retaining the full contents of a sensitive document. Agree the evidence needed for support and cost analysis, then restrict its access and lifetime.

Choose complexity only where it helps

Try the simplest viable path on representative work. A rule or a database lookup may answer a structured question without a language model. A smaller model may handle a bounded classification task while a more capable one is reserved for difficult cases. Those are hypotheses to evaluate, not automatic cost wins.

FrugalGPT explores cost and quality through model cascades; RouteLLM studies routing between stronger and weaker models using preference data. Both make routing a concrete design option. Neither supplies a universal saving: the traffic, quality bar and cost of a wrong route still belong in your evaluation.

Routing creates its own failure mode: the system must recognise which case belongs on which path. Include routing mistakes, escalation and repeated work in the comparison. A cheap first attempt followed by a second full attempt can be slower and more expensive than selecting the right path at the start.

Check when a cached answer can still be used. It must be valid for the current user, document version and operational data. If the result depends on permissions or current inventory, the cache key and invalidation rules need to account for those changes.

Limit the cost and duration of each run

Define the maximum useful work a run can do. Bound retries, tool calls and elapsed time, and provide a clear result when the budget is exhausted. A person should see whether the workflow completed, stopped safely or needs attention.

Compare candidate changes on the same approved cases. Report accepted outcomes, review effort, latency and recurring cost together. If a cheaper configuration raises rework, show that tradeoff instead of presenting the model savings alone.

After launch, review the expensive and rejected cases with the business. They may point to a missing data source, an unsuitable task or a product change more valuable than another model optimisation.

MeasureWhy it belongs
Accepted outcomesDefines useful work
Review and reworkCaptures the effort left with people
End-to-end latencyReflects the actual wait
Recurring costIncludes tools, infrastructure and retries

Sources & further reading

  1. Google SRE — Monitoring Distributed Systems
  2. Chen, Zaharia & Zou — FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
  3. Ong et al. — RouteLLM: Learning to Route LLMs from Preference Data

Primary sources inform the technical background. The examples and proposed working methods are MGLO’s own; Business scenarios are identified in the text.