Skip to main content

Optimization

Track the full cost of accepted AI agent work.

Measure the whole workflow, not only the model call.

Aestus joins spend, time, retries, review, rework, and acceptance so you can compare changes against the same quality bar.

The invoice stops before the work does.

Model spend is only one input. The accepted result also carries context preparation, tool use, retries, review, waiting, coordination, and rework. Cut one line without seeing the rest and the total can rise.

enterprise message volume, year over year
reasoning-token consumption per organization, twelve months
320×
OpenAI, 2025

Ship the governed customer import

Accepted

Completion cost by accepted run

Indexed cost per accepted run · run 1 = 100

0255075100
−45% vs Run 1

Run

Optimization events

  1. Run 3 · −8Context reused
  2. Run 6 · −9Smaller model · acceptance unchanged
  3. Run 8 · −6Policy check moved pre-execution

Where the cost came down

Run 1Run 10

Model
30−15
Context preparation
18−11
Retries
12−8
Rework
8−4
Coordination
10−3
Tools
10−2
Review time
12−2
Total
1005545

Fictional example — the 100-to-55 curve is not measured savings or a customer result. Measure your own workflow to establish a baseline.

Completion cost values: Run 1: 100; Run 2: 94; Run 3: 86; Run 4: 82; Run 5: 78; Run 6: 69; Run 7: 66; Run 8: 60; Run 9: 58; Run 10: 55. Optimization events: Run 3: Context reused; Run 6: Smaller model · acceptance unchanged; Run 8: Policy check moved pre-execution. Saved by component: Model 30 → 15; Context preparation 18 → 7; Retries 12 → 4; Rework 8 → 4; Coordination 10 → 7; Tools 10 → 8; Review time 12 → 10.

A router sees a call. Aestus records the outcome.

Use call-level routing and the workflow-level acceptance record together.

The router's view

One API call, one price.

  • Blind to quality; every downgrade is a bet.

The Aestus view

The full run, with acceptance attached.

  • Compare candidate routes only after they meet the same evaluator and review bar.

Same standard. A cost change you can verify.

Five places to test a lower total cost.

Each change keeps the result and its acceptance evidence in view.

  1. Context: stop rebuilding the brief.

    Before

    Context rebuilt at every handoff.

    • Request, agent run, review: each one missing pieces.

    With Aestus

    One context bundle travels forward.

    • New decisions added, nothing rebuilt.

    Payoff: less time rebuilding context.

  2. Routing: prove the fit for each step.

    Before

    Every step on the premium model.

    • Maximum spend, even for routine work.

    With Aestus

    Each candidate runs against the same bar.

    • Compare model, agent, tool, or person routes on quality, time, and cost.

    Payoff: evidence for the route you choose.

  3. Change control: test before promotion.

    Before

    A change ships on intuition.

    • No replay evidence, no canary, and no clear rollback point.

    With Aestus

    Replay, evaluate, canary, then promote.

    • Keep the candidate only when its evidence clears the release bar.

    Payoff: same job, less work, same standard.

  4. Prevention: stop avoidable work early.

    Before

    The violation surfaces at Publish.

    • After spend, review, rollback, remediation.

    With Aestus

    "Missing approval for external publish" fires at Build.

    • 2 agent runs and 3 downstream steps avoided.

    Payoff: avoidable work never starts.

    A 100-agent fleet at 2,000 actions a day is ~73M runtime evaluations a year; every check moved to design time comes off that bill. Aestus cost model — illustrative.

  5. Coordination: one path to acceptance.

    Before

    Which version is accepted? Who reviewed?

    • Salesforce, Notion, Slack, GitHub: each holds a piece.

    With Aestus

    One record holds the accepted version and its reviewer.

    • Salesforce, Notion, Slack, GitHub: each feeds the same record.

    Payoff: one answer to both questions.

Review evidence before changing a workflow.

  • Run history with outcomes

    Compare status, evaluations, acceptance, time, cost, and failures.

  • Evidence-bound candidates

    Review the change, supporting runs, and stop rule.

  • A controlled improvement loop

    Replay, review, canary, and roll back under a stop rule.

Example — fictional team and data. Not a live run or customer result.

Routing decisions

Ship the governed customer import

Accepted
  1. 1 · Goal set

    Why: Cached · Goal accepted

    Model: HumanTokens: Cost: 8
  2. 2 · Spec drafted

    Why: Cheaper · Spec bar met

    Model: SonnetTokens: 18.4kCost: 14
  3. 3 · Code written

    Why: Cached · Acceptance unchanged

    Model: HaikuTokens: 41.2kCost: 9
  4. 4 · Checks run

    Why: Retried once · Tests passed

    Model: Aestus ReviewTokens: 12.6kCost: 7
  5. 5 · Reviewed & approved

    Why: Checked early · P-12 approved

    Model: HumanTokens: Cost: 10
  6. 6 · Shipped & learned

    Why: Cached · Accepted artifact

    Model: AestusTokens: 3.1kCost: 7

Total · Run 10 of 10

Run 1 · 100

Total cost 55, down from 100 on run 1

The metric Aestus owns: cost per completed workflow.

Only an accepted result connects AI spend to a business outcome. Measure inputs and runs, but make the acceptance record the denominator.

  • Input efficiency

    Per call: model, tools, context, compute, people.

  • Workflow efficiency

    Per run: steps, retries, review, coordination, waiting, rework.

  • Completion efficiency

    Per accepted outcome: cost, time, quality, governance.

What is modeled, and what is measured.

The curve is illustrative. Your product view uses your runs. Your evaluators and reviewers define quality. Modeled values stay labeled, and recorded values link to evidence. Falsifiable by design.

We have a model gateway and a cost dashboard.

Keep them. Aestus adds the rest of the cost record: tools, retries, time, review, coordination, rework, and whether the final result was accepted.

Measure one workflow's baseline.

Talk to us