Transfer customer settlement funds
Held by policy · P-12 · $80,000
Waiting on Priya Raman · 10:03
Aestus is one record for work done by people and AI agents: why it started, who did it, what it produced, who approved it, what it cost, and whether the result was accepted.
AI use is rising fast. The harder job is turning that activity into work the business can accept, measure, and trust.
Example — fictional team and data. Not a live run or customer result.
Ship the governed customer import
AES-RUN-184 · Launch Operations
Choose a fictional research brief or code handoff. Read the draft, check the rules, and follow the correction to a review decision.
Fictional review examples. All notes, people, outputs, and decisions are made up. No customer data or live action is used.
Read both full transcriptsFictional review examples. All notes, people, outputs, and decisions are made up. No customer data or live action is used.
Three fictional interview notes. These people are not Aestus customers.
Brief: Summarize the three fictional interview notes. Link each claim to a quote. Keep names private.
Cover all three notes and attach a matching quote to each summary claim.
Use only anonymous note labels; include no person or company names in the output.
Source material:
Note A: “I need to see the source beside the claim.”
Note B: “I cannot tell which draft is ready to use.”
Note C: “Tell me the checks before the work starts.”
Research lead (fictional) · Brief v1
The three anonymous source notes below are the complete input. The reviewer will check both rules before using the summary.
Research agent (fictional) · Summary v1
Note A: The reviewer needs the source beside each claim. Quote A: “I need to see the source beside the claim.”
Note B: The handoff needs a clear version. Quote B: “I cannot tell which draft is ready to use.”
Note C: The reviewer needs the checks before work starts. Quote: missing.
Research reviewer (fictional) · Summary v1
All three notes are summarized, but the third claim has no quote. A plausible summary is not enough.
Rule 1: Not met: the Note C claim has no matching quote attached.
Rule 2: Met: the output uses only anonymous note labels.
Research reviewer (fictional) · Summary v1 — sent back
Decision: send back Summary v1. Add the exact Note C quote to its claim. Keep the anonymous labels. This version is not accepted.
Research agent (fictional) · Summary v2
Note A: The reviewer needs the source beside each claim. Quote A: “I need to see the source beside the claim.”
Note B: The handoff needs a clear version. Quote B: “I cannot tell which draft is ready to use.”
Note C: The reviewer needs the checks before work starts. Quote C: “Tell me the checks before the work starts.”
Research reviewer (fictional) · Summary v2 — accepted
Decision: accept Summary v2 against these two rules. The reviewer checked all three claims against their supplied notes. Summary v1 remains sent back; the decision applies to v2 only.
This fictional decision shows the review method. It does not prove that AI output is always correct.
Rule 1: Met: all three notes have a summary claim and a matching quote.
Rule 2: Met: the output contains no person or company names.
A separate fictional review of the Meridian customer-import release handoff. This is not the original 71-minute journey or its P-12 approval story.
Brief: Prepare the customer-import-v1.0.0 release handoff for Maya. Show the required checks and explain how to return to the prior release.
Include results for invalid-row rejection and duplicate-safe retry checks.
Include rollback instructions and name the person who owns the rollback decision.
Source material:
Fictional check report: invalid-row rejection passed; duplicate-safe retry passed.
Fictional release owner: Maya. No repository, deploy, or external tool is connected to this example.
Maya, release owner (fictional) · Brief v1
A release handoff needs more than a passed check. Maya will compare the written handoff with both rules before accepting it.
Coding agent (fictional) · Release handoff v1
Release handoff v1: customer-import-v1.0.0. Check report: an invalid row is rejected; retrying a valid row does not create a duplicate. Both checks pass in this fictional report.
Rollback instructions: missing.
Maya, release reviewer (fictional) · Release handoff v1
The required check results are present. The handoff does not say how to return to the prior release.
Rule 1: Met: invalid-row rejection and duplicate-safe retry results are included.
Rule 2: Not met: rollback instructions and their decision owner are missing.
Maya, release reviewer (fictional) · Release handoff v1 — sent back
Decision: send back Release handoff v1. Add rollback steps and the decision owner. Passed checks do not replace the missing handoff requirement.
Coding agent (fictional) · Release handoff v2
Release handoff v2: customer-import-v1.0.0. Check report: an invalid row is rejected; retrying a valid row does not create a duplicate. Both checks pass in this fictional report.
Rollback instructions: pause new imports, restore the prior release, then check the saved row count before imports resume. Maya owns the rollback decision. These are fictional handoff notes, not instructions for a real system.
Maya, release reviewer (fictional) · Release handoff v2 — accepted
Decision: accept Release handoff v2 against these two rules. The reviewer checked the report and rollback notes. Release handoff v1 remains sent back; the decision applies to v2 only.
Accepting this fictional document does not deploy code, authorize a payment, or prove production safety.
Rule 1: Met: both required check results remain in the handoff.
Rule 2: Met: rollback steps are present and Maya owns the decision.
62% of organizations are testing agents, but only 39% report enterprise-level EBIT impact. Most of that group attributes less than 5% of EBIT to AI (McKinsey, 2025). Four missing links keep activity from becoming accepted work.
Output grows, but nobody can show which goal it served or what should stop when priorities change.
Briefs, runs, decisions, and reviews sit in separate tools. Each person rebuilds the story.
A model bill says what a call cost. It does not say whether the final result was accepted.
80% report unintended agent actions (SailPoint, 2025). A log alone cannot show whether an action was allowed.
Aestus joins the plan, the work, the run, the decision, and the outcome. Start with one workflow, not a migration.
1 · Strategize & Plan
Connect goals to the projects, tasks, runs, and artifacts that serve them. The reason for the work stays attached as it moves.
People and connected agents report progress on the task, so the goal view stays tied to live work.
Move from a goal to the work, runs, and artifacts beneath it without rebuilding the chain.
Example — fictional team and data. Not a live run or customer result.
2 · Execute
Keep the brief, assigned worker, discussion, artifact, review, and acceptance decision on the same card.
Assign work to a person, agent, tool, or workflow. Keep each handoff and decision with the task.
Lay out steps, choose participants, set approval points, and inspect a dry run before release.
Connect Claude Code, Codex, Cursor, GitHub, Slack, or any supported MCP client to the same task lifecycle.
Example — fictional team and data. Not a live run or customer result.
Launch Operations · AES / Board
10:11
Featured · Sandboxes early access
Early-access teams run Claude Code and Codex in remote sandboxes with organization-set access to repositories, networks, credentials, tools, models, and budgets.
Start Claude Code or Codex from your terminal. The session runs in the governed sandbox.
Name what the session can reach. Other network and credential requests are refused and recorded.
Review the terminal record, tools, policy refusals, and cost with the work.
Example — fictional team and data. Not a live run or customer result.
Developer terminal
Live3 · Accelerate Results
Aestus records model and tool spend, time, retries, review, rework, and acceptance for the complete workflow.
See which workflow and producer reached acceptance, what it cost, and where time or retries accumulated.
Replay, evaluate, and canary a candidate. Promote it only after the evidence meets your quality bar.
Ship the governed customer import
AcceptedIndexed cost per accepted run · run 1 = 100
Run
Optimization events
Where the cost came down
Run 1Run 10
Fictional example — the 100-to-55 curve is not measured savings or a customer result. Measure your own workflow to establish a baseline.
Completion cost values: Run 1: 100; Run 2: 94; Run 3: 86; Run 4: 82; Run 5: 78; Run 6: 69; Run 7: 66; Run 8: 60; Run 9: 58; Run 10: 55. Optimization events: Run 3: Context reused; Run 6: Smaller model · acceptance unchanged; Run 8: Policy check moved pre-execution. Saved by component: Model 30 → 15; Context preparation 18 → 7; Retries 12 → 4; Rework 8 → 4; Coordination 10 → 7; Tools 10 → 8; Review time 12 → 10.
4 · Secure Everything
Aestus keeps the active policy, approval decision, actor, action, and evidence with the run that produced the result.
Review the workflow definition, tool access, and approval path before a new version goes on duty.
A governed action can continue, wait for a named approver, or stop under the pinned policy boundary.
Keep the actor, authority, policy, action, and hash-chained receipt in one verifiable record.
Example — fictional team and data. Not a live run or customer result.
Dry-run report
Ship the governed customer import
rev 14AES-184
01Goal set
Would runMCMaya Chenwould read the goal brief
est. 0:20 · 8 credits
02Spec drafted
Would runAestus Planningwould write the import spec
est. 8:00 · 14 credits
03Code written
Would runClaude Codewould call GitHub · open PR #184
est. 16:00 · 9 credits
04Checks run
Would runAestus Reviewwould call GitHub · run the checks
est. 17:00 · 7 credits
05Reviewed & approved
Would holdPRPriya Ramanwould hold for sign-off under P-12
Every transfer.create call waits for Priya Raman
waits · 10 credits
06Shipped & learned
Would runMCMaya Chenwould call Slack · accept customer-import-v1.0.0
est. 8:00 · 7 credits
Aestus joins those pieces around one accepted result. It connects to the tools you already run.
Use one real workflow to establish the cost, time, quality, and approval baseline. Then decide what is worth improving.