Home/Autonomy & gates
Deep dive / how much rope, and who holds it

An agent is not a chatbot wearing a finance skin.

It reads live state, plans a sequence it can defend, writes back through your APIs, then checks what changed. Autonomy is a dial your team turns — one notch at a time.

Reversible by design Evidence on every step Raised only by you
01 / the loop

Sense, plan, act, verify — and one place a person stands.

The gate sits on act, never on sense or plan. The agent may think as hard as it likes.

Exhibit A· the agent loopthe gate is on act, only
A four-step loop — sense, plan, act, verify — with a human approval gate placed only between plan and act. 01Sense 02Plan 03Act 04Verify THEN DO IT AGAIN HUMAN GATE
A copilot answers a question and stops. These close the loop — and everything interesting is in the fourth verb.

Sense

Primary records on your side of the boundary — billing exports, tag trees, token meters, IDocs, PO tables. No scraped dashboards.

Plan

Steps committed up front with the evidence behind each and the option it rejected. Below your bar, it halts and says why.

Act — gated

Executes through your APIs and your IAM, with a reversal path defined before it can fire.

02 / the rungs

Nothing arrives switched on at full autonomy.

Each class of action starts on the bottom rung and climbs only after your team has evaluated it on your own numbers.

The agent proposes; a person carries it out. Nothing executes. Every disagreement is logged — the cheapest possible place to discover the thing you were wrong about.

Reversible by designEvidence on every stepRaised only by you

Moving an action up a rung is your team’s call, made on your evidence. gothink.ai never widens an agent’s authority on its own initiative, in a release note, or by default.

03 / what “gated” means in practice

Reading is free. Changing state is not.

Runs freely

Anything that only observes

Detection, attribution, forecasting and drafting are continuous and unattended — a signal is worth little if it arrives a fortnight late.

Design targetoverspend surfaced in < 1 min
Approvalnone required
Waits for a person

Anything that changes state

A commitment purchase, a posted entry, a lifted payment block. Proposed, evidenced, and held for a named human on your audit trail.

Who approvesa named person, inside your environment
Before it can firea reversal path exists
Guardrails

Budgets and blast radius

Caps on steps, spend, tokens and records touched per run. Allow-listed operations, rate limits on writes, and a kill switch your team owns.

Trace

Replayable, down to the step

Inputs, retrieved passages, tool calls, tokens, cost, latency and the approver’s name — held in your systems.

Evaluation

Scored before it ships

The ground-truth set becomes a regression suite that runs on every prompt edit, model upgrade and tool change.

The same idea, one layer down: the confidence gate on data

Actions have approval thresholds. Individual data fields have confidence thresholds — and they work the same way: a score on every field, a bar you set per field and per document type, and anything doubtful routed to a steward instead of into the database.

Raise the bar and more goes to review; lower it and more goes straight through. Either way the trace records which happened.

See the confidence gate

The ask / start with a scoped pilot

Turn the work you keep discovering by hand into a governed capability.

Three steps, one quarter, measured against your live numbers.

1

Pick a pilot estate

One business unit, one question answered by hand, or one document family.

2

Run a discovery sprint

Two weeks: baseline, target, and a deployment plan for your environment.

3

Prove the number

Measured in your environment, evaluated by your team, before anything scales.