Agents

Agents shouldn't send email without an approval step

Aug 25, 2026 · 3 min read

An agent that drafts your email is an assistant. An agent that sends unreviewed email under your name is a liability with your signature attached. The gap between those two products is not model quality. It is one approval step between the proposal and the send.

A send is not like other actions

Most agent mistakes are cheap. A bad search wastes a second, and a bad draft wastes a click. A bad send leaves your infrastructure, lands in someone else's inbox, and carries a real person's name. There is no undo and no quiet fix in the next deploy. Whoever received it now believes your colleague wrote it, and in every way that matters, your colleague did.

Agents fail in ways the demo never shows. They reply to the wrong thread, misread tone, invent a discount that doesn't exist, or send the internal version of a message to an external address. The lesson is not to keep agents away from email; it is to separate the act of writing from the authority to send.

Approval gates as an API primitive

The pattern fits in one sentence: the agent proposes an action, the action sits in a pending state, a human approves or rejects it, and only an approved action executes. The design decision that matters is where the gate lives. A prompt that says "always ask before sending" is a suggestion the model can lose halfway through a long context. A pending state that cannot execute without an approval record is a mechanism, enforced by the API no matter what the model does.

Model the proposal as a first-class resource with an ID, a payload, a requester, and a status. The rest falls out of that choice. A review queue is a list of pending proposals. An approval is a status transition with an actor and a timestamp attached. A rejection carries a reason, and the agent can read that reason before its next attempt.

Which actions deserve a gate

Gating everything is as bad as gating nothing. If a human must approve every read, they will stop reading the approvals by Thursday, and a rubber stamp protects nobody. The useful line runs between actions that leave your system and actions that stay inside it.

  • Gate sends. Outbound email under a person's name is the canonical case, and the one with the least forgiving audience.
  • Gate deletes. Removing a thread or a calendar event destroys information you may want for the audit later.
  • Gate external bookings. An agent that books a meeting on someone else's calendar commits another person's time.
  • Leave reads, drafts, and searches ungated. Nothing leaves your system, so there is nothing to approve, and this is exactly the work you want the agent doing at volume.

The approval record is half the value

The brake is the visible half of the gate. The record is the half you keep. When someone asks why an email went out, you answer with who proposed it, who approved it, when, and what the payload was at the moment of approval. Without that record, the honest answer is "the model decided," which satisfies neither your customer nor your lawyer.

The record also changes how you tune the agent. Rejection reasons are labeled data about where it misjudges. A cluster of rejections about tone tells you something no benchmark will.

Friction is the feature, for now

The obvious objection is that approvals slow the agent down. They do, on purpose. Speed is what you earn after trust, not before it. Run a gated agent for a quarter and you will know which action types never get rejected, and then you can loosen policy per action with evidence instead of optimism.

Horato ships this shape directly: agent tools that propose instead of execute, with approvals and history recorded per run. However you build it, build the gate before the agent gets the keys.