Designing an MCP Tool Around a Side Effect You Can't Take Back

Author
Arthur TeboulCo-Founder, DocuTrak
8 min readSummarize:OpenAIClaudeMistral AIGoogle GeminiGitHub CopilotPerplexity

When a points system counts an event twice, you have a bug with a cleanup path. You find the duplicate, reverse the award, and the leaderboard is right again. An email to a real person has no cleanup path. Once it lands, the best you can do is send another one explaining the first.

That asymmetry shapes how I design agent tools. At DokuTrak we build client document collection for accountants and other professionals: a request lists the documents needed, the client uploads through one secure link, and reminders run until the checklist is complete or the link expires. The DokuTrak MCP connector, which is open source, lets the professional's own agent, in Claude Desktop or Claude Code, create such a request. (MCP, the Model Context Protocol, is the open standard an assistant uses to discover and call tools on a server.) Creating one emails the client.

We saw how fast that goes wrong in an internal test on September 15, 2026. The whole instruction was "ask X for their ID by September 30", and Claude went straight from that line to a created and sent Document Request, with no recap and no consent. Version 0.1.1 of the connector added the consent step. This article covers the rest: the email as an event that should happen at most once, and the windows where the call can fail.

Short version: split "create the record" from "cause the effect" so a failed email is reported, not swallowed. Word every error result so the model doesn't call the non-idempotent tool again. Know which unit you would deduplicate on. And where you can, move the side effect out of the call, so the tool becomes safe to repeat.

Why the default endpoint was the wrong primitive

Our API's POST /v1/requests accepts a sendEmail flag, and it defaults to true. With the default, the service creates the request and sends the email in the same handler. If the email step throws, the error is logged and the handler carries on, under a comment that states the policy plainly: // Don't fail the request creation if email fails. The caller gets the created request back as a success.

For a person clicking a button, that's defensible: the dashboard shows the request. For an agent it's a trap: the model reports the client as asked, and the client has heard nothing. Our own web app doesn't lean on that default: it passes sendEmail: false and sends in a separate call. The connector does the same.

runTool(async () => {
  const created = await api.post<RequestDetail>('/v1/requests', {
    recipientEmail: input.recipient_email,
    // …
    deadline: toDeadlineIso(input.deadline),
    sendEmail: false,
    // …
  });
  const id = created.data.id;

  try {
    const sent = await api.post<{ expiresAt?: string }>(`/v1/requests/${id}/send`, {});
    return jsonResult({
      requestId: id,
      // …
      sent: true,
      uploadLinkExpiresAt: sent.data?.expiresAt ?? null,
    });
  } catch (error) {
    const reason = error instanceof ApiError
      ? `${error.title} (${error.status}): ${error.detail}`
      : describe(error);
    return errorResult(
      `Document Request ${id} was created but the email did not go out — ${reason}. ` +
        'It is visible in the DokuTrak dashboard, where it can be sent.'
    );
  }
})

Two calls turn one opaque outcome into two observable ones. runTool wraps every handler, so a tool never throws: every failure comes back as a result flagged isError, with text the model reads.

An aside on toDeadlineIso: it turns a bare date such as 2026-09-30 into noon UTC, "so the stored day is the day given". Midnight UTC is the previous evening across the Americas, so code converting it back to a local date would land a day early, and a due date in a sent email can't be corrected. Trophy's guide to streak time zones and DST covers the same class of bug from the streak side.

Three failure windows, three outcomes

Splitting the call gives each failure a place, and each needs a different answer.

Window What exists afterwards What the agent sees What should happen next
Creation refused (validation, auth, a service rule) Nothing: no request, no email isError with the service's own detail Fix the input and call again, but only after a real refusal
Created, send refused A pending request whose first email the service reports as not sent isError naming the request id Send it from the dashboard; don't create a second one
Send succeeded, answer lost A request and a sent email A failure that isn't one Nothing, because the job is done, but the agent can't know that

The "created, send refused" window is the one the split was built for: the failure is reported instead of swallowed, and the text points at a request that already exists. That request isn't inert, though. It sits in pending with a live upload link, and the reminder cadence counts from when that link was created, so if nobody sends it from the dashboard, the first reminder will still reach the client.

The "send succeeded, answer lost" window is where honesty matters. The connector sets no timeout of its own, and its catch block treats any error from /send as a refused send. If the connection dropped after the service had already emailed the client, the text would say the email did not go out, and that would be wrong. Either way, the agent holds a failure for an action that happened. The same can happen one step earlier: if the answer to the creation call is lost, a pending request exists that the agent doesn't know about.

Why an MCP error result invites a retry

The MCP specification is explicit about what error results are for. On the tools page, tool execution errors "contain actionable feedback that language models can use to self-correct and retry with adjusted parameters".

That's right for most tools. For a non-idempotent one, every isError reads to the model as an offer to try again, so the wording has to point somewhere else. Ours names the created request and sends the professional to the dashboard to send it.

What it doesn't do is tell the model, in so many words, not to call create_request again. A model reading "was created" will usually infer it. For an email I'd rather not rely on "usually": an explicit "do not create it again" belongs in that string.

Pick the unit of deduplication

Deduplicate on the natural unit of the event. For us that unit is "this request, to this person", and the two ways an agent can repeat itself treat it very differently.

Retrying /send for the same request is contained. The service keeps one link row per request, and each send writes a fresh token over it (refreshForRequest updates the row keyed by the request id). The client gets another email, but the link in the earlier one stops working. There is one live link per request, however many times it's sent.

Retrying create_request is not contained. It makes a new request with its own link and its own reminder cadence, so the client can end up chased twice for the same documents. There is no idempotency key on POST /v1/requests today.

Charlie Hopkins-Brinicombe's post on preventing points farming sets the bar I'd want: "The same event submitted multiple times produces the same result as submitting it once." His key, lesson-${lessonId}, works because the lesson is the natural unit: completing it once is the event. Trophy's API takes that key as an Idempotency-Key header and doesn't process a repeat. Nothing in our call identifies "this request, to this person" yet, so the "answer lost" window has no shipped fix.

One option the agent already has is to read before retrying: get_request searches by client name, email or title. The tools allow it; the description doesn't ask for it yet.

One difference from Trophy's case decides where our guard sits. A double-counted event can be reversed afterwards; a double-sent email can't. So the confirmation comes before the call, and deduplication can only limit the damage.

Make the other tool idempotent by removing its side effect

The connector's second write tool, request_replacement, handles a request with rejected files. It could have emailed the client straight away. We chose not to, and that choice is what makes the annotation true:

annotations: { readOnlyHint: false, destructiveHint: false, idempotentHint: true, openWorldHint: false },
// …
note: 'No email was sent by this call. The Client will be reached by the reminder cadence.',

The call re-flags the rejected files, sets the request to pending_resubmission, and puts it back on the reminder cadence. The reminders are what reach the client. A second call leaves the same state behind. If nothing is rejected, the service refuses with a 400 and changes nothing.

It isn't idempotent in every sense: each call inserts a request_resubmission audit row, so two calls leave two rows. A tool is idempotent in effect when a repeat leaves the same external state and the person on the other end notices nothing. That's the version that matters for an agent retrying, and moving the side effect out of the call is what made it true.

MCP tool annotations and _meta as the contract

create_request is annotated readOnlyHint: false, destructiveHint: true, idempotentHint: false, openWorldHint: true, and carries _meta: { 'anthropic/requiresUserInteraction': true }.

idempotentHint: false is information for the client, not protection. The same spec page says "clients MUST consider tool annotations to be untrusted unless they come from trusted servers", and nothing obliges a client to act on the hint. The protection comes from the other two. As recorded in our decision record for 0.1.1, Claude Desktop prompts for a destructive tool, and Claude Code prompts on each call that carries the key, even in its auto and bypass modes. Other clients may do neither.

The permission prompt proves someone was there; the recap is what they read. The description asks the model to show who the email goes to, the due date, every document worded as the client will see it, and any message, then wait for a yes "even when the request already looks complete". It's the last look anyone gets at the event before it exists.

A checklist for tools with a side effect you can't reverse

  1. Does the endpoint report the side effect's failure, or swallow it? If it swallows it, trigger the effect in a second call.
  2. Does each failure window have its own answer?
  3. Does every error result name what already exists, so the model doesn't repeat the non-idempotent call?
  4. What is the natural unit of deduplication, and does anything in the call carry it?
  5. Could the side effect move to a scheduler, so the tool is safe to repeat?

FAQ

What does idempotentHint actually do?

It tells the MCP client whether calling the tool again with the same arguments has any additional effect. A client may use it when deciding to retry or to ask the user first, but it makes nothing idempotent, and the spec tells clients to treat annotations from untrusted servers as untrusted.

Should an agent retry a failed non-idempotent tool call?

Not blindly. First check, with a read tool if one exists, whether the effect already happened. If the error names an object that was created, act on that object instead of creating another.

Sources


About the author

Arthur Teboul founded DokuTrak, where accountants, lawyers and brokers ask their clients for documents and let the reminders do the chasing. Both code blocks above come from the dokutrak-mcp repository on GitHub.

Author
Arthur TeboulCo-Founder, DocuTrak

Get the latest on gamification

Product updates, best practices, and insights on retention and engagement — delivered straight to your inbox.

The gamification layer for consumer apps

Drop-in gamification features you can ship this sprint. Increase retention and user engagement without sacrificing your roadmap.

Designing an MCP Tool Around a Side Effect You Can't Take Back