An Integration Architect Does Not Let the Agent Own the Write

A copilot “fixed” a stuck order by calling create-order again. The chat said the first call timed out. The warehouse already had the order. The second call created a twin. Nobody had sent an Idempotency-Key. The model was not the bug. The missing door was.

An integration architect treats the agent as another caller. It must enter through the same contract as any partner API. Conversational memory is not the system of record.

Think of a night-shift clerk who can open the warehouse system and also keeps a notebook. You do not let the notebook invent SKUs. You give the clerk the same receiving dock, the same stamp, and the same rule: one delivery note, one receipt. If the stamp machine jams, you check the dock book. You do not write a second delivery note because the clerk’s notebook says “I think it failed.”

This sits next to AI architecture boundaries and decisions that survive security review. The door I actually ship is the config-driven facade in integration-azure (commit 2539242), composed like the rest of the connector catalog.


ByteByteGo-style diagram: user asks an agent, the agent calls a facade door with auth schema and Idempotency-Key, the bus carries one CloudEvent, and the system of record owns the write. A dashed path shows chat memory must not skip the door.
Figure 1. The agent proposes. The system of record keeps the truth. The facade accepts once only when the caller sends Idempotency-Key. Without that header, every POST publishes. Click image to zoom.

What an integration architect changes when the caller is an agent

Partner SFTP and partner HTTP already have a shape. An agent feels different because it talks, retries, and sounds sure. I do not give it a private write path because of that. I ask the same three questions I ask of any new system: who may call, what payload is legal, and what happens if the caller tries twice.

Caller May do Must not own
Partner file drop Land a file on the agreed path The order master
Partner API POST a schema’d event The pipeline ledger
Interactive copilot Read, draft, propose a payload A write that stock or money depends on
Background agent Call the same door, with its own identity A side channel into the ERP “because the tool can”

Microsoft’s shared-responsibility note for agents is the same split in their words: a plain LLM waits for a human to act on the answer; an agent acts through tools. You still own tool permissions, identity, and what those tools are allowed to change. See AI agent shared responsibility.

Azure’s Well-Architected pattern says the orchestration layer may hold ephemeral state for the request and should not persist it past that. Design retries and idempotency on purpose. A chat thread is ephemeral state. I do not promote it to a ledger. Source: Architecture pattern for AI workloads on Azure.

The door is a facade, not a new agent tool

In functions/facade, one HTTP trigger serves every configured route. The trigger itself is anonymous (AuthorizationLevel.Anonymous, catch-all {*path}). Each route’s auth block decides who gets in. Adding an endpoint is a config change, not a new Function. That is the point of a door: the agent does not get its own Function App.

The sample route checks a delegated scope. The key is optional. This is functions/facade/config/facade.sample.yaml, trimmed to the fields that matter:

defaults:
  idempotency:
    header: Idempotency-Key
    ttl: "24:00:00"
routes:
  - id: order-created
    auth:
      mode: aadJwt
      audience: "api://integration-facade"
      requiredScopes: [events.publish]
    response:
      onSuccess: { status: 202 }
      onIdempotentReplay:
        status: 200
        body: { status: duplicate }
      onSinkError:
        status: 503
        headers: { Retry-After: "30" }

Auth modes in FacadeOptions.cs are None, FunctionKey, AadJwt, ApiKey, and Hmac. AadJwt checks issuer, audience, lifetime, and the scp claim. It does not read roles. A background agent token that only has app roles will fail requiredScopes: [events.publish]. I do not pretend the sample route is an agent-identity route. I add a route whose auth matches the token I will actually send.

Entra’s planning guide is the identity half of that sentence. For most agents, an agent identity is the right object. Autonomous work has no user present and uses application permissions (admin consent). Interactive work has a signed-in user and delegated permissions; the token subject is the user, with the agent as actor. I do not paste a user’s password into a tool. Guide: Plan your agent identity architecture.

If the caller sends the header, the facade records {routeId}:{idempotencyKey} in Cosmos before it publishes. A replay maps a response and returns. It does not call the sinks again.

// functions/facade/Functions/FacadeFunction.cs
var isFirstSeen = await idempotencyStore.TryRecordAsync(
    route.Id, idempotencyKey, idempotency.Ttl, ct);
if (!isFirstSeen)
{
    var replayMapped = mapper.Map(route.CloudEvent, context);
    return ResponseBuilder.Build(
        request, route.Response.OnIdempotentReplay, /* … */);
}

Two limits I want on the table, because an integration architect who hides them will ship the twin-order bug in a nicer costume.

  • No header, no dedupe. The block runs only when defaults.idempotency is set and the header is present. An agent tool that “just POSTs” can double-publish.
  • The key is recorded before the sinks succeed. TryRecordAsync runs, then sinks publish. If a sink fails, the key is already stored. A retry with the same key takes the replay path and does not publish again. A 503 plus a loyal retry can look like “duplicate” while the event never landed. I would not call an agent write-tool done until a sink failure releases that key, or the replay path checks whether a sink actually accepted.

Also: facade does not write the pipeline ledger. The README is explicit. The ledger starts downstream in subscription, after the event is on a queue. A 202 means the door accepted a CloudEvent. It does not mean the ERP committed. I say that in the tool contract so the agent does not tell the user “done” on a 202.

Read, draft, or write

Eng managers can use this table without reading the Function. Engineers can turn each row into a tool permission.

Action Agent Gate
Read a status Yes, scoped Least privilege. No write scope on a read tool.
Draft a payload Yes A human or a policy sees the diff before send.
Write stock, money, or customer master Only through the door Schema, auth, and Idempotency-Key. 202 only if sinks accept — not “the model decided.”
Retry after timeout Same key, new POST Not a lookup. First create publishes. 202 only if sinks accept. 409 replay does not publish. No header: no dedupe — do not do that. 503 after a burned key: stop and verify.

I also want a correlation id on the event, not only in the chat. The sample CloudEvent copies x-correlation-id, or falls back to the trace id, into an extension. That is how I find the write later. The chat transcript will not be there at 2 a.m.


ByteByteGo-style diagram: a timeout retry is another POST with the same Idempotency-Key, not a Cosmos lookup. First create publishes sinks and returns 202 only if those sinks accept. A sink failure returns 503 and burns the key. A Cosmos 409 replay does not publish. No header means no dedupe.
Figure 2. A retry is another POST. The key is a write, not a lookup. 202 only if the sinks accept. Replay does not publish. Click image to zoom.

Failure modes I stop blaming on the model

What it looks like What actually happened What I change
Twin order after a timeout Tool retried with no Idempotency-Key Key required on every write tool. Same key on retry.
“Duplicate” but nothing downstream Key stored, then sink failed, retry took replay Release the key on sink failure, or verify acceptance before replay.
Agent says the ERP updated Door returned 202. Ledger not written yet. Tool result says accepted, not committed.
Background agent gets 401 Route requires scp. App token has roles. Match auth to the token shape. Do not reuse a user refresh token.
Write skipped the bus Tool called the ERP client directly One door. Same schema as the partner API. See the MCP gateway for the tool side of that rule.

Partial failure has a second edge. Sample sets onPartialFailure: fail. If one sink accepts and another fails, the request fails, and the key is already burned. The accepted sink is not replayed. That is why I keep write routes to one sink unless I have a story for the other.

FAQ

Is the agent a system of record?

No. It is a caller. The system of record is the ERP, CRM, or ledger that still owns the truth after the chat is gone. An integration architect draws that line before the first tool ships.

Does chat memory count as a ledger?

No. It is ephemeral. The facade idempotency store is also not the pipeline ledger. It is a short-lived Cosmos row, {routeId}:{key}, with the sample TTL of 24 hours. The pipeline ledger starts in subscription.

Can an autonomous agent call the sample order-created route?

Not if it presents an app-only token. That route’s aadJwt check looks for events.publish in scp. App roles will not satisfy it. Give the agent its own route, or extend the check, and use an agent identity rather than a copied user secret.

What does HTTP 202 mean here?

The sample success response. FacadeFunction returns it only after sink results pass the success rule. With onPartialFailure: fail (the sample and the default), any failed sink returns 503, not 202. So 202 means the configured sinks accepted the CloudEvent publish. It still does not mean the ERP committed. Tell the user “accepted,” then confirm downstream.

Where should the retry live?

In the tool contract, with a stable Idempotency-Key, on a new POST. There is no lookup API. TryRecordAsync creates {routeId}:{key} or gets a Cosmos 409. Not in a prompt that says “try again if it fails.” If the first attempt already returned 503, that key is burned and a same-key retry will not publish. Stop and verify in the system of record. Do not mint a new key and POST a second order.

I still compose that door in YAML with the other connectors, rather than forking a Function per agent. The capstone is compose, don’t fork. The agent is a new caller on an old dock. The stamp does not move into the notebook.

Similar Posts