LLM agents in production: orchestration, tools, and idempotency
“Agent” became a marketing word for any chatbot with tools. In production, an agent is a stateful control loop: observe context, choose an action (tool or message), apply side effects, repeat — under budgets, auth, and idempotency rules the model was never trained to guarantee.
This study is about designing loops that survive real users, duplicate webhooks, and partial outages — without pretending the LLM is the source of truth for business state.
Separate orchestration from generation
The failure mode of v1 agents is a single thread where the model decides both what to do next and how to say it. CRM funnels, approvals, and compliance paths need deterministic transitions owned by code.
A durable split:
| Layer | Owns | Should not own | |-------|------|----------------| | Orchestrator (state machine, workflow engine) | Current step, allowed tools, guards, handoff | Wording of the reply | | LLM | Natural language for the current step, structured tool args within schema | Whether payment was captured | | Tools | Side effects with auth and validation | Open-ended “do whatever” APIs |
The model proposes; the orchestrator commits state changes after tools succeed. If create_lead fails, the conversation stays in COLLECT_CONTACT, not in a fictional “done” state.
Tools are APIs, not prompt decorations
Every tool exposed to the model is a public surface:
- Narrow schemas (enums, max length, required fields). Reject out-of-schema calls before they hit the database.
- Server-side authorization on every tool invocation — never trust the model’s implied user id.
- Timeouts and rate limits per org; voice and SMS surfaces need stricter caps than internal admin chat.
- Return structured errors the model can read (“duplicate SKU”, “office closed”) instead of stack traces.
Idempotency-Key, dedupe on (org_id, client_request_id)) are non-optional when webhooks retry.State: where it lives and what gets logged
Conversation state belongs in your store (Postgres, workflow engine), not only in the model’s context window. Persist at minimum:
- Current orchestration step and version of the flow definition
- Entity ids created in this session (lead, ticket, draft quote)
- Channel metadata (WhatsApp thread id, call id, e-mail Message-Id)
When context gets long, summarize under rules (structured fields first, free text second) instead of dumping raw chat until the model forgets the active step.
Reliability patterns that actually get used
Human-in-the-loop for irreversible or high-value actions: send quote, change price, delete data. The agent prepares; a human or policy engine approves.
Compensation, not hope: if step 3 succeeded and step 4 failed, workflow code enqueues cleanup or marks partial state — do not ask the model to “undo.”
Circuit breakers on tool backends: after N failures, switch to degraded mode (collect callback number, create ticket) instead of infinite tool loops.
Concurrency: two tabs or a phone call plus webchat on the same lead need a single writer model (row lock, session id, or merge policy). LLMs do not resolve race conditions.
Async boundaries
Not every effect belongs in the request that answered the user. Pattern:
1. Synchronous path: validate, persist intent, return user-visible reply. 2. Publish event (Kafka, SQS, etc.) for enrichment, CRM sync, second-pass LLM, outbound campaigns. 3. Consumer applies with at-least-once semantics and idempotent handlers.
Voice and live chat stay inside tight latency; heavy work scales on workers. The agent loop should know which tools are inline vs enqueue-only.
Observability for agents
Standard APM is necessary but not sufficient. Add domain spans:
orchestrator.transitionwith from/to statetool.invokewith name, latency, success, org id (no PII in span attrs unless policy allows)llm.completionwith model, token counts, finish reason- Retrieval spans if RAG feeds the step (chunk ids, hit count)
Security and abuse
Agents multiply attack surface: prompt injection via retrieved docs, tool injection via user content, exfiltration through “summarize this URL.” Mitigations that ship:
- Treat retrieved text as untrusted; system instructions cannot be overridden by chunk content (instruction hierarchy + output validation where needed).
- Allowlisted domains for browse/fetch tools; no arbitrary SSRF from agent tools.
- Output filters for secrets patterns before sending to external channels.
Testing agents
Unit-test orchestrator transitions with no LLM. Contract-test tools against fixtures. Integration tests with recorded tool responses. LLM evaluations on golden dialogs for tone and tool selection, not for math.
Snapshot-testing model prose rots quickly; snapshot-testing state sequences (IDENTIFY → COLLECT → CREATE_LEAD) stays stable.
Closing
Production agents are workflow software with an LLM front end. State machines, strict tools, idempotent side effects, async workers, and decision logging are what make them operable. The model’s job is to communicate within the current step — not to be the database of record.
If you are adding voice or multichannel later, keep one orchestration graph and swap channel adapters; parallel agent implementations per channel diverge in a quarter.