An agent receives a complaint, checks the order and requests a refund. The payment provider completes the operation, but its response gets lost. The agent sees a timeout. It tries again.

In this hypothetical example, retry rules, request identifiers and stored state determine whether the customer gets a second refund. Even a model that understood the entire conversation needs an interface that can establish what already happened. That is one reason to return to Designing Data-Intensive Applications when designing an AI system.

O’Reilly lists February 2026 as the publication date for the second edition, by Martin Kleppmann and Chris Riccomini. Its contents include transactions, replication, stream processing, vector embeddings and durable workflow execution. The boar on the cover remains a recognisable marker for a book about what happens to data beyond a single function. Publisher and contents.

Here we connect those topics to AI applications and several established references in the discussion of their architecture. This is Lazarus Systems commentary based on public materials, without a review of every chapter of the book.

Decorative AI visualisation. It does not depict the architecture discussed here.

Giving a model tools adds dependencies

In The Shift from Models to Compound AI Systems, published on 18 February 2024, Matei Zaharia and co-authors described systems that combine models with retrieval, tools and further model calls. The term “compound AI” captures something a demo can obscure: the result depends on several components working together. BAIR.

In our example, the model interprets the complaint, search retrieves the policy, a database holds the order and the payment provider issues the refund. Each component has its own state and can stop responding at a different point. Assessing the quality of the generated reply covers only part of that process.

Our practical conclusion is to map a single operation before choosing another model. Identify who owns each write. Where is the authoritative refund record? Which copies exist only to support search? Who reconciles a disagreement? Without those answers, even assigning a failure to the right component becomes difficult.

Compound AI does not justify adding components indefinitely. A separate evaluator model, queue or vector database brings maintenance costs. In our view, each addition should solve a named problem and have a test showing that it improves on a simpler design.

RAG can retrieve an outdated policy

The 2020 paper by Patrick Lewis and co-authors combined text generation with information retrieved from an external index. It is a foundational reference for retrieval-augmented generation, or RAG. The RAG paper.

An enterprise implementation still needs a designed path from source document to answer. Suppose the deadline for submitting complaints changes. The new policy is already in the CMS, but the indexing job has stalled. The assistant retrieves the old version and summarises it accurately. An evaluation that only checks whether the answer follows the supplied passage may accept it.

Treating the index as derived data helps identify the problem. The source document, chunks, embeddings and cached answers have their own update cycles. We suggest keeping the document identifier and version with each chunk, measuring indexing lag and defining what happens when that lag exceeds an acceptable threshold. The application could fetch the source directly, report that a current source is unavailable or refer the case to a person.

Revoking access raises a related issue. Deleting a source file leaves questions about the index, cache and previous conversations. Access controls should apply during retrieval, before data enters the model’s context. Removal procedures should also cover derived copies under the chosen retention policy. A prompt asking the model to keep secrets does not provide that control.

Our assessment of RAG is therefore conditional: it gives a team a way to update knowledge outside model training, but the team must maintain that update path. An acceptance test should include a changed document, a delayed index and revoked access, followed by an inspection of the system’s response to each case.

Refunding after a lost response

Return to the complaint. If the retried request gets a fresh identifier, the payment provider may treat it as another refund. We would design the tool to use a stable idempotency key for the same business operation. Here, idempotency means that retrying that operation does not create another effect. The recipient must recognise the key, and the application must account for the recipient’s retention window and conflict rules.

Writing “completed” to the application database does not automatically cover an external payment. A gap can arise between issuing the refund and recording the result. The application therefore needs a state for an unknown outcome, a way to query the provider by identifier and a process for reconciling the records. Retrying without that information can become another financial operation.

This applies transaction and partial-failure concepts to an AI agent; it does not promise that one pattern will solve every integration. If an external API offers neither deduplication nor outcome lookup, the scope for safe automation is narrower. A team may need to stop ambiguous cases for manual resolution.

Agents, workflows and durable memory

In Building effective agents, published on 19 December 2024, Anthropic distinguishes workflows with predefined code paths from agents whose models choose the next steps. Its authors advise adding complexity when it improves the result. The page now also notes that tooling has changed since publication. Anthropic, 2024.

For complaint handling, we would use fixed rules for authorisation and refund execution. The model could interpret the customer’s account, find missing information and prepare a reply. Freedom to choose the next step should follow from the difficulty of the task. Code can check the refund limit and permission to act regardless of how persuasively a model explains its decision.

The link to data systems becomes more explicit in Anthropic’s description of Managed Agents, published on 8 April 2026. It describes separate interfaces for the session log, the agent’s control loop and the execution environment. A durable log lets work resume after the control process fails. Anthropic, 2026.

We would apply that principle even without using the service: an agent process should be able to disappear without losing the record of completed actions. Conversation history helps a model continue a task. The operation record must also establish whether a tool executed a request. A conversation summary can omit the detail needed to verify that outcome.

Durable storage has costs too. It may contain customer data, tool results and confidential documents. Teams need to define what they record, who can access it and how long it stays. We are not proposing an unrestricted archive of everything.

What to add to AI acceptance criteria

DDIA supplies a vocabulary for state, delays and failures. An AI project also needs semantic evaluation: did the model understand the intent, use the appropriate source and choose a permitted action? A correct database write does not establish that the decision was correct. Both kinds of checks are needed.

Before ending a pilot after a few successful conversations, we propose five trials:

  1. Change a source document and check when the new version appears in answers.
  2. Revoke a user’s access and repeat a question about a previously accessible document.
  3. Break the connection after an operation completes and check that a retry does not duplicate its effect.
  4. Stop the agent process halfway through a task and resume it from durable state.
  5. Compare a simple workflow with an agent on the same tasks, including quality, whole-case cost, completion time and human interventions.

These are proposed Lazarus Systems tests, without claims of measured results. Define the expected behaviour and stopping condition before running each one. Bring the record of one such failure to the architecture review, together with evidence that the system could establish the operation’s state and finish the case safely.

LAZARUS SYSTEMSLet’s talk about your project ↗