Define the unit of work

Start with one business action and write its beginning and end. For an internal knowledge assistant, success might mean a user receives an answer supported by accessible documents, or a clear explanation that the available material does not answer the question. For an agent, success may include a verified update to another system. These are different contracts and should be evaluated separately.

An illustrative purchasing assistant exposes the distinction. Reading an order is a permission. Recommending a correction is an output. Applying that correction is an action with an owner and an audit trail. A fluent conversation can conceal those transitions unless the architecture names them.

ReAct explores interleaving reasoning and actions in language-model tasks. That makes the boundary especially visible: a system can gather information and act across several steps. The paper is a research method, not an operating contract for your business. The application still has to define permitted actions and acceptable outcomes.

Draw the dependencies before choosing the model

List the sources of truth, the identity checks, the tools and the failure destinations. Identify what happens when a dependency is slow, empty, stale or unavailable. The model may still answer when a retrieval service fails; the application needs to decide whether that answer should be shown at all.

Sculley and colleagues describe how machine learning systems accumulate technical debt through data dependencies, feedback loops and the surrounding infrastructure. Their paper supports looking beyond the isolated model. In an architecture review, we turn that concern into named owners and visible interfaces: who can change the input, who notices a change and which downstream decisions depend on it.

Build one complete path

A first release should include the user, the evidence and the failure route. An attractive front end attached to an unobserved model call is not enough to learn whether the workflow holds together. Use a narrow scope so the team can afford to build the operational pieces early.

The ML Test Score treats production readiness as a collection of checks around data, models, infrastructure and monitoring. Its useful lesson for a design review is the breadth of the question. A model comparison cannot stand in for checking the service that will contain it.

For the purchasing example, a useful pilot could retrieve the order, compare a small set of fields and present the evidence in a review queue. It would record the run identifier, relevant versions, reviewer outcome and any failed step. It could avoid modifying the accounting system entirely until the review process is understood.

This is also where the team should agree what it will measure. A correct recommendation that arrives after the decision has been made is not useful. A fast recommendation that requires a long investigation may cost more than the previous process. Quality, latency, review effort and operating cost belong in the same release discussion.

A short failure walkthrough

Draw six boxes on a page: request, identity, source, model, action, operator. Start at each box in turn and remove something the next box expects. The exercise is deliberately small. A source returns an old revision. A user loses access while a run is queued. The destination accepts a change and the response disappears.

For each interruption, write what the user sees and what the operator knows. If those answers are different, make the gap intentional. A user may see a short message while an authorised operator sees a run identifier and a safe reconciliation path. Neither should be told that a business action failed merely because a network request timed out.

Design the day after launch

Name the person who can pause the workflow and the person who can change it. Document how work continues when the AI service is unavailable. Decide what run records to retain, who can inspect them and how corrections enter the next evaluation.

The release readout should say what was tested, which cases remain unsupported and what would cause the team to roll back. A limited but explicit operating boundary gives the business something it can reason about. Expansion then becomes a decision based on use, rather than a consequence of the software being technically available.

Before adding another feature, ask someone outside the implementation to follow that record. Can they identify the source revision, the decision and the owner without a private explanation? If they cannot, the next increment may be an operating surface rather than a more capable model.

QuestionAn answer the team can inspect
What is allowed?A permission and approval map
What counts as working?Representative cases and acceptance criteria
What happened?A run record tied to source and version
Who takes over?A named operator and a recovery procedure

Sources & further reading

  1. Sculley et al. — Hidden Technical Debt in Machine Learning Systems
  2. Breck et al. — The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction
  3. Yao et al. — ReAct: Synergizing Reasoning and Acting in Language Models

Primary sources inform the technical background. The examples and proposed working methods are MGLO’s own; Business scenarios are identified in the text.