Picture a customer asking where an order is. The business has a planned date in one system, a supplier update in someone's inbox, a check recorded on a paper form and a coordinator who knows that the customer changed the specification yesterday. An AI agent receives the question and produces a fluent answer using one of those sources. The wording may sound confident while the answer is wrong.

The problem is not solved by adding more prompt instructions. The agent needs to know which order it is handling, which record is authoritative, whether the latest update has been reviewed, what it is allowed to tell the customer and when it should wait for a person. Those are properties of the business workflow and its information system.

Be clear about what you mean by an agent

People use 'AI agent' to describe a range of tools. A conversational assistant may answer a question or draft text. A conventional automation follows explicit rules such as 'when a job is marked ready, notify the planner'. An agent can interpret a task, use information or tools and choose a next action within the permissions it has been given. The boundaries vary by design, so define what the system will actually do before planning it.

Some tasks do not require an agent. A rule-based workflow is often easier to understand when the inputs and conditions are predictable. A text assistant can help someone draft or summarise while that person remains responsible for the record. A more autonomous agent may make sense when the task involves several steps of interpretation and a person can review exceptions. Choose the simplest approach that handles the work reliably.

The word 'agent' should never conceal the operating detail. Say which event starts the task, what information it can use, what tools it can call, whether it can change a record or contact someone, and what happens after it finishes. If those statements are difficult to make, the workflow probably needs more definition before any model is chosen.

Start with a workflow the business already understands

Choose a small, repeatable piece of work with an owner and a result that matters. Examples might include checking an incoming request for missing information, preparing a draft response from an approved job record, classifying a document for staff review or flagging a job whose status conflicts with its plan. Avoid starting with broad instructions such as 'manage production', 'run HR' or 'handle customer service'. Those describe whole areas of responsibility, not a task with a clear boundary.

Walk through several actual examples with the team. Include a routine case, a delayed case, an incomplete record and an exception where someone had to make a judgement. For each example, note what the person saw, what they decided, which rule applied, what they recorded and who needed to know. This is how the team reveals tacit knowledge that a prompt cannot safely invent.

A good first task has four properties

1. The input is identifiable

The agent can resolve which customer, order, job or request it is handling and retrieve the current record. Identifiers should be unambiguous, and the system should surface relevant changes rather than relying on a name copied into a message. If the record cannot be found or two sources disagree, the task should stop and route the mismatch to a person.

2. The action has a defined limit

'Help with the order' is open-ended. 'Check that the order has a delivery date and draft a message asking for the missing date' is bounded. State whether the agent may read, draft, update, approve or send. A permission to prepare a change should not automatically include permission to commit it. Give each task the least authority it needs to do its job.

3. The result can be checked

A person or system should be able to see whether the task was completed and what evidence supports the result. A vague answer such as 'the job looks fine' is difficult to review. A useful result might identify the record checked, the fields that matched, the discrepancy found and the proposed next step. That makes review faster and helps the business understand why the agent acted.

4. Uncertainty has somewhere to go

A safe workflow defines how the agent behaves when the record is missing, the instruction conflicts with a rule, the confidence is low or the action could affect a customer or quality decision. It can stop, create a review task, ask a responsible person or provide a draft without sending it. A visible handoff is better than quietly guessing or failing in a way no one notices.

Build six foundations around the task

A current record

Give the agent access to the record the team already treats as authoritative, or establish a shared record if none exists. Define which fields matter and who is allowed to change them. Include the status, dates and notes the task needs; exclude unrelated personal or commercially sensitive information. If the agent needs to read messages or documents, state how their contents are tied back to the right job and what is retained.

A named owner

Every task should belong to a role in the operation. Someone decides what the task means, approves changes to its rules, reviews exceptions and investigates poor results. The owner does not have to inspect every routine result forever, but the business needs a person with the authority to change or stop the workflow.

Bounded permissions

Separate access to read information from permission to change it or communicate externally. Start with the narrowest access that makes the task useful. Use role-based access and protect credentials. Keep high-impact actions behind an approval until the business has evidence that the proposed action is dependable and the customer or quality risk is understood.

A human review point

Decide which cases can proceed automatically, which need sampling and which always need a named reviewer. A person should see enough context to judge the recommendation rather than simply click approve. Record the decision and any correction. Review is a designed part of the workflow, not a final button bolted onto an unpredictable process.

An exception route

Write down what happens when an agent cannot find a record, encounters conflicting dates, sees an unfamiliar document, receives an unusual request or reaches a tool error. Set a clear destination, response time and fallback owner. Make stalled tasks visible. If an exception just disappears into a shared inbox, the business has automated the creation of another queue rather than resolved it.

A record of what happened

Keep an appropriate record of the task, the information used, the action proposed or taken, the human decision and any error. The level of detail depends on the process and the data involved. A useful history lets a manager answer 'what changed and why?' without asking the agent to reconstruct the past from memory.

A worked example: checking a job update

Imagine a team receives a progress update from the shop floor. The process owner wants the system to compare the update with the current job plan, flag a likely delay and prepare an internal note. The job record contains the agreed due date, current stage, assigned owner and latest approved change. The agent can read those fields and draft a note. It cannot alter the promise to the customer or mark a quality check as passed.

If the update matches the plan, the system records that the check ran and leaves the next step with the named owner. If the update suggests the job may be late, it creates a review task with the comparison and source update attached. If the job record is missing its due date, the agent routes it as a data issue. The team can see what happened, and the agent's role remains narrow.

This example is an illustration of an agent workflow, not a claim that every project or client system already uses AI. The important design is the connection between the task, the real job record, the action limit and the person who owns an exception. Those foundations also make conventional automation and manual work clearer.

Choose between a rule, an assistant and an agent

Match the tool to the work
PatternGood fitExample boundary
Rule-based automationInputs and conditions are stable and explicit.Notify the owner when a due date passes.
AI assistanceA person benefits from a draft, summary or classification.Prepare a customer update for a named employee to review.
Agent workflowThe task involves bounded interpretation and several permitted steps.Compare a progress note with the job plan, then route a mismatch.
Human decisionThe choice depends on context, accountability or material judgement.Approve a revised customer commitment or quality exception.

The categories can sit in one workflow. A rule may trigger the work, an AI assistant may prepare context, an agent may complete a bounded check and a person may approve the consequential decision. Avoid choosing an agent because it sounds more advanced. Choose the pattern that makes work clearer, faster or more dependable at an acceptable level of risk.

Test the whole workflow, including its failures

Before launch, build a set of real examples from the workflow owner. Include normal inputs, incomplete records, duplicate requests, changed instructions, conflicting dates, unusual wording and tool failures. Ask experienced staff to show how they currently decide each case. Use these examples to check whether the agent retrieves the right record, follows the right rule, produces a reviewable result and escalates the cases that need judgement.

Test against the action, not just the answer's fluency. A draft can sound professional and still quote a stale date. A classification can look plausible and still route a quality issue to the wrong team. Review correctness, omissions, false alarms, missed exceptions and whether the human reviewer has enough context. Keep examples that failed; they reveal what the workflow and instructions need to handle.

Decide what level of performance is acceptable before switching the workflow on. A task that drafts internal notes may tolerate a different error rate from one that changes a delivery promise or updates a record used for a regulated check. Where errors have larger consequences, keep stronger review, narrower permissions and more extensive test cases. The process owner should sign off on the boundaries, not just the model behaviour.

Measure value as well as accuracy

Record the current process before introducing an agent. How long does the task take? How many requests need rework? How often does the team chase for missing information? How long do exceptions remain unresolved? Choose measures that reflect the business outcome, then compare them after people have used the new process. If a task becomes faster but creates extra review work or more customer corrections, the system has moved effort rather than removed it.

  • Time from task arriving to a usable result.
  • Share of routine cases completed without correction.
  • Number and type of cases sent for human review.
  • Missed exceptions, incorrect changes and customer-impacting errors.
  • Staff time spent reviewing, correcting and supporting the workflow.
  • Whether people use the system or return to private workarounds.

Review these measures with the people doing the work. A good result includes a process they can understand and trust. If the agent is technically active but staff do not use it, learn why. Its information may be stale, the review may take longer than the original task or the workflow may not match what happens in practice.

Roll it out in a controlled sequence

First, observe without taking action

Run the workflow against real examples without changing records or sending messages. Compare the output with what the responsible employee did. This reveals gaps in context and evaluation while the existing process remains in control.

Then, prepare work for a person

Let the agent create drafts, classifications or review tasks. The team decides whether to accept or correct them. Capture why corrections happen and improve the rules or source data before expanding the agent's authority.

Automate only the actions the evidence supports

Once the workflow performs consistently on the cases that matter, consider allowing narrow low-consequence actions to happen automatically. Keep explicit review for decisions that affect customers, money, quality or safety. Define how the team can pause the workflow and return to the manual route.

Protect the business information around the agent

Treat an agent as a user of business systems with a particular purpose and access. Decide which data it genuinely requires, where the data is processed, what is retained and which people can review the output. Avoid placing credentials, customer details or confidential documents in an unapproved tool. Follow the same access, retention and supplier checks the business applies to other software.

Prompt instructions alone are not an access-control system. Enforce important permissions in the application or connected tools. Separate test data from live data where possible, keep a person able to stop material actions and watch for unexpected activity. Document who approved the workflow and how the business can disable or change it if the process, provider or risk changes.

A readiness check before you build

  • Can you name one task, its owner and the outcome it should produce?
  • Can the system identify the current record and the information the task needs?
  • Are read, draft, update and send permissions clearly separated?
  • Do you know which cases must be reviewed or escalated?
  • Can a manager see what the agent did and why?
  • Have you recorded a baseline and agreed what a useful result looks like?
  • Can the business pause the workflow and continue its important work?

A 'no' is useful. It tells you what foundation to put in place first, whether that is cleaning up records, agreeing an owner, making an exception route visible or connecting two systems. Build those pieces as part of the operation instead of expecting an agent to infer them.

Give the agent a job the business can manage

The strongest starting point is usually small and specific: one workflow, one accountable owner, a clear source record and a result the team can check. Build the information, permissions and review path around it. Keep the agent's actions proportionate to the evidence and expand only when the business can see what it is doing and how the work is improving.

If your team is exploring AI, bring one process that is slow, difficult to track or repetitive enough to invite automation. Map who owns the work, where its records live and what happens when a case is unusual. That conversation will reveal whether the next step is an agent, a conventional automation, a better-connected system or simply a clearer process.

How a build runs, stage by stage

Back to the notes