Capability map

Capability 07 · Technology

AI & Agentic Workflows

Design the authority before giving the agent tools.

I design agents that reason over business context, use bounded tools, pause when ambiguity becomes consequential and make every action explainable and auditable.

Request → context → reasoning → tool → draft → review → commit → auditOne request, from an unstructured email to a committed CRM record. Select a step to see what the agent reads, which tool it may use, how much authority it has—and what it leaves in the audit trail.

  1. Read & reason
  2. Draft · reversible
  3. Human authority
  4. Record
Step 04 · Read & reason

Identify products

Matches two lines to the catalogue; one reference could be two products.

Reads
Catalogue, cross-references, this customer’s order history
Tool
matchProducts
Authority
Recommend
Leaves in the audit trail
Per-line match, alternatives, ambiguity flag
An incoming RFQ email becomes a CRM record. The agent reads and drafts; a person holds the commit; the system writes and records—and the trail explains every step.
01What agentic workflows meansNot a chatbot

Not “add a chatbot”. A bounded reasoning component in a real process.

I treat AI as one participant in an operating process: it reads the context it is given, uses the tools it is allowed, stops where consequence begins and leaves a trail a person can follow.

It is not
  • Adding a chatbot beside the CRM
  • Giving a model unrestricted API access
  • Using a confidence score as permission
  • Assuming AI can repair fragmented data
  • Automating every decision
  • Letting prompts define governance
  • Treating a demo as production architecture
That produces impressive demos and ungoverned behaviour.
It is
  1. Business taskThe real work the agent helps complete—and what “done” means.
  2. ContextRelevant, trusted and timely context, assembled—not dumped.
  3. Tool accessOnly the actions the task needs, each behind a contract.
  4. AuthorityRead, recommend, draft, execute, approval or forbidden—per action.
  5. AmbiguityRetrieve, ask, escalate or stop—decided before it happens.
  6. ReviewHuman judgement placed where consequences are real.
  7. EvidenceEvery recommendation carries its sources and its uncertainty.
  8. EvaluationScored per dimension on a fixed set—never one accuracy number.
  9. FallbackA defined path when the agent, a tool or a model fails.
  10. OutcomesMeasured as process results, owned by the business.
Four principles, applied below
  1. 01Confidence is not authority.

    A score says how sure the model is—not what it may do.

    Confidence vs consequence
  2. 02The agent should know what it is allowed to do before it decides what it wants to do.

    Authority is designed per action, before tools exist.

    Authority model
  3. 03AI cannot fix unreliable identity.

    It can rank candidates; a trusted key or a person confirms them.

    Context model
  4. 04Human-in-the-loop is a design decision, not a fallback for bad AI.

    Review sits where consequence sits, even at high confidence.

    Human-in-the-loop
02My agentic design approachTen steps · prompts come later

From business task to safe improvement.

Ten moves from the work the agent should help complete to the loop that improves it without silently changing its behaviour. Prompts come after the fourth.

01Start from the business task

Step 01 · The question

What real work should the agent help complete?

A task with a definition of done—not “an assistant”.

For example
  • Classify an inbound request
  • Identify the customer
  • Summarize account context
  • Prepare a quote draft
  • Propose the next action
  • Triage a service request
ProducesTask definition

Second-orderWithout a task, there is nothing to evaluate—only demos

  1. Start from the business task
    Step 01 · The question

    What real work should the agent help complete?

    A task with a definition of done—not “an assistant”.

    For example
    • Classify an inbound request
    • Identify the customer
    • Summarize account context
    • Prepare a quote draft
    • Propose the next action
    • Triage a service request
    ProducesTask definition

    Second-orderWithout a task, there is nothing to evaluate—only demos

  2. Define the decision surface
    Step 02 · The question

    What information does the agent need to reason correctly?

    Relevant, trusted, timely—assembled, not dumped.

    Define
    • Structured system data
    • Unstructured text
    • Historical context
    • Policies
    • Prior actions
    • User intent
    ProducesContext model

    Second-orderMore context is not always better context

  3. Define authority
    Step 03 · The question

    What may the agent do?

    Per action, before any tool exists.

    Classify each action
    • Read
    • Recommend
    • Draft
    • Execute
    • Require approval
    • Forbidden
    ProducesAuthority matrix

    Second-orderThe agent should know what it is allowed to do before it decides what it wants to do

  4. Define tools
    Step 04 · The question

    Which system actions are exposed to the agent?

    Each tool is a contract, not an endpoint.

    For example
    • Search account
    • Create draft opportunity
    • Retrieve product
    • Update classification
    • Send notification
    • Request approval
    ProducesTool contract

    Second-orderAn agent cannot misuse a capability it was never given

  5. Define confidence and ambiguity behaviour
    Step 05 · The question

    What should happen when the agent is unsure?

    Decided per action and consequence.

    Choose between
    • Continue
    • Ask the user
    • Retrieve more context
    • Escalate
    • Stop
    ProducesAmbiguity policy

    Second-orderConfidence is not authority

  6. Define human review
    Step 06 · The question

    Where does human judgement become mandatory?

    Where consequence is, not where the model is weak.

    Define
    • Consequential decisions
    • Irreversible actions
    • Financial commitments
    • Customer communication
    • Low-confidence matches
    ProducesReview model

    Second-orderHuman-in-the-loop is a design decision, not a fallback for bad AI

  7. Design execution
    Step 07 · The question

    How does the agent move through the workflow?

    A stateful run with stopping points.

    Define
    • Plan
    • Tool use
    • Intermediate state
    • Retries
    • Stopping points
    • Escalation
    ProducesAgent workflow

    Second-orderA good agent knows when to stop

  8. Define evidence and audit
    Step 08 · The question

    Can a person understand why the agent acted?

    The trail is designed, not logged by accident.

    Store
    • Input
    • Context
    • Tools used
    • Recommendation
    • Evidence
    • Human decision
    • Final action
    ProducesAudit model

    Second-orderA recommendation without evidence asks reviewers to trust, not to judge

  9. Evaluate
    Step 09 · The question

    How do we know the agent is good enough?

    Scored per dimension on a fixed set—never one number.

    Define
    • Evaluation set
    • Expected outcome
    • Extraction accuracy
    • Entity matching
    • Action correctness
    • Escalation quality
    • False autonomy
    • Latency and cost
    ProducesEvaluation framework

    Second-orderOne aggregate accuracy hides the one failure that matters

  10. Improve safely
    Step 10 · The question

    How does the agent learn without silently changing behaviour?

    Every change is versioned and regression-tested.

    Define
    • Failure review
    • Prompt and policy versioning
    • Evaluation regression
    • Human corrections
    • Controlled release
    ProducesImprovement loop

    Second-orderA prompt edit is a release

03Typical problemsSymptoms, and what is usually missing

The symptoms arrive as “the AI got it wrong”.

Usually it did what nothing prevented it from doing. Underneath, a task, an identity rule, an authority level or an evaluation set was never designed.

  • Requests arrive as email and attachments; people retype them into the CRM.

    Usually missingA defined intake task with an extraction contract
  • The agent picks a customer when two accounts look the same.

    Usually missingAn identity rule, and a stop when it is not met
  • Low-confidence output is committed because nothing said otherwise.

    Usually missingAn ambiguity policy tied to consequence
  • The agent’s credentials can do far more than its task needs.

    Usually missingAn authority matrix and least-privilege tools
  • Reviewers approve AI output without seeing why it was proposed.

    Usually missingA review card with evidence and uncertainty
  • Business policy lives in prompts nobody versions.

    Usually missingPolicies as data, prompts as versioned configuration
  • Quality is judged from a good demo and a few anecdotes.

    Usually missingAn evaluation set scored per dimension
  • Nobody can reconstruct why the agent acted.

    Usually missingAn audit model and an execution trace
04What I designTangible artefacts

Artefacts that make an agent safe to put in a real process.

Each one answers a question a demo never has to.

BoundariesWhat the agent may know and do
  • Agent workflowWhich steps, which actors, where it stops?
  • Context architectureWhat the agent knows, from which source, how fresh?
  • Authority matrixRead, recommend, draft, execute, approve or forbid—per action
  • Tool catalog & contractsInput, output, preconditions, side effect, reversibility
  • Guardrail modelWhat is checked before and after every action?
JudgementWhere people decide
  • Human review modelWhat the reviewer sees, and what each decision does
  • Ambiguity policyRetrieve, ask, escalate or stop—when?
  • Confidence policyHow confidence and consequence combine
  • Evidence modelWhich sources support each recommendation?
  • Fallback & escalation modelWhat happens when the agent or a tool fails?
AssuranceHow we know it works—and keeps working
  • Evaluation setScored cases per dimension, not one accuracy number
  • Audit trailInput, context, tools, evidence, decision, action
  • Observability modelWhich trace shows what the agent did, step by step?
  • Improvement loopHow corrections become evaluation cases
  • Release governanceVersioned prompts and policies, gated by regression
05Context modelRetrieve · filter · structure · reason

More context is not always better context.

An agent reasons over what it is given. Six layers of context, read two ways: dumped into one prompt, or assembled from the systems that own each fact.

What the agent knows
LayerRetrieveFilterStructure
User input“Prepare this RFQ for the account owner”The instruction and who gave itOnly what the user may seeTask, requester, permissions
Unstructured contentEmail, attachment, message threadThis message and its attachmentSignatures, disclaimers and old replies removedLines: reference, quantity, date
Structured business dataCRM, ERP, data platformAccounts for the sender’s domain; open opportunitiesActive, legally valid accounts onlyCandidates with keys and status
Policy & rulesRFQ definition, substitution ruleThe versioned policies for this taskOnly policies in force todayRules the tools enforce, not prose
Historical contextPast orders, prior RFQsThis customer’s last 12 months of linesSame product family onlyFrequencies per reference
Tool resultsSearch, match, availabilityResults of the tools called in this runLatest call per questionTyped results with timestamps
Reason

The agent reasons over typed, sourced, current context—each item with its key and timestamp.

  • Relevant

    Needed for this task and this step—not everything the agent could read.

  • Trusted source

    From the system that owns the fact, with its key.

  • Timely state

    Fresh enough for the decision, with its timestamp.

AI cannot fix unreliable identity. It can rank candidates. A trusted key—or a person—confirms them.

06Agent authority modelRead → forbidden, per action

The agent should know what it is allowed to do before it decides what it wants to do.

Every action the agent could take has an authority level set by consequence and reversibility—not by how capable the model is. Select an action to see why.

Agent actions and the authority each one is given (example pattern)
ActionReadRecommendDraftExecuteHuman approvalForbidden
Execute
Read
Execute
Execute
Recommend
Human approval
Human approval
Human approval
Forbidden
Human approval
Selected action

Submit quote

Human approval

A commercial commitment in the customer’s hands.

Reversible
No
Consequence
High

An example pattern, not a universal rule: authority is set per organization, per action.

07Tool boundaryContracts, not endpoints

A tool is an architecture contract, not an API endpoint.

The agent can only do what its tools allow—so each tool states its purpose, input, output, permissions, preconditions, side effect, failure and reversibility, and enforces them in code.

Tool catalog
Deliberately not tools
  • approveDiscount

    A commercial decision right. The agent can see the policy; it cannot exercise it.

  • createErpCustomer

    Legal and financial creation runs through the orchestrated process with finance validation.

  • mergeAccounts

    Master-data authority stays with data stewards; the agent may only suggest candidates.

Tool contract

createOpportunityDraft

Agent authority: ExecuteA person commits later—in review

Purpose
Prepare a reviewable draft from confirmed data
Input
Account key, product lines, source request ID
Output
Draft ID and every value with its source
Permissions
Create drafts only; no stage, owner or amount beyond the draft
Preconditions
Account confirmed or flagged for review; request not already drafted
Side effect
Creates a draft only—invisible to pipeline, forecast and customer
Failure
Duplicate request ID → return the existing draft; validation error → stop with reason
Reversibility
Fully reversible: drafts expire if not approved

Agent workflows are stateful processes

The same run model as any orchestration: states, valid transitions, who moves each one—agent, system, person, timer—and terminal outcomes. Select a state.

Run path
Failure
Recovery
Run path

Waiting review

A person holds the decision; the draft and evidence are ready.

Can move to
  • CommittedApproved (possibly edited)Person
  • RejectedRejected with a reason codePerson
  • Context readyReviewer requests more contextPerson

Entered from Reasoning, Escalated.

08Human-in-the-loopApprove · edit · reject · more context

Human-in-the-loop is a design decision, not a fallback for bad AI.

The reviewer sees what the agent believes, why, the source evidence, the uncertainty and exactly what approval will do—and every decision is recorded with a reason.

The agent prepared · synthetic

RFQ from a distributor · three lines

What it believesThis is an RFQ from an existing distributor for two known products and one probable substitute.

Proposed values with confidence and evidence
ValueConfidenceEvidence
CustomerDistributor account · regional entityHighSender domain and signature match the account; four prior RFQs from this contact
Line 1Product A-100 × 200HighExact catalogue reference in the attachment
Line 2Product B-220 × 50HighCross-reference to a competitor code, bought twice last year
Line 3Product C-310X × 20Low“C310” matches two products; the X variant was ordered once

If approvedCreate an opportunity at Qualified for the account owner, with three lines; line 3 as C-310 unless edited. No email is sent.

The reviewer decides
Reason code
Wrong product variant
What happens
Line 3 becomes C-310X; the edited draft is committed.
What it teaches
The correction becomes a labelled evaluation case for product matching.
  • Accountability

    A named person made the consequential decision, on recorded evidence.

  • Learning data

    Every edit and rejection, with its reason code, is a labelled example.

  • Auditability

    Anyone can see what the agent proposed and what the human changed.

09Confidence vs consequenceTwo axes, four behaviours

Confidence is not authority.

Confidence says how sure the model is. Consequence says what happens if it is wrong. Pick an action and move its confidence: in the high-consequence row, no score reaches “execute”.

Pick an action
Stop · escalate

Unsure and consequential: stop and hand over everything gathered.

Human approval

Sure, but consequential: prepare it perfectly—a person commits.

Retrieve · ask · escalate

Cheap to get more context. Retrieve, ask the requester, or hand to a person.

Execute

Low consequence and high confidence: act, visibly and reversibly.

Send the quote to the customer · 97 %Human approval High confidence does not grant authority: a person commits.

10Evaluation & observabilityPer dimension · inspectable runs

A good demo is not an evaluation.

Agent quality is scored on a fixed set, per dimension, against its own target—never as one “AI accuracy” number. And every run can be read step by step.

260 scored cases · synthetic results
  1. ClassificationDid it identify the request correctly?97 %target 95 %0
  2. ExtractionDid it capture the required data?95 %target 90 %+4
  3. Entity matchingDid it identify the right account and products?93 %target 90 %+4
  4. ReasoningDid it choose the right next step?93 %target 90 %+1
  5. Tool useDid it call the right tool with the right arguments?96 %target 95 %-1
  6. AuthorityDid it stop when it had to?98 %target 100 % · gate-2
  7. EvidenceDid it support its recommendation?96 %target 95 %0
  8. OutcomeDid the process end as expected?88 %target 85 %+2

Blocked. Better extraction and matching do not compensate for a single missed stop: the authority gate is 100 %.

One run, as an execution trace

Time, actor, action, evidence and outcome for every step. AI operations should be inspectable.

  1. SystemRequest receivedEvidence: Message from a distributor domain, one attachmentOutcome: Run R-7F3 created
  2. AgentContext retrievedEvidence: Account candidates, 12 months of orders, RFQ policy v4Outcome: Context ready
  3. AgentAccount matchedEvidence: Domain + signature + four prior RFQsOutcome: One candidate, high confidence
  4. AgentProduct 1 matchedEvidence: Exact catalogue referenceOutcome: A-100 × 200
  5. AgentProduct 2 ambiguousEvidence: “C310” matches C-310 and C-310XOutcome: Flagged; not resolved by the agent
  6. AgentHuman review requestedEvidence: Draft, evidence and one open questionOutcome: Task for the account owner
  7. HumanProduct correctedEvidence: Reviewer chose C-310X · reason: wrong variantOutcome: Correction stored as evaluation case
  8. SystemDraft committedEvidence: Approval reference on the recordOutcome: Opportunity created
11Reference casesFictional & composite

Five agents, each with a boundary.

Synthetic B2B scenarios. The RFQ intake agent is the flagship—the agent design behind the existing lab, reference architecture and process case, which it links to rather than repeats.

  1. Case 01Selected workAuthority, tools & human commit

    Designing an AI-assisted RFQ intake agent

    RFQs arrive as free-text email and attachments, and an agent is expected to “handle them”—with nobody having said what it may do.

    Architecture questionWhen is a customer match trusted, what happens with an ambiguous product—and what may the agent never do, however confident it is?

    human stopEmailClassMatchLinesDraftReviewCommit
    MailboxAI agentCRMProduct catalogueConnects toProcess ArchitectureSystems & CRM Architecture
  2. Case 02Context assembly, fact vs inference

    Preparing account context before a commercial visit

    Salespeople prepare visits by opening six screens—or not at all—and an AI summary is proposed that may invent what it cannot find.

    Architecture questionWhich sources are authoritative, how fresh must they be—and how does the briefing separate fact from inference?

    human stopVisitContextSignalsSummaryFocusReview
    CRMERPData platformCalendarConnects toRevenue OperationsData & Integrations
  3. Case 03Taxonomy, routing & mixed intent

    Turning commercial email into structured work

    A shared commercial inbox mixes RFQs, pricing questions, complaints, technical questions, leads and order questions—and response time depends on who looks first.

    Architecture questionWhich classes are safe to route automatically, what happens with mixed intent—and may the agent create tasks or send replies?

    human stopEmailClassifyEntitiesOwnerDraftRoute
    MailboxAI agentCRMService deskConnects toProcess ArchitectureAutomation
  4. Case 04Evidence-backed summary, manager decides

    AI-assisted account review with evidence

    Monthly account reviews start with managers assembling numbers; the discussion that follows has little time and no shared evidence.

    Architecture questionWhat may the agent prepare for a performance review—and where must it stop short of commercial judgement?

    human stopDataSignalsContextSummaryReviewAction
    Data platformCRMPlanningConnects toRevenue OperationsData & Integrations
  5. Case 05Suggest with evidence, steward decides

    Using AI to assist data quality without giving it master-data authority

    The CRM has probable duplicate accounts, inconsistent names and missing classifications—and a proposal to let a model fix them.

    Architecture questionWhere does AI help data quality—and why must merge, rename and reclassification stay with a person?

    human stopScanSuggestEvidenceQueueStewardApply
    CRMERPData platformSteward queueConnects toData & IntegrationsSystems & CRM Architecture
12Questions that change the designThe second layer

Every AI request has a second layer.

The visible request is where the design starts, not where it ends.

Start from a request
Visible requirement
“Add an AI assistant to the CRM.”
First layer

A chat panel that can answer questions about any record.

That is an interface. It is not yet a workflow.

4 of 5 questions surfaced

Second layer — what the request leaves unsaid
  1. “Answer anything” has no definition of done and cannot be evaluated.

13Connected capabilitiesWhere agents sit

Process defines the task. Systems hold the truth. The agent works in between.

AI & Agentic Workflows

  1. Process Architecture

    The process defines the task, the decision rights and the review points the agent works within.

  2. Systems & CRM Architecture

    Systems of record own identity and state; the agent reads them and writes only through bounded tools.

  3. Data & Integrations

    Context is only as good as its source, key and freshness—AI cannot fix unreliable identity.

  4. Automation

    Agent workflows are stateful runs: the same states, guards, retries and recovery as any orchestration.

  5. Revenue Operations

    Account review and visit preparation turn governed signals into evidence-backed briefings.

  6. Digital Operating Model

    Someone owns the agent after go-live: its policies, evaluation and release.

Next capability · 08Revenue Operations

An agentic workflow is not a model with access. It is a designed process with

  1. A task
  2. Context
  3. Authority
  4. Tools
  5. Ambiguity rules
  6. Review
  7. Evidence
  8. Evaluation
  9. Release control

A good agent knows when to stop.

Design the authority before giving the agent tools.

Start with the boundary

An agent that demos well—and nobody has decided what it may do?

  • Which of its actions are reversible?
  • What does it do when two customers look the same?
  • How would you know it got worse after a prompt change?
Start a conversation