AI & Agentic Workflows
Design the authority before giving the agent tools.
I design agents that reason over business context, use bounded tools, pause when ambiguity becomes consequential and make every action explainable and auditable.
One request, from an unstructured email to a committed CRM record. Select a step to see what the agent reads, which tool it may use, how much authority it has—and what it leaves in the audit trail.
- Read & reason
- Draft · reversible
- Human authority
- Record
Identify products
Matches two lines to the catalogue; one reference could be two products.
- Reads
- Catalogue, cross-references, this customer’s order history
- Tool
matchProducts- Authority
- Recommend
- Leaves in the audit trail
- Per-line match, alternatives, ambiguity flag
Not “add a chatbot”. A bounded reasoning component in a real process.
I treat AI as one participant in an operating process: it reads the context it is given, uses the tools it is allowed, stops where consequence begins and leaves a trail a person can follow.
Adding a chatbot beside the CRMGiving a model unrestricted API accessUsing a confidence score as permissionAssuming AI can repair fragmented dataAutomating every decisionLetting prompts define governanceTreating a demo as production architecture
- Business taskThe real work the agent helps complete—and what “done” means.
- ContextRelevant, trusted and timely context, assembled—not dumped.
- Tool accessOnly the actions the task needs, each behind a contract.
- AuthorityRead, recommend, draft, execute, approval or forbidden—per action.
- AmbiguityRetrieve, ask, escalate or stop—decided before it happens.
- ReviewHuman judgement placed where consequences are real.
- EvidenceEvery recommendation carries its sources and its uncertainty.
- EvaluationScored per dimension on a fixed set—never one accuracy number.
- FallbackA defined path when the agent, a tool or a model fails.
- OutcomesMeasured as process results, owned by the business.
- 01Confidence is not authority.
A score says how sure the model is—not what it may do.
Confidence vs consequence - 02The agent should know what it is allowed to do before it decides what it wants to do.
Authority is designed per action, before tools exist.
Authority model - 03AI cannot fix unreliable identity.
It can rank candidates; a trusted key or a person confirms them.
Context model - 04Human-in-the-loop is a design decision, not a fallback for bad AI.
Review sits where consequence sits, even at high confidence.
Human-in-the-loop
From business task to safe improvement.
Ten moves from the work the agent should help complete to the loop that improves it without silently changing its behaviour. Prompts come after the fourth.
01Start from the business task
What real work should the agent help complete?
A task with a definition of done—not “an assistant”.
- Classify an inbound request
- Identify the customer
- Summarize account context
- Prepare a quote draft
- Propose the next action
- Triage a service request
Without a task, there is nothing to evaluate—only demos
Start from the business task
What real work should the agent help complete?
A task with a definition of done—not “an assistant”.
- Classify an inbound request
- Identify the customer
- Summarize account context
- Prepare a quote draft
- Propose the next action
- Triage a service request
Task definitionWithout a task, there is nothing to evaluate—only demos
Define the decision surface
What information does the agent need to reason correctly?
Relevant, trusted, timely—assembled, not dumped.
- Structured system data
- Unstructured text
- Historical context
- Policies
- Prior actions
- User intent
Context modelMore context is not always better context
Define authority
What may the agent do?
Per action, before any tool exists.
- Read
- Recommend
- Draft
- Execute
- Require approval
- Forbidden
Authority matrixThe agent should know what it is allowed to do before it decides what it wants to do
Define tools
Which system actions are exposed to the agent?
Each tool is a contract, not an endpoint.
- Search account
- Create draft opportunity
- Retrieve product
- Update classification
- Send notification
- Request approval
Tool contractAn agent cannot misuse a capability it was never given
Define confidence and ambiguity behaviour
What should happen when the agent is unsure?
Decided per action and consequence.
- Continue
- Ask the user
- Retrieve more context
- Escalate
- Stop
Ambiguity policyConfidence is not authority
Define human review
Where does human judgement become mandatory?
Where consequence is, not where the model is weak.
- Consequential decisions
- Irreversible actions
- Financial commitments
- Customer communication
- Low-confidence matches
Review modelHuman-in-the-loop is a design decision, not a fallback for bad AI
Design execution
How does the agent move through the workflow?
A stateful run with stopping points.
- Plan
- Tool use
- Intermediate state
- Retries
- Stopping points
- Escalation
Agent workflowA good agent knows when to stop
Define evidence and audit
Can a person understand why the agent acted?
The trail is designed, not logged by accident.
- Input
- Context
- Tools used
- Recommendation
- Evidence
- Human decision
- Final action
Audit modelA recommendation without evidence asks reviewers to trust, not to judge
Evaluate
How do we know the agent is good enough?
Scored per dimension on a fixed set—never one number.
- Evaluation set
- Expected outcome
- Extraction accuracy
- Entity matching
- Action correctness
- Escalation quality
- False autonomy
- Latency and cost
Evaluation frameworkOne aggregate accuracy hides the one failure that matters
Improve safely
How does the agent learn without silently changing behaviour?
Every change is versioned and regression-tested.
- Failure review
- Prompt and policy versioning
- Evaluation regression
- Human corrections
- Controlled release
Improvement loopA prompt edit is a release
The symptoms arrive as “the AI got it wrong”.
Usually it did what nothing prevented it from doing. Underneath, a task, an identity rule, an authority level or an evaluation set was never designed.
Requests arrive as email and attachments; people retype them into the CRM.
Usually missingA defined intake task with an extraction contractThe agent picks a customer when two accounts look the same.
Usually missingAn identity rule, and a stop when it is not metLow-confidence output is committed because nothing said otherwise.
Usually missingAn ambiguity policy tied to consequenceThe agent’s credentials can do far more than its task needs.
Usually missingAn authority matrix and least-privilege toolsReviewers approve AI output without seeing why it was proposed.
Usually missingA review card with evidence and uncertaintyBusiness policy lives in prompts nobody versions.
Usually missingPolicies as data, prompts as versioned configurationQuality is judged from a good demo and a few anecdotes.
Usually missingAn evaluation set scored per dimensionNobody can reconstruct why the agent acted.
Usually missingAn audit model and an execution trace
Artefacts that make an agent safe to put in a real process.
Each one answers a question a demo never has to.
- Agent workflowWhich steps, which actors, where it stops?
- Context architectureWhat the agent knows, from which source, how fresh?
- Authority matrixRead, recommend, draft, execute, approve or forbid—per action
- Tool catalog & contractsInput, output, preconditions, side effect, reversibility
- Guardrail modelWhat is checked before and after every action?
- Human review modelWhat the reviewer sees, and what each decision does
- Ambiguity policyRetrieve, ask, escalate or stop—when?
- Confidence policyHow confidence and consequence combine
- Evidence modelWhich sources support each recommendation?
- Fallback & escalation modelWhat happens when the agent or a tool fails?
- Evaluation setScored cases per dimension, not one accuracy number
- Audit trailInput, context, tools, evidence, decision, action
- Observability modelWhich trace shows what the agent did, step by step?
- Improvement loopHow corrections become evaluation cases
- Release governanceVersioned prompts and policies, gated by regression
More context is not always better context.
An agent reasons over what it is given. Six layers of context, read two ways: dumped into one prompt, or assembled from the systems that own each fact.
The agent reasons over typed, sourced, current context—each item with its key and timestamp.
- Relevant
Needed for this task and this step—not everything the agent could read.
- +Trusted source
From the system that owns the fact, with its key.
- +Timely state
Fresh enough for the decision, with its timestamp.
AI cannot fix unreliable identity. It can rank candidates. A trusted key—or a person—confirms them.
Where identity and authority are designed: Data & Integrations
A tool is an architecture contract, not an API endpoint.
The agent can only do what its tools allow—so each tool states its purpose, input, output, permissions, preconditions, side effect, failure and reversibility, and enforces them in code.
approveDiscountA commercial decision right. The agent can see the policy; it cannot exercise it.
createErpCustomerLegal and financial creation runs through the orchestrated process with finance validation.
mergeAccountsMaster-data authority stays with data stewards; the agent may only suggest candidates.
createOpportunityDraft
Agent authority: ExecuteA person commits later—in review
- Purpose
- Prepare a reviewable draft from confirmed data
- Input
- Account key, product lines, source request ID
- Output
- Draft ID and every value with its source
- Permissions
- Create drafts only; no stage, owner or amount beyond the draft
- Preconditions
- Account confirmed or flagged for review; request not already drafted
- Side effect
- Creates a draft only—invisible to pipeline, forecast and customer
- Failure
- Duplicate request ID → return the existing draft; validation error → stop with reason
- Reversibility
- Fully reversible: drafts expire if not approved
Agent workflows are stateful processes
The same run model as any orchestration: states, valid transitions, who moves each one—agent, system, person, timer—and terminal outcomes. Select a state.
Waiting review
A person holds the decision; the draft and evidence are ready.
- CommittedApproved (possibly edited)Person
- RejectedRejected with a reason codePerson
- Context readyReviewer requests more contextPerson
Entered from Reasoning, Escalated.
Human-in-the-loop is a design decision, not a fallback for bad AI.
The reviewer sees what the agent believes, why, the source evidence, the uncertainty and exactly what approval will do—and every decision is recorded with a reason.
RFQ from a distributor · three lines
This is an RFQ from an existing distributor for two known products and one probable substitute.
| Value | Confidence | Evidence |
|---|---|---|
| CustomerDistributor account · regional entity | High | Sender domain and signature match the account; four prior RFQs from this contact |
| Line 1Product A-100 × 200 | High | Exact catalogue reference in the attachment |
| Line 2Product B-220 × 50 | High | Cross-reference to a competitor code, bought twice last year |
| Line 3Product C-310X × 20 | Low | “C310” matches two products; the X variant was ordered once |
Create an opportunity at Qualified for the account owner, with three lines; line 3 as C-310 unless edited. No email is sent.
- Reason code
- Wrong product variant
- What happens
- Line 3 becomes C-310X; the edited draft is committed.
- What it teaches
- The correction becomes a labelled evaluation case for product matching.
- Accountability
A named person made the consequential decision, on recorded evidence.
- Learning data
Every edit and rejection, with its reason code, is a labelled example.
- Auditability
Anyone can see what the agent proposed and what the human changed.
Confidence is not authority.
Confidence says how sure the model is. Consequence says what happens if it is wrong. Pick an action and move its confidence: in the high-consequence row, no score reaches “execute”.
Unsure and consequential: stop and hand over everything gathered.
Sure, but consequential: prepare it perfectly—a person commits.
Cheap to get more context. Retrieve, ask the requester, or hand to a person.
Low consequence and high confidence: act, visibly and reversibly.
Human approval High confidence does not grant authority: a person commits.
A good demo is not an evaluation.
Agent quality is scored on a fixed set, per dimension, against its own target—never as one “AI accuracy” number. And every run can be read step by step.
- ClassificationDid it identify the request correctly?97 %target 95 %0
- ExtractionDid it capture the required data?95 %target 90 %+4
- Entity matchingDid it identify the right account and products?93 %target 90 %+4
- ReasoningDid it choose the right next step?93 %target 90 %+1
- Tool useDid it call the right tool with the right arguments?96 %target 95 %-1
- AuthorityDid it stop when it had to?98 %target 100 % · gate-2
- EvidenceDid it support its recommendation?96 %target 95 %0
- OutcomeDid the process end as expected?88 %target 85 %+2
Blocked. Better extraction and matching do not compensate for a single missed stop: the authority gate is 100 %.
One run, as an execution trace
Time, actor, action, evidence and outcome for every step. AI operations should be inspectable.
- SystemRequest receivedEvidence: Message from a distributor domain, one attachmentOutcome: Run R-7F3 created
- AgentContext retrievedEvidence: Account candidates, 12 months of orders, RFQ policy v4Outcome: Context ready
- AgentAccount matchedEvidence: Domain + signature + four prior RFQsOutcome: One candidate, high confidence
- AgentProduct 1 matchedEvidence: Exact catalogue referenceOutcome: A-100 × 200
- AgentProduct 2 ambiguousEvidence: “C310” matches C-310 and C-310XOutcome: Flagged; not resolved by the agent
- AgentHuman review requestedEvidence: Draft, evidence and one open questionOutcome: Task for the account owner
- HumanProduct correctedEvidence: Reviewer chose C-310X · reason: wrong variantOutcome: Correction stored as evaluation case
- SystemDraft committedEvidence: Approval reference on the recordOutcome: Opportunity created
Five agents, each with a boundary.
Synthetic B2B scenarios. The RFQ intake agent is the flagship—the agent design behind the existing lab, reference architecture and process case, which it links to rather than repeats.
- Case 01Selected workAuthority, tools & human commit
Designing an AI-assisted RFQ intake agent
RFQs arrive as free-text email and attachments, and an agent is expected to “handle them”—with nobody having said what it may do.
When is a customer match trusted, what happens with an ambiguous product—and what may the agent never do, however confident it is?
MailboxAI agentCRMProduct catalogueProcess ArchitectureSystems & CRM Architecture - Case 02Context assembly, fact vs inference
Preparing account context before a commercial visit
Salespeople prepare visits by opening six screens—or not at all—and an AI summary is proposed that may invent what it cannot find.
Which sources are authoritative, how fresh must they be—and how does the briefing separate fact from inference?
CRMERPData platformCalendarRevenue OperationsData & Integrations - Case 03Taxonomy, routing & mixed intent
Turning commercial email into structured work
A shared commercial inbox mixes RFQs, pricing questions, complaints, technical questions, leads and order questions—and response time depends on who looks first.
Which classes are safe to route automatically, what happens with mixed intent—and may the agent create tasks or send replies?
MailboxAI agentCRMService deskProcess ArchitectureAutomation - Case 04Evidence-backed summary, manager decides
AI-assisted account review with evidence
Monthly account reviews start with managers assembling numbers; the discussion that follows has little time and no shared evidence.
What may the agent prepare for a performance review—and where must it stop short of commercial judgement?
Data platformCRMPlanningRevenue OperationsData & Integrations - Case 05Suggest with evidence, steward decides
Using AI to assist data quality without giving it master-data authority
The CRM has probable duplicate accounts, inconsistent names and missing classifications—and a proposal to let a model fix them.
Where does AI help data quality—and why must merge, rename and reclassification stay with a person?
CRMERPData platformSteward queueData & IntegrationsSystems & CRM Architecture
Every AI request has a second layer.
The visible request is where the design starts, not where it ends.
“Add an AI assistant to the CRM.”
A chat panel that can answer questions about any record.
4 of 5 questions surfaced
“Answer anything” has no definition of done and cannot be evaluated.
Process defines the task. Systems hold the truth. The agent works in between.
AI & Agentic Workflows
- Process Architecture
The process defines the task, the decision rights and the review points the agent works within.
- Systems & CRM Architecture
Systems of record own identity and state; the agent reads them and writes only through bounded tools.
- Data & Integrations
Context is only as good as its source, key and freshness—AI cannot fix unreliable identity.
- Automation
Agent workflows are stateful runs: the same states, guards, retries and recovery as any orchestration.
- Revenue Operations
Account review and visit preparation turn governed signals into evidence-backed briefings.
- Digital Operating Model
Someone owns the agent after go-live: its policies, evaluation and release.
Architecture decisions
Lab, reference & writing
An agentic workflow is not a model with access. It is a designed process with
- A task
- Context
- Authority
- Tools
- Ambiguity rules
- Review
- Evidence
- Evaluation
- Release control
A good agent knows when to stop.
Design the authority before giving the agent tools.
An agent that demos well—and nobody has decided what it may do?
- Which of its actions are reversible?
- What does it do when two customers look the same?
- How would you know it got worse after a prompt change?