Designing evidence-grounded AI preparation for commercial decisions
A fictional, composite case: AI prepares, a person decides. The moment changes—a visit, a review—the architecture does not.
Independently created. Contains no employer or client implementation detail, internal names or figures.
The case in brief
Current reality
Before a visit or an account review, the context a person needs is spread across accounts, pipeline, orders, invoices, published performance, recovery actions, service cases and notes—and the proposed fix, “summarize the account”, reads fluently even when a source is missing, stale or wrong.
What must become true
Before a commercial decision, AI prepares a concise briefing from governed sources: the published version of every signal, freshness respected, every factual statement cited, fact kept apart from inference, anomaly from explanation and open action from recommendation, questions surfaced instead of invented causes—no false business fact, no silent update to any record, and commercial judgement left with the person responsible.
Design question
What may AI assemble and infer from commercial context—and how do we ensure that evidence, uncertainty and human judgement remain distinguishable?
01The reality · current state
What happens today—or in the pilot.
Commercial decisions start with context assembly. Before a visit, a salesperson needs the account, its contacts, recent activity, open opportunities, orders and invoices, performance and what was promised last time. Before the monthly performance review, a manager needs the published numbers, what changed since the previous review, the recovery plan and which actions are overdue. The information exists—in several systems, at different freshness. People open several screens, build an export or a presentation, reconstruct what changed by hand, work from memory—or go in unprepared. A generative summary is attractive and risky: it reads well even when a source is missing, stale or wrong.
- Preparation depends on each person’s habits and the time they have
- A pilot visit summary mentioned an order that had been cancelled
- A pilot review page recommended a price reduction nobody could support
- Numbers in the summary did not match the published dashboard
- Conflicts between CRM and ERP were silently averaged
- Nobody could tell which statements were facts and which were guesses
- Summaries pasted into notes became “facts” later
What the apparent request hides
The apparent request is “summarize the account”. It hides the questions that decide whether the summary can be trusted.
- Which sources have standing?An authoritative record, a derived signal and a note are not the same kind of truth.
- Which version is current?A performance number is only meaningful as a published version with an as-of date.
- What is fact—and what is inference?“Orders fell” and “the customer switched supplier” read alike in fluent prose.
- What changed?Change needs two points in time compared at the same definition.
- What counts as an anomaly?An anomaly is a pattern worth attention, not an explanation of it.
- What may the model suggest?Questions and areas of attention—or pricing and strategy?
- Which statements need a citation?Every factual or changed-state statement, or none of them can be trusted.
- May it write anything back?A preparation task that writes becomes a source of unreviewed facts.
- What if the evidence conflicts?Two systems disagree; an average of them is a number nobody recorded.
- What if identity is uncertain?Context from two possible customers, blended, is worse than no context.
02Work decomposition
Reasoning, rules and authority—kept apart.
Before any model reads anything, the work is split by what it needs: reasoning a model can do, rules a system enforces, and authority that stays with the person who decides.
AI reasoning
- Summarize recent history
- Compare now with the previous interaction or review
- Surface anomalies
- Suggest areas of attention
- Generate questions
- Type every statement
Deterministic system rules
- Resolve one governed account key
- Enforce the requester’s visibility
- Select the published version of each signal
- Apply each source’s freshness rule
- Attach source, version and timestamp
- Exclude sources without standing
Human authority
- Visit objective
- Interpretation of what is heard
- Diagnosis
- Pricing and strategy
- Commitments to the customer
- Every record change
Reasoning does not imply authority: the model reasons over context the rules assembled, and its output is a brief, not a record.
One architecture, two moments
The context contract, statement types, tools and limits are the same; only the task and the person who decides change.
Visit preparation
A visit on the calendar for a known account, the day before
- Relevant changes since the previous interaction
- Open topics and commitments
- Performance context
- Recent transactions and activity
- Questions worth asking
- Unresolved issues
- Decides the visit objective
- Conducts the conversation
- Interprets what is heard
- Records what actually happened
- Decides the follow-up
Account and performance review
An account on the list for the monthly performance review
- Governed, published performance
- Changes since the previous review
- Open recovery actions
- Anomalies
- The evidence for each
- Questions for discussion
- Diagnoses
- Decides
- Escalates
- Changes the strategy
- Closes or modifies recovery work
03Context contract
Every source has standing; every line has a type.
Context is architecture. Every source must be relevant, trusted, timely, permitted, identified and structured before it reaches the model—more tokens do not replace governed context. The contract is written per source before any prompt: what it is, what standing it has, how fresh it must be, which entity it refers to, which version counts, what AI may use it for and how the reader traces a statement back to it. Context is never flattened into “what the model received”: a governed analytical fact, an operational fact, a person’s account of a meeting, an external item and the model’s own inference have different standing—and the brief shows which is which. No real field or object names.
Context contract
| Source | Standing | Freshness | Identity | Version | Allowed use | Citation / traceability |
|---|---|---|---|---|---|---|
| Published performance & segment | Governed analytical fact | Monthly, versioned | Governed account key | Published version only | Fact as published · comparison | Signal, version and as-of date |
| Account & contacts | Authoritative fact | Live | Governed account key | — | Fact | Record link |
| Open opportunities | Operational fact | Live | Account key | — | Fact · comparison | Record link and stage as of today |
| Orders & invoices | Authoritative fact (ERP) | Daily | ERP customers linked to the account | — | Fact · comparison · inference input | Document reference and refresh time |
| Recovery plan & actions | Operational fact | Live | Account key | — | Open action | Action, owner and due date |
| Service cases | Operational context | Live | Account key | — | Inference input · question generation | Case reference |
| Visit and review notes | User-authored account of an interaction | As written | Account key | — | Inference input · question generation—never a fact about the customer | Author and date |
| Public company news | External context · separate provenance | Weekly | Legal entity, matched and confirmed | — | Question generation only | Source, date and reliability |
| AI inference | Interpretation—never fact | — | The brief it appears in | — | Shown labelled; never cited as a source | The facts it rests on |
| Email threads | Excluded | — | — | — | Not read: consent and relevance not established | — |
Statement types
Every statement in a brief has a type. Presenting an inference as a fact is the failure the architecture exists to prevent.
| Type | What it asserts | Evidence it needs | Rule |
|---|---|---|---|
| Fact | Something recorded now | Source, version and timestamp | Never inferred |
| Change | A difference between two points in time | Both values and both dates | Compared at the same definition |
| Anomaly | A pattern worth attention | The data that shows it | Not an explanation |
| Open action | Committed work that is still open | The action, its owner and due date | Not a recommendation |
| Inference | A possible explanation | The facts it rests on | Always labelled; never promoted to fact |
| Question | What the evidence cannot answer | Why it is asked | Replaces an unsupported cause |
04Agent workflow
Each step with its actor, its tool and its authority.
CRMERPData platformCalendar
- 01Trigger
A visit on the calendar or an account on the review list—never a free-form request to “look into” a customer.
- 02Resolve identity
One governed account key, within the requester’s visibility; no briefing until it is unambiguous.
- 03Assemble context
Each source read through its contract: standing, freshness and version attached to every value.
- 04Published signals & open actions
Only published signal versions; open actions with owners and due dates.
- 05AI preparation
Summarizes, compares, flags anomalies and writes questions; every statement typed and cited.
- 06Human decision
The salesperson or manager reads, challenges and decides; records are written by people.
05Key dimension · Context contract, typed evidence, read-only authority
Governed context in, a typed and cited brief out
Context is assembled by rules before the model reasons. The model’s output is a brief whose every line has a type and a trail back to its evidence. The decision sits outside the architecture’s reach.
- SystemSources with standing
- Account
- Transactions
- Pipeline
- Published signals
- Actions & plans
- Activity
- SystemContext contract
- Identity
- Freshness
- Version
- Standing
- Permission
- AgentAI preparation
- Fact
- Change
- Anomaly
- Open action
- Inference
- Question
- AgentEvidence-grounded brief
- Every line typed
- Every fact cited
- Gaps stated
- HumanHuman decision
- Objective
- Diagnosis
- Follow-up
- Visit objective
- Diagnosis
- Strategy and pricing
- Customer commitment
- Every record change
Briefing excerpt · every line typed and traced
| Statement | Type | Source | Version | Evidence |
|---|---|---|---|---|
| Invoiced value is below the published year-to-date target. | Fact | Performance to target | Published version · as of the 3rd | Governed performance model |
| Open pipeline is lower than at the previous review. | Change | Open opportunities | Live · against the review snapshot | Pipeline snapshot taken at the last review |
| One product family has had no recent orders. | Anomaly | Order lines | Daily · refreshed this morning | Orders by product family, last 90 days |
| Two recovery actions are overdue. | Open action | Recovery plan | Live | Due dates on the plan’s actions |
| Demand may have shifted to another supplier. | Inference | Derived from the anomaly above | — | Not confirmed by the customer: labelled, never stated as fact |
| Has the customer qualified another supplier? | Question | Suggested for the conversation | — | Asked because the evidence does not support a cause |
A model’s confidence is not evidence. A statement is trusted because it traces to a source, a version and a record—and a cause the evidence does not support becomes a question.
07Ambiguity policy
When the agent is unsure.
- Stale dataStop
Marked stale or left out—never presented as current
- Conflicting dataContinue
Both values shown as a conflict, never averaged; if the published numbers themselves conflict, the brief is not produced
- Missing contextContinue
The brief states what is missing instead of filling the gap
- Unknown causeAsk
Phrased as a question for the conversation or the account owner
- Identity uncertaintyStop
No brief until one account is confirmed—context from two possible customers is never mixed
- No evidenceStop
The claim is not made
08Human review
Review where ambiguity becomes consequential.
Human review is not “a person checks the AI”. The person receives a brief built for judgement—and keeps everything that is judgement.
| Review point | Why here | What the person sees | What they can do | What happens next |
|---|---|---|---|---|
| Before the visit or the review | The brief feeds a commercial decision; the decision, not the brief, carries the consequence | Typed statements with source, version and evidence; what is missing, stale or conflicting; suggested questions | Use, challenge or dismiss any line; ask for more context; report a wrong statement | The person decides and records what happened; a reported statement becomes an evaluation case |
- Statement type
- Fact, change, anomaly, open action, inference or question—on every line
- Evidence
- The record behind each factual or changed-state statement
- Source
- Which system or note it came from, and its standing
- Freshness & version
- As-of date, refresh time or published version
- Uncertainty
- What is missing, stale, conflicting or unconfirmed
- Questions & suggestions
- Labelled, and free to dismiss
- Interpretation
- Diagnosis
- Strategy
- Commitment
- Operational updates
09Evaluation
Scored per dimension—never one accuracy number.
| Dimension | Question | Measure | Target |
|---|---|---|---|
| Evidence coverageEvidence & correctness | Is every factual and changed-state statement cited? | Statements with a valid source, version and evidence | 100 % (synthetic design target) |
| FaithfulnessEvidence & correctness | Does each statement match its source? | Reviewer-verified statements | ≥ 98 % (synthetic design target) |
| Statement typeTask quality | Is each statement typed correctly—no inference presented as fact? | Correct type on the labelled set | ≥ 95 % (synthetic design target) |
| Freshness complianceEvidence & correctness | Were stale sources left out or marked? | Briefs honouring every freshness rule | 100 % (synthetic design target) |
| Entity correctnessEvidence & correctness | Is all context from the right customer? | Briefs with a single, correct identity | 100 % (synthetic design target) |
| Out-of-bounds adviceAuthority compliance | No pricing, strategy or record changes? | Out-of-bounds statements | 0 — release gate |
| UsefulnessTask quality | Did the person use it? | Questions and attention areas kept or acted on | Tracked, no target yet |
10The trade-offs
Credible options, judged against these premises.
Let the model search every system freely
Exploration by an analyst
Cost: Unknown sources, stale data, no citations, mixed identitiesTwo separate agents, one per moment
Moments with nothing in common
Cost: Two context models and two sets of limits that drift apartA static report, no AI
Stable, simple accounts
Cost: Change, conflict and anomalies still read between screensDeterministic assembly, typed and cited preparation, read-only tools
Decisions that depend on recent change
Cost: A context contract, a citation format and an evaluation set to maintainDesign decision
Assemble governed context deterministically before any reasoning; let AI summarize, compare, surface anomalies and generate questions; type and cite every output; expose no autonomous operational write authority during preparation; and leave commercial decisions and record changes with people.
AI prepares; a person decides.
11Audit, governance & learning
Consequential activity can be reconstructed afterwards.
A brief that shaped a decision can be reconstructed afterwards—without a logging interface anyone has to read every day.
- The account key and the requester’s visibility
- Every source read, with its version or refresh time
- The tools called and what each returned
- The brief as shown, every statement with its type and citation
- Statements the person reported as wrong
- The decision itself—recorded by the person where it is made, never by the brief
Reported statements improve the next brief. They do not change what the architecture may do.
- The evaluation set
- The context-contract review
- Freshness and exclusion rules
- Prompt changes, released only against the evaluation set
- Grant write access
- Promote an inference to a fact
- Change a source’s standing without its owner
The owner of the preparation service, with a periodic review alongside the sales and data owners
12The second layer
Questions that change the design.
Sources
- Which sources are authoritative—and which are deliberately excluded?
- Which version of a signal does the brief use?
- How fresh must each source be?
Truthfulness
- How is inference kept apart from fact?
- How are citations shown?
- What happens when sources conflict?
Use
- Who may write what after the visit or review?
- How is usefulness measured?
- How are wrong statements reported and turned into evaluation cases?
13Decisions & outputs
What the work produces.
- 01Context contract
- 02Statement types & citation format
- 03Bounded read tools
- 04Briefing and review-page templates
- 05Ambiguity policy
- 06Evaluation set