AI case / 01 · Context contract, typed evidence, read-only authority

Designing evidence-grounded AI preparation for commercial decisions

A fictional, composite case: AI prepares, a person decides. The moment changes—a visit, a review—the architecture does not.

Fictional scenario

Independently created. Contains no employer or client implementation detail, internal names or figures.

The case in brief

Current reality

Before a visit or an account review, the context a person needs is spread across accounts, pipeline, orders, invoices, published performance, recovery actions, service cases and notes—and the proposed fix, “summarize the account”, reads fluently even when a source is missing, stale or wrong.

What must become true

Before a commercial decision, AI prepares a concise briefing from governed sources: the published version of every signal, freshness respected, every factual statement cited, fact kept apart from inference, anomaly from explanation and open action from recommendation, questions surfaced instead of invented causes—no false business fact, no silent update to any record, and commercial judgement left with the person responsible.

Design question

What may AI assemble and infer from commercial context—and how do we ensure that evidence, uncertainty and human judgement remain distinguishable?

01The reality · current state

What happens today—or in the pilot.

Commercial decisions start with context assembly. Before a visit, a salesperson needs the account, its contacts, recent activity, open opportunities, orders and invoices, performance and what was promised last time. Before the monthly performance review, a manager needs the published numbers, what changed since the previous review, the recovery plan and which actions are overdue. The information exists—in several systems, at different freshness. People open several screens, build an export or a presentation, reconstruct what changed by hand, work from memory—or go in unprepared. A generative summary is attractive and risky: it reads well even when a source is missing, stale or wrong.

  • Preparation depends on each person’s habits and the time they have
  • A pilot visit summary mentioned an order that had been cancelled
  • A pilot review page recommended a price reduction nobody could support
  • Numbers in the summary did not match the published dashboard
  • Conflicts between CRM and ERP were silently averaged
  • Nobody could tell which statements were facts and which were guesses
  • Summaries pasted into notes became “facts” later

What the apparent request hides

The apparent request is “summarize the account”. It hides the questions that decide whether the summary can be trusted.

  • Which sources have standing?An authoritative record, a derived signal and a note are not the same kind of truth.
  • Which version is current?A performance number is only meaningful as a published version with an as-of date.
  • What is fact—and what is inference?“Orders fell” and “the customer switched supplier” read alike in fluent prose.
  • What changed?Change needs two points in time compared at the same definition.
  • What counts as an anomaly?An anomaly is a pattern worth attention, not an explanation of it.
  • What may the model suggest?Questions and areas of attention—or pricing and strategy?
  • Which statements need a citation?Every factual or changed-state statement, or none of them can be trusted.
  • May it write anything back?A preparation task that writes becomes a source of unreviewed facts.
  • What if the evidence conflicts?Two systems disagree; an average of them is a number nobody recorded.
  • What if identity is uncertain?Context from two possible customers, blended, is worse than no context.

02Work decomposition

Reasoning, rules and authority—kept apart.

Before any model reads anything, the work is split by what it needs: reasoning a model can do, rules a system enforces, and authority that stays with the person who decides.

AI reasoning

  • Summarize recent history
  • Compare now with the previous interaction or review
  • Surface anomalies
  • Suggest areas of attention
  • Generate questions
  • Type every statement

Deterministic system rules

  • Resolve one governed account key
  • Enforce the requester’s visibility
  • Select the published version of each signal
  • Apply each source’s freshness rule
  • Attach source, version and timestamp
  • Exclude sources without standing

Human authority

  • Visit objective
  • Interpretation of what is heard
  • Diagnosis
  • Pricing and strategy
  • Commitments to the customer
  • Every record change

Reasoning does not imply authority: the model reasons over context the rules assembled, and its output is a brief, not a record.

One architecture, two moments

The context contract, statement types, tools and limits are the same; only the task and the person who decides change.

Visit preparation

WhenA visit on the calendar for a known account, the day before

AI prepares
  • Relevant changes since the previous interaction
  • Open topics and commitments
  • Performance context
  • Recent transactions and activity
  • Questions worth asking
  • Unresolved issues
The salesperson
  • Decides the visit objective
  • Conducts the conversation
  • Interprets what is heard
  • Records what actually happened
  • Decides the follow-up

Account and performance review

WhenAn account on the list for the monthly performance review

AI prepares
  • Governed, published performance
  • Changes since the previous review
  • Open recovery actions
  • Anomalies
  • The evidence for each
  • Questions for discussion
The manager
  • Diagnoses
  • Decides
  • Escalates
  • Changes the strategy
  • Closes or modifies recovery work

03Context contract

Every source has standing; every line has a type.

Context is architecture. Every source must be relevant, trusted, timely, permitted, identified and structured before it reaches the model—more tokens do not replace governed context. The contract is written per source before any prompt: what it is, what standing it has, how fresh it must be, which entity it refers to, which version counts, what AI may use it for and how the reader traces a statement back to it. Context is never flattened into “what the model received”: a governed analytical fact, an operational fact, a person’s account of a meeting, an external item and the model’s own inference have different standing—and the brief shows which is which. No real field or object names.

Context contract

Context contract: for each source, its standing, freshness, identity, version, allowed use and citation
SourceStandingFreshnessIdentityVersionAllowed useCitation / traceability
Published performance & segmentGoverned analytical factMonthly, versionedGoverned account keyPublished version onlyFact as published · comparisonSignal, version and as-of date
Account & contactsAuthoritative factLiveGoverned account key—FactRecord link
Open opportunitiesOperational factLiveAccount key—Fact · comparisonRecord link and stage as of today
Orders & invoicesAuthoritative fact (ERP)DailyERP customers linked to the account—Fact · comparison · inference inputDocument reference and refresh time
Recovery plan & actionsOperational factLiveAccount key—Open actionAction, owner and due date
Service casesOperational contextLiveAccount key—Inference input · question generationCase reference
Visit and review notesUser-authored account of an interactionAs writtenAccount key—Inference input · question generation—never a fact about the customerAuthor and date
Public company newsExternal context · separate provenanceWeeklyLegal entity, matched and confirmed—Question generation onlySource, date and reliability
AI inferenceInterpretation—never fact—The brief it appears in—Shown labelled; never cited as a sourceThe facts it rests on
Email threadsExcluded———Not read: consent and relevance not established—

Statement types

Every statement in a brief has a type. Presenting an inference as a fact is the failure the architecture exists to prevent.

Statement types: what each asserts, the evidence it needs and its rule
TypeWhat it assertsEvidence it needsRule
FactSomething recorded nowSource, version and timestampNever inferred
ChangeA difference between two points in timeBoth values and both datesCompared at the same definition
AnomalyA pattern worth attentionThe data that shows itNot an explanation
Open actionCommitted work that is still openThe action, its owner and due dateNot a recommendation
InferenceA possible explanationThe facts it rests onAlways labelled; never promoted to fact
QuestionWhat the evidence cannot answerWhy it is askedReplaces an unsupported cause

04Agent workflow

Each step with its actor, its tool and its authority.

Systems and partiesCRMERPData platformCalendar

  1. 01
    Trigger

    A visit on the calendar or an account on the review list—never a free-form request to “look into” a customer.

    SystemRead
  2. 02
    Resolve identity

    One governed account key, within the requester’s visibility; no briefing until it is unambiguous.

    SystemReadRead account context
  3. 03
    Assemble context

    Each source read through its contract: standing, freshness and version attached to every value.

    SystemReadRead recent transactions
  4. 04
    Published signals & open actions

    Only published signal versions; open actions with owners and due dates.

    SystemReadRead published signals
  5. 05
    AI preparation

    Summarizes, compares, flags anomalies and writes questions; every statement typed and cited.

    AgentRecommend
  6. 06
    Human decision

    The salesperson or manager reads, challenges and decides; records are written by people.

    HumanHuman approval

05Key dimension · Context contract, typed evidence, read-only authority

Governed context in, a typed and cited brief out

Context is assembled by rules before the model reasons. The model’s output is a brief whose every line has a type and a trail back to its evidence. The decision sits outside the architecture’s reach.

From governed context to a human decision
  1. SystemSources with standing
    • Account
    • Transactions
    • Pipeline
    • Published signals
    • Actions & plans
    • Activity
  2. SystemContext contract
    • Identity
    • Freshness
    • Version
    • Standing
    • Permission
  3. AgentAI preparation
    • Fact
    • Change
    • Anomaly
    • Open action
    • Inference
    • Question
  4. AgentEvidence-grounded brief
    • Every line typed
    • Every fact cited
    • Gaps stated
  5. HumanHuman decision
    • Objective
    • Diagnosis
    • Follow-up
Human authority—outside the architecture’s reach
  • Visit objective
  • Diagnosis
  • Strategy and pricing
  • Customer commitment
  • Every record change

Briefing excerpt · every line typed and traced

Briefing excerpt: each statement with its type, source, version and evidence
StatementTypeSourceVersionEvidence
Invoiced value is below the published year-to-date target.FactPerformance to targetPublished version · as of the 3rdGoverned performance model
Open pipeline is lower than at the previous review.ChangeOpen opportunitiesLive · against the review snapshotPipeline snapshot taken at the last review
One product family has had no recent orders.AnomalyOrder linesDaily · refreshed this morningOrders by product family, last 90 days
Two recovery actions are overdue.Open actionRecovery planLiveDue dates on the plan’s actions
Demand may have shifted to another supplier.InferenceDerived from the anomaly above—Not confirmed by the customer: labelled, never stated as fact
Has the customer qualified another supplier?QuestionSuggested for the conversation—Asked because the evidence does not support a cause

A model’s confidence is not evidence. A statement is trusted because it traces to a source, a version and a record—and a cause the evidence does not support becomes a question.

06Tool & authority boundary

Bounded read tools—never raw system access.

The agent does not call platform interfaces. It calls four read capabilities whose contracts fix purpose, inputs, identity and permission scope, output, freshness and failure behaviour. None of them has a side effect—and that is visible in the contract, not promised in a prompt. Tool names are descriptive, not implementation names.

Tool contracts: purpose, inputs and scope, output, freshness or version, side effects and failure behaviour
ToolPurposeInputs & scopeReturnsFreshness / versionSide effectsOn failure
Read account contextWho the account is and who is involvedOne governed account key · only what the requester may seeAccount, contacts, owner, last interactionLiveNoneUnknown or ambiguous key → stop
Read published signalsPerformance and segment as publishedThe same key and a period · the requester’s visibilityValues with signal, version and as-of datePublished version onlyNoneNo published version → left out, and said so
Read open actionsOpen tasks, commitments and recovery actionsThe same key · the requester’s visibilityActions with owner and due dateLiveNoneUnavailable → actions stated as unknown
Read recent transactionsOrders and invoices in a windowERP customers linked to the key and a date window · the requester’s visibilityLines by product family with document referencesDailyNoneOlder than its rule → excluded with a note

Authority by action

Actions and the authority the agent has for each
ActionAuthorityWhy
Retrieve allowed contextReadThrough the four read tools, within the requester’s visibility
Summarize, compare and identify changeRecommendOutput is a brief, not a record
Flag anomalies and propose questionsRecommendLabelled suggestions a person may dismiss
Suggest areas of attentionRecommendAttention, not a decision
Change opportunity state or account dataForbiddenNo write tools in a preparation task
Write visit notesForbiddenNotes are a person’s record of what happened
Close tasks or change recovery plansForbiddenPlan outcomes are decided in review
Recommend pricing or commit customer strategyForbiddenCommercial judgement belongs to the person responsible
Create a business fact that did not existForbiddenEvery fact must trace to a record

07Ambiguity policy

When the agent is unsure.

  • Stale dataStop

    Marked stale or left out—never presented as current

  • Conflicting dataContinue

    Both values shown as a conflict, never averaged; if the published numbers themselves conflict, the brief is not produced

  • Missing contextContinue

    The brief states what is missing instead of filling the gap

  • Unknown causeAsk

    Phrased as a question for the conversation or the account owner

  • Identity uncertaintyStop

    No brief until one account is confirmed—context from two possible customers is never mixed

  • No evidenceStop

    The claim is not made

08Human review

Review where ambiguity becomes consequential.

Human review is not “a person checks the AI”. The person receives a brief built for judgement—and keeps everything that is judgement.

Review points: why review sits there, what the person sees, what they can do and what happens next
Review pointWhy hereWhat the person seesWhat they can doWhat happens next
Before the visit or the reviewThe brief feeds a commercial decision; the decision, not the brief, carries the consequenceTyped statements with source, version and evidence; what is missing, stale or conflicting; suggested questionsUse, challenge or dismiss any line; ask for more context; report a wrong statementThe person decides and records what happened; a reported statement becomes an evaluation case
The person receives
Statement type
Fact, change, anomaly, open action, inference or question—on every line
Evidence
The record behind each factual or changed-state statement
Source
Which system or note it came from, and its standing
Freshness & version
As-of date, refresh time or published version
Uncertainty
What is missing, stale, conflicting or unconfirmed
Questions & suggestions
Labelled, and free to dismiss
Remains their responsibility
  • Interpretation
  • Diagnosis
  • Strategy
  • Commitment
  • Operational updates

09Evaluation

Scored per dimension—never one accuracy number.

Evaluation dimensions: what each proves, its question, measure and target
DimensionQuestionMeasureTarget
Evidence coverageEvidence & correctnessIs every factual and changed-state statement cited?Statements with a valid source, version and evidence100 % (synthetic design target)
FaithfulnessEvidence & correctnessDoes each statement match its source?Reviewer-verified statements≥ 98 % (synthetic design target)
Statement typeTask qualityIs each statement typed correctly—no inference presented as fact?Correct type on the labelled set≥ 95 % (synthetic design target)
Freshness complianceEvidence & correctnessWere stale sources left out or marked?Briefs honouring every freshness rule100 % (synthetic design target)
Entity correctnessEvidence & correctnessIs all context from the right customer?Briefs with a single, correct identity100 % (synthetic design target)
Out-of-bounds adviceAuthority complianceNo pricing, strategy or record changes?Out-of-bounds statements0 — release gate
UsefulnessTask qualityDid the person use it?Questions and attention areas kept or acted onTracked, no target yet

10The trade-offs

Credible options, judged against these premises.

Rejected

Let the model search every system freely

Exploration by an analyst

Cost: Unknown sources, stale data, no citations, mixed identities
Rejected

Two separate agents, one per moment

Moments with nothing in common

Cost: Two context models and two sets of limits that drift apart
Situational

A static report, no AI

Stable, simple accounts

Cost: Change, conflict and anomalies still read between screens
Selected

Deterministic assembly, typed and cited preparation, read-only tools

Decisions that depend on recent change

Cost: A context contract, a citation format and an evaluation set to maintain

Design decision

Assemble governed context deterministically before any reasoning; let AI summarize, compare, surface anomalies and generate questions; type and cite every output; expose no autonomous operational write authority during preparation; and leave commercial decisions and record changes with people.

Key principleAI prepares; a person decides.

11Audit, governance & learning

Consequential activity can be reconstructed afterwards.

A brief that shaped a decision can be reconstructed afterwards—without a logging interface anyone has to read every day.

Reconstructable for every consequential run
  1. The account key and the requester’s visibility
  2. Every source read, with its version or refresh time
  3. The tools called and what each returned
  4. The brief as shown, every statement with its type and citation
  5. Statements the person reported as wrong
  6. The decision itself—recorded by the person where it is made, never by the brief

Reported statements improve the next brief. They do not change what the architecture may do.

Outcomes feed
  • The evaluation set
  • The context-contract review
  • Freshness and exclusion rules
  • Prompt changes, released only against the evaluation set
They never
  • Grant write access
  • Promote an inference to a fact
  • Change a source’s standing without its owner

Governed byThe owner of the preparation service, with a periodic review alongside the sales and data owners

12The second layer

Questions that change the design.

Sources

  1. Which sources are authoritative—and which are deliberately excluded?
  2. Which version of a signal does the brief use?
  3. How fresh must each source be?

Truthfulness

  1. How is inference kept apart from fact?
  2. How are citations shown?
  3. What happens when sources conflict?

Use

  1. Who may write what after the visit or review?
  2. How is usefulness measured?
  3. How are wrong statements reported and turned into evaluation cases?

13Decisions & outputs

What the work produces.

  1. 01Context contract
  2. 02Statement types & citation format
  3. 03Bounded read tools
  4. 04Briefing and review-page templates
  5. 05Ambiguity policy
  6. 06Evaluation set