AI case / 05 · Suggest with evidence, steward decides

Using AI to assist data quality without giving it master-data authority

A fictional agentic case: AI that makes data stewards faster without becoming the data steward.

Fictional scenario

Independently created. Contains no employer or client implementation detail, internal names or figures.

01The outcome

Stewards work a queue of suggestions ranked by impact, each with evidence; decisions are recorded with reason codes and feed back into the matcher; nothing changes in master data without a steward.

The reality

The CRM has probable duplicate accounts, inconsistent names and missing classifications—and a proposal to let a model fix them.

Architecture question

Where does AI help data quality—and why must merge, rename and reclassification stay with a person?

Key decision

Give the agent recommend authority only: it suggests with evidence into a steward queue; merge, rename and reclassification remain steward actions executed by the system.

ContextYears of manual entry left the CRM with near-duplicate accounts, names in several spellings and many accounts without an industry or segment classification. Some duplicates are linked to ERP customers, campaigns and open opportunities. The identity model (see Systems & CRM) says which key defines a customer; the data is not yet consistent with it.

Systems and partiesCRMERPData platformSteward queue

02The reality · current state

What happens today—or in the pilot.

  • Probable duplicates split pipeline and activity
  • Names are spelled several ways, so searches miss accounts
  • Many accounts have no classification, so segments are incomplete
  • A model was proposed to merge duplicates automatically
  • Merges cannot be cleanly undone once ERP links move
  • Stewards have no queue—only spreadsheets

03Agent workflow

Each step with its actor, its tool and its authority.

  1. 01
    Scan

    A scheduled scan selects accounts by rule: similar names, shared domains, missing fields.

    SystemRead
  2. 02
    Suggest

    Duplicate candidates, a normalized name, a probable classification.

    AgentRecommend
  3. 03
    Evidence

    Why: shared tax ID, same domain, same address, similar history—and what differs.

    AgentRecommend
  4. 04
    Steward queue

    Suggestions ranked by impact: open pipeline, ERP links, campaigns.

    SystemExecutecreateStewardTask
  5. 05
    Steward decision

    Accept, edit or reject with a reason code.

    HumanHuman approval
  6. 06
    Apply

    The system applies accepted changes through the governed merge or update process.

    SystemExecute

04Key dimension · Suggest with evidence, steward decides

What the agent may suggest—and what only a steward decides

Each finding type has a clear split between suggestion and decision, and the evidence that must accompany it.

Data-quality findings: what the agent may suggest, what a steward decides, and the evidence required
FindingAgent maySteward decidesEvidence shown
Duplicate accountsSuggest candidate pairs with similarity evidenceWhether to merge, which record survives, what happens to ERP linksTax ID, domain, address, contact overlap—and differences
Inconsistent namesPropose a normalized display nameThe legal name always stays as registered; display names by steward ruleVariants found and their sources
Missing classificationRecommend an industry or segment with a reasonThe classification that drives segments and routingProducts bought, website description, similar accounts
Incomplete descriptionsDraft a description from public and internal sourcesWhether it is accurate enough to saveCited sources for every claim

AI cannot fix unreliable identity. It can make the steward’s queue shorter and better evidenced; the key and the decision stay with people.

05Authority

What the agent may do—and what it may not.

Actions and the authority the agent has for each
ActionAuthorityWhy
Read accounts, contacts and linksReadNeeded to find and evidence candidates
Suggest duplicates, names, classificationsRecommendInto a steward queue with evidence
Create steward tasksExecuteReversible, owned work items
Merge accountsForbiddenMaster-data authority; rarely reversible
Change legal name or classificationForbiddenDrives identity, segments and routing

06Ambiguity policy

When the agent is unsure.

  • Candidate linked to different ERP customersEscalate

    Flagged high impact; steward and finance decide together

  • Weak evidence only (name similarity)Stop

    Not suggested—noise erodes steward trust

  • Classification sources disagreeContinue

    Suggested with both reasons shown

07Evaluation

Scored per dimension—never one accuracy number.

Evaluation dimensions: question, measure and target
DimensionQuestionMeasureTarget
PrecisionAre suggested duplicates real?Accepted ÷ suggested≥ 85 %
EvidenceDoes each suggestion show why—and what differs?Suggestions with complete evidence100 %
Impact rankingDo high-impact cases come first?Share of accepted suggestions in the top of the queueTracked
AuthorityNo master-data change without a steward?Unapproved changes0 — release gate

08The trade-offs

Credible options, judged against these premises.

Rejected

Model merges duplicates above a threshold

None for customer master data

Cost: Irreversible changes across CRM and ERP
Situational

Rules-only matching

Clean keys and consistent entry

Cost: Misses the fuzzy cases that matter
Selected

AI suggestions with evidence into a steward queue

Messy data with a defined identity model

Cost: Steward capacity and a queue to run

09The second layer

Questions that change the design.

Authority

  1. Who owns the decision to merge?
  2. Which fields may never be changed by suggestion?
  3. How are ERP-linked records handled?

Evidence

  1. What evidence is required for a suggestion?
  2. How are differences shown?
  3. When is a suggestion too weak to show?

Learning

  1. How do rejections improve the matcher?
  2. How is steward capacity planned?
  3. How is data quality measured over time?

10Decisions & outputs

What the work produces.

  1. 01Finding types and evidence rules
  2. 02Steward queue and ranking
  3. 03Authority split
  4. 04Reason codes
  5. 05Evaluation set
  6. 06Data-quality trend