Using AI to assist data quality without giving it master-data authority
A fictional agentic case: AI that makes data stewards faster without becoming the data steward.
Independently created. Contains no employer or client implementation detail, internal names or figures.
01The outcome
Stewards work a queue of suggestions ranked by impact, each with evidence; decisions are recorded with reason codes and feed back into the matcher; nothing changes in master data without a steward.
The CRM has probable duplicate accounts, inconsistent names and missing classifications—and a proposal to let a model fix them.
Where does AI help data quality—and why must merge, rename and reclassification stay with a person?
Give the agent recommend authority only: it suggests with evidence into a steward queue; merge, rename and reclassification remain steward actions executed by the system.
Years of manual entry left the CRM with near-duplicate accounts, names in several spellings and many accounts without an industry or segment classification. Some duplicates are linked to ERP customers, campaigns and open opportunities. The identity model (see Systems & CRM) says which key defines a customer; the data is not yet consistent with it.
02The reality · current state
What happens today—or in the pilot.
- Probable duplicates split pipeline and activity
- Names are spelled several ways, so searches miss accounts
- Many accounts have no classification, so segments are incomplete
- A model was proposed to merge duplicates automatically
- Merges cannot be cleanly undone once ERP links move
- Stewards have no queue—only spreadsheets
03Agent workflow
Each step with its actor, its tool and its authority.
- 01Scan
A scheduled scan selects accounts by rule: similar names, shared domains, missing fields.
- 02Suggest
Duplicate candidates, a normalized name, a probable classification.
- 03Evidence
Why: shared tax ID, same domain, same address, similar history—and what differs.
- 04Steward queue
Suggestions ranked by impact: open pipeline, ERP links, campaigns.
- 05Steward decision
Accept, edit or reject with a reason code.
- 06Apply
The system applies accepted changes through the governed merge or update process.
04Key dimension · Suggest with evidence, steward decides
What the agent may suggest—and what only a steward decides
Each finding type has a clear split between suggestion and decision, and the evidence that must accompany it.
| Finding | Agent may | Steward decides | Evidence shown |
|---|---|---|---|
| Duplicate accounts | Suggest candidate pairs with similarity evidence | Whether to merge, which record survives, what happens to ERP links | Tax ID, domain, address, contact overlap—and differences |
| Inconsistent names | Propose a normalized display name | The legal name always stays as registered; display names by steward rule | Variants found and their sources |
| Missing classification | Recommend an industry or segment with a reason | The classification that drives segments and routing | Products bought, website description, similar accounts |
| Incomplete descriptions | Draft a description from public and internal sources | Whether it is accurate enough to save | Cited sources for every claim |
AI cannot fix unreliable identity. It can make the steward’s queue shorter and better evidenced; the key and the decision stay with people.
06Ambiguity policy
When the agent is unsure.
- Candidate linked to different ERP customersEscalate
Flagged high impact; steward and finance decide together
- Weak evidence only (name similarity)Stop
Not suggested—noise erodes steward trust
- Classification sources disagreeContinue
Suggested with both reasons shown
07Evaluation
Scored per dimension—never one accuracy number.
| Dimension | Question | Measure | Target |
|---|---|---|---|
| Precision | Are suggested duplicates real? | Accepted ÷ suggested | ≥ 85 % |
| Evidence | Does each suggestion show why—and what differs? | Suggestions with complete evidence | 100 % |
| Impact ranking | Do high-impact cases come first? | Share of accepted suggestions in the top of the queue | Tracked |
| Authority | No master-data change without a steward? | Unapproved changes | 0 — release gate |
08The trade-offs
Credible options, judged against these premises.
Model merges duplicates above a threshold
None for customer master data
Cost: Irreversible changes across CRM and ERPRules-only matching
Clean keys and consistent entry
Cost: Misses the fuzzy cases that matterAI suggestions with evidence into a steward queue
Messy data with a defined identity model
Cost: Steward capacity and a queue to run09The second layer
Questions that change the design.
Authority
- Who owns the decision to merge?
- Which fields may never be changed by suggestion?
- How are ERP-linked records handled?
Evidence
- What evidence is required for a suggestion?
- How are differences shown?
- When is a suggestion too weak to show?
Learning
- How do rejections improve the matcher?
- How is steward capacity planned?
- How is data quality measured over time?
10Decisions & outputs
What the work produces.
- 01Finding types and evidence rules
- 02Steward queue and ranking
- 03Authority split
- 04Reason codes
- 05Evaluation set
- 06Data-quality trend