Designing a reconciler for stale downstream state
A fictional automation case: the execution side of reconciliation. The comparison model is in the Data & Integrations cases; this one is about acting on it safely.
Independently created. Contains no employer or client implementation detail, internal names or figures.
01The outcome
Downstream state converges on expected state every day; automated closes stay within guardrails; anything unusual goes to a person first.
Upstream eligibility changes, but campaign memberships, tasks and recovery processes stay behind—and cleaning them up is a quarterly manual job.
How does an automation close what should no longer exist—safely, and without closing real work?
Run reconciliation as a scheduled automation with classified actions and guardrails, rather than as a periodic manual cleanup.
Accounts move in and out of segments and eligibility every month. Entry automations work; exits are sometimes missed—an outage, a partial run, a rule change. Stale campaign memberships and tasks accumulate until someone cleans them up by hand.
Daily, after the eligibility and segment publication for the day.
- Stale memberships exited
- Stale tasks closed with a reason—unless protected
- Unusual volumes held for review
- A drift report for the owner
02The reality · current state
What the automation does today.
- Stale campaign memberships keep sending the wrong messages
- Open tasks for accounts that recovered clutter owners’ lists
- Recovery plans stay open for ineligible accounts
- A bulk cleanup once closed tasks that sellers were working
- Nobody knows how much drift exists at any time
03Target steps
Each step with its mode, its actor and its guard.
- 01Derive expected state
From current segments and eligibility.
- 02Compare with actual
Memberships, tasks and plans by account.
- 03Classify differences
Match, missing, stale (unexpected), conflicting.
- 04Act within guardrails
Close, exit or create—keyed and limited.
- 05Owner review
Held items and the drift report.
04Key dimension · Automated cleanup with guardrails
Expected against actual—then act, within limits
The reconciler classifies every difference and acts only where the rule is clear and the change is small and reversible.
| Key | Expected downstream | Actual downstream | Outcome |
|---|---|---|---|
ACC-A12 | Retention campaign | Retention campaign | MatchBoth sides agree. |
ACC-B07 | Recovery task open | — | MissingEntry automation failed during an outage. |
ACC-C33 | No campaign | Retention campaign | UnexpectedAccount recovered; exit never ran. |
ACC-D19 | Not eligible | Recovery plan open | UnexpectedAccount left the population. |
ACC-E02 | Review call | Recovery call (in progress) | DifferenceSegment improved while the seller was working the task. |
- Match
Downstream matches expected.
- Owner
- Nobody
- Action
- Record the check.
- Difference
Present, but not what the segment requires.
- Owner
- Account owner
- Action
- Protected if in progress: suggest, don’t close.
- Missing
Expected, not present.
- Owner
- Automation
- Action
- Create with the same key the entry would have used.
- Unexpected
Present, no longer expected.
- Owner
- Automation
- Action
- Close or exit with reason and run ID—within the blast-radius limit.
- Dry run first
New rules report what they would do for a period before acting.
- Blast radius
An unusual share of closes holds the run for review.
- Protected work
Anything in progress or recently touched is suggested, never closed.
- Reversibility
Every automated close carries a reason and the run ID.
A reconciler that can close anything will one day close everything. The guardrails are the design.
05Execution safety
If it can run twice, it is designed to run twice.
Cleanup touches many records in several systems; it must be observable, limited and reversible.
- Dry run
- Default for new rules: report only
- Blast radius
- More than a set share of a population → hold for review
- Protected records
- Tasks in progress or touched in the last days are never auto-closed
- Action keys
- Account + action + run date
- Reversible
- Closed with a reason and run ID, so they can be reopened
06Exceptions & recovery
Not every failure is a retry.
- Segment publication late DependencyExpected state not readySkip the run; no comparison against stale expectationsData team
- Unusual close volume ValidationBlast-radius limit exceededHold the run for reviewCommercial operations
- Campaign platform rate limit Transient · retryLimit errors on exitsBack off; continue next batchAutomation
- Rule misconfigured ConfigurationDry-run report looks wrongKeep the rule in dry runRule owner
07The trade-offs
Credible options, judged against these premises.
Quarterly manual cleanup
Tiny populations
Cost: Drift for months; risky bulk editsFix every entry automation instead
When outages never happen
Cost: Exits still get missed; no measure of driftDaily reconciler with guardrails
State that drives work in several systems
Cost: Guardrail thresholds need an owner08The second layer
Questions that change the design.
Expected state
- Where does expected state come from?
- What if it is late or incomplete?
- Which downstream records are in scope?
Acting safely
- What may the reconciler close without asking?
- What is protected?
- How large a change is too large?
Ownership
- Who reviews held items?
- How is an automated close undone?
- How is drift reported over time?
09Decisions & outputs
What the work produces.
- 01Expected-state derivation
- 02Difference classes and actions
- 03Guardrails
- 04Dry-run mode
- 05Drift report