Automation case / 01 · Run record, state machine & safe resume

Designing a monthly commercial performance orchestrator

A fictional automation case: turning one big scheduled job into an observable, resumable monthly run.

Fictional scenario

Independently created. Contains no employer or client implementation detail, internal names or figures.

01The outcome

One run per period that anyone can open to see its state, current step, progress, failures and owner—resumable from the failed step, and finished only when the result is reconciled.

The reality

A monthly process with eight dependent steps runs as one scheduled job—so any failure means starting again, and some accounts get their actions twice.

Architecture question

How is one monthly run identified, resumed and finished—exactly once?

Key decision

Model the cycle as a run record with a state machine and step-level guards, rather than one scheduled job or a chain of triggers.

ContextEvery month a commercial cycle must refresh performance data, compute KPIs, classify customers into segments, activate those segments in the CRM and campaigns, generate follow-up actions and confirm the result. The first version was one scheduled job that called every step in sequence. When it failed half-way, an administrator restarted it from the beginning.

Trigger

The monthly period opens and the data refresh for it has completed.

Done means
  • Every eligible account has a current segment for the period
  • Actions exist once per account and segment
  • The run is Success with a reconciliation report
  • The business owner has the summary
Systems and partiesSchedulerData platformCRMCampaign platform

02The reality · current state

What the automation does today.

  • Nobody can say which month a run belongs to, or whether it finished
  • A restart after a failure repeats steps that already succeeded
  • Accounts receive duplicate tasks when action generation runs twice
  • The job starts before the data refresh has finished
  • Failures are visible only in a technical log
  • A test run and a production run look identical

03Target steps

Each step with its mode, its actor and its guard.

  1. 01
    Data refresh

    Wait for the period’s refresh to complete; it is a dependency, not a step we run.

    OrchestratedData platformGuardRefresh complete for the period
  2. 02
    KPI computation

    Compute performance to target for the period.

    AutomatedData platformGuardKPIs not yet computed for run key
  3. 03
    Customer classification

    Exactly one segment per eligible account.

    AutomatedData platformGuardPopulation count = classified count
  4. 04
    Segment activation

    Publish the period’s segments as a new version.

    AutomatedData platformGuardVersion not yet published
  5. 05
    CRM writeback

    Upsert in batches with continuations; resumes from the last completed batch.

    OrchestratedIntegration → CRMGuardBatch cursor
  6. 06
    Action generation

    Create tasks and campaign memberships once per account and segment.

    AutomatedCRM automationGuardAccount + segment + entry period key
  7. 07
    Reconciliation

    Compare expected with actual segments and actions.

    AutomatedReconcilerGuardDifferences classified
  8. 08
    Finalization

    Review the summary and close the run—or abort it with a reason.

    HumanCommercial operationsGuardReconciliation within tolerance

04Key dimension · Run record, state machine & safe resume

One run, eight steps, one state machine

The run record is the orchestrator’s memory: which period, which step, which batch, what failed, who owns it.

Run path
Failure
Recovery
Run path

Running

A step or batch is executing.

Can move to
  • WaitingAwaiting a dependency or the owner’s reviewSystem
  • FailedStep failed beyond its retry budgetSystem

Entered from Created, Waiting, Resume.

Steps, their dependencies, guards and failure behaviour
StepDepends onGuardOn failure
1 · Data refreshExternal refreshWait with a budgetWaiting → Failed after budget; owner: data team
2 · KPI computationStep 1Once per run keyRetry; then Failed
3 · ClassificationStep 2Population = classifiedFail before publishing
4 · Segment activationStep 3Version published onceResume re-publishes the same version
5 · CRM writebackStep 4Batch cursorResume from the last completed batch
6 · Action generationStep 5Account + segment + entry periodRerun creates nothing twice
7 · ReconciliationStep 6Once per runDifferences go to their owners
8 · FinalizationStep 7Tolerance metOwner decides: close or abort

Can step 5 start twice? No: the step guard sees it already running for this run, and each batch is keyed by the cursor.

05Execution safety

If it can run twice, it is designed to run twice.

Execution patternScheduled run record + state machine + queue with continuations

The cycle outlives one transaction, depends on a refresh it does not control, and must resume at the step that failed.

Run key
Period + scope (e.g. 2026-09 · all companies)
Start guard
A second run for the same key is refused atomically
Step guard
Each step checks its own “already done for this run” state
Effect keys
Account + segment + entry period for every task and membership
Dry run
Same run, same checks, no side effects; produces the would-be report
Abort
A terminal state with a reason; completed steps stay recorded

06Exceptions & recovery

Not every failure is a retry.

  • Refresh late DependencyDependency not complete at startWait with a budget, then escalateData team
  • CRM limit reached mid-writeback Transient · retryLimit error on a batchBack off; continue from the cursorAutomation
  • Run started twice DuplicateRun key already existsRefuse the second startNobody: absorbed
  • Account without owner ConfigurationOwner lookup empty in action generationHolding queue; continue the restCommercial operations
  • Classification incomplete ValidationPopulation ≠ classified countFail before activationSales operations

07The trade-offs

Credible options, judged against these premises.

Rejected

One scheduled job calling every step

Short, local, idempotent work

Cost: Restart from zero; duplicates; no visibility
Rejected

A chain of triggers, one per step

Two or three local steps

Cost: No run identity; failures break the chain silently
Selected

Run record, state machine, guards and continuations

Long, dependent, recurring work

Cost: A run model and a small run view to build

08The second layer

Questions that change the design.

Identity & start

  1. How is one monthly run identified?
  2. How is a duplicate start prevented?
  3. How are dependencies represented?

Failure & resume

  1. What happens when step 4 fails?
  2. Can step 5 start twice?
  3. How does the process resume?

Completion & control

  1. When is the run considered complete?
  2. How does a dry run differ from production?
  3. How is abort represented—and who sees run status?

09Decisions & outputs

What the work produces.

  1. 01Run model & run key
  2. 02State machine
  3. 03Step dependency & guard table
  4. 04Continuation design
  5. 05Dry-run and abort semantics
  6. 06Run status view