Skip to content

ADR-0013 — Repository Reset State Machine

Context

Repository Reset (the org-wide effort to collapse dozens of stray branches, 148+ open MRs, and drifted release lines back to a stable integration lane) repeatedly stalled on the same failure mode: repositories stayed perpetually "in progress" with no way to tell who was supposed to act next. Reports conflated several different questions — is this repo still in scope, what workflow position is it at, who currently owns the next action, and is it actually ready to be worked given its dependencies — into a single ad hoc status string. That made "blocked" ambiguous (blocked forever, or blocked pending one specific person's action, or simply not yet reachable in dependency order?) and let agents either stall indefinitely or advance a stage that wasn't theirs to advance. A further conflation surfaced once the Stage/Owner split was in place: "ready to work" (a dependency-graph fact) kept getting mixed with "executable right now" (an authority fact) — a repository can be fully unblocked in the dependency DAG while still correctly non-executable by Engineering because ownership has transferred elsewhere.

Decision

Model Repository Reset as three orthogonal systems, each answering exactly one question, with no system reaching into another's state:

                  Fleet Scheduler (DAG)
              answers: READY = dependencies == COMPLETE
                           │
                           ▼
               Repository State Machine
        answers: current Stage, current Owner (derived)
 Not Started → Engineering → Awaiting Operator →
   Awaiting Pipeline → Awaiting Governance → Complete
                           ▲
                           │
                    Authority Model
        answers: can the current Owner execute
                 the next transition?
     Engineering │ Operator │ Pipeline │ Governance

Fleet Scheduler (DAG). Answers only: is this repository's dependency set COMPLETE? It has no knowledge of git-write access, CI, approvals, hooks, or authentication — those are not its concern.

Repository State Machine. Tracks three independent dimensions instead of one status field: - Lifecycle (is it still in scope): ACTIVE · RETIRED · ARCHIVED - Stage (workflow position, exactly one at a time): Not Started → Engineering → Awaiting Operator → Awaiting Pipeline → Awaiting Governance → Complete - Owner (who may act next) is derived from Stage, not tracked separately.

Authority Model. Answers only: can the current Owner execute the next transition right now? (E.g. Engineering has no authenticated session → Engineering-executable = NO, independent of whether the repo is DAG-READY or what Stage it's in.)

Rule

  1. Stage transitions require recorded evidence, never inference:
Transition Required evidence New owner
Engineering → Awaiting Operator Commit SHA recorded (or explicit operator gate if engineering cannot commit) Operator
Awaiting Operator → Awaiting Pipeline Push or MR SHA recorded Pipeline
Awaiting Pipeline → Awaiting Governance Canonical validation + green pipeline recorded Governance
Awaiting Governance → Complete release→main merge SHA recorded None
  1. Ownership invariant: only the current authority may advance a repository's stage. Engineering cannot advance Pipeline. Pipeline cannot advance Governance. Governance cannot advance Engineering. Advancing a stage without its required evidence is a governance violation of the same class as bypassing a single-authority rule elsewhere in this standard (e.g. authentication) — a plausible-sounding shortcut overriding a rule with no exceptions.

  2. Readiness and executability are independent facts — the top invariant of this model. A repository may be READY in the scheduler (all dependencies COMPLETE) while simultaneously non-executable by Engineering, because readiness is determined only by dependencies and executability is determined only by current authority. Neither system may infer the other's answer.

  3. A blocked authority is a transfer, not a global halt. When engineering cannot proceed (e.g. no authenticated session to push/merge with), the correct record is not "blocked" in the abstract — it is: current Stage, current Owner, the transition reason, and the specific wake condition that returns ownership. The repository continues to occupy a well-defined position in the state machine; it does not become an undefined "stuck" repo.

  4. Fleet metrics are three independent aggregations, never maintained independently of the underlying tables and never merged into one composite score:

  5. Scheduler: count of READY repositories, count blocked by dependency
  6. State Machine: count per Stage (Engineering / Awaiting Operator / Awaiting Pipeline / Awaiting Governance / Complete)
  7. Authority: count Engineering-executable, count Operator-owned, count Pipeline-owned, count Governance-owned

A narrative summary is not a substitute for these tables.

  1. Work never disappears — it completes or transfers authority. There is no terminal state other than Complete (or Retired/Archived at the Lifecycle level) that represents an agent simply stopping.

Consequences

  • Every repository occupies exactly one Stage with exactly one implied Owner at any time; "blocked" is replaced by "Stage X, Owner Y, waiting on evidence Z."
  • "Ready to work" and "executable right now" are never conflated again: a READY repository sitting in Awaiting Operator is correctly reported as non-executable by Engineering, not as a false blocker or a false green light.
  • Dashboards and fleet-health reports become deterministic aggregations instead of narrated status, matching the Evaluator Pattern's separation of evidence-collection from decision-making (ADR-0005).
  • Agents working Repository Reset must record the transition evidence (SHA, pipeline status, merge SHA) at each stage change, not just narrate progress.
  • This ADR formalizes a methodology that was previously tracked only in conversational/session memory; memory remains a convenience index pointing here, not the canonical source.