Catalyst reads two config layers. The setup script (setup-catalyst.sh) writes both for you, so you
rarely edit them by hand. This page covers the keys you’re most likely to touch.
.catalyst/config.json — plain project info. Safe to commit to git.
~/.config/catalyst/ — machine-local Layer-2. Never commit these files. Three siblings
compose the merged view (earlier wins):
cluster-secrets.json — shared, byte-identical across all cluster nodes (bot OAuth creds,
smeeChannel, livenessAnchorIssue, etc.). Written by cluster-sync from cluster-bots.sops.json
and by catalyst-join from the bundle. Read first.
node.json — per-node (host name, cloudFeed/githubFeed mode, orchestration tuning, etc.).
Written by catalyst-join (non-clobber; preserves operator overrides). Read second.
config.json — backward-compat fallback. Read last. Keys still here on nodes that haven’t
run catalyst-join or cluster-sync yet; migrated keys are gradually drained by those tools.
The projectKey links Layer-1 and the per-project config-{projectKey}.json (legacy file for
per-team API keys; separate from the three Layer-2 siblings above).
As work moves, Catalyst updates the ticket’s Linear status for you. stateMap says which status
name to use for each step (research, inProgress, inReview, done, and so on). Set a key to
null to skip that update.
You usually don’t edit this by hand. When you run setup-catalyst.sh with a Linear token, it reads
your real status names and fills stateMap in. Pointing stateMap at a status that doesn’t exist
makes the next update fail, so only edit it if your status names are unusual.
Agent notes — research, plans, handoffs — are written to a git-backed thoughts repo, one per
GitHub organization. Three keys, and the first two are easy to confuse:
HumanLayer profile name — the key under .thoughts.profiles in humanlayer.json
directory
Subdirectory inside the thoughts repo for this project’s notes (defaults to the repo’s own basename)
user
Username passed to the thoughts CLI; null uses the CLI default
org and profile are different things and need not match. org is a GitHub owner; profile
is a local alias. Nor does org have to match the org of the project’s own code repo — the example
above is the real shape: code under github.com/groundworkapp/Adva, thoughts under
github.com/rightsite-cloud/thoughts, reached through the HumanLayer profile adva.
provision-thoughts.sh reads org from each registered project’s own .catalyst/config.json. If
org is absent it falls back to profile, then to the org segment of the project’s checkout path —
each fallback prints a WARN naming the project, because both are guesses that are only right when
the names happen to coincide. Set org explicitly on every project; a missing org degrades loudly
but never aborts a node join.
The orchestration.dispatchMode key picks how Catalyst runs each ticket:
execution-core — the autonomous daemon. It watches your board, picks up ready tickets, and
runs them with no command from you. This is the away-from-keyboard mode.
phase-agents — runs each ticket as ten short background jobs, one per step.
oneshot-legacy — one long-running job per ticket. The older default.
Which substrate runs a phase worker: bg (a claude --bg background job, today’s behavior), oneshot-legacy, or sdk (the in-process Claude Agent SDK — not yet implemented; falls back to bg, CTL-1365b). Resolution: CATALYST_EXECUTOR env → this key → node-class default (all classes → bg today). Distinct from dispatchMode.
orchestration.maxParallel
3
How many tickets run at once
orchestration.worktreeDir
~/catalyst/wt/<projectKey>
Where worktrees are created
orchestration.pluginDirs
unset
Path(s) to the plugin checkout(s) workers run from (<checkout>/plugins/dev). Set by setup-plugin-source.sh; resolved by phase-agent-dispatch and refreshed by catalyst-stack hotpatch / merge-to-main. String or :-joined array. May also live in the machine config (Layer 2); the CATALYST_PLUGIN_DIRS env var overrides both.
Which process pulls plugin checkouts current on this node: broker (the drift-check timer does git fetch + reset --hard, today’s default) or updater (the standalone catalyst-updater LaunchAgent owns the pull; the broker keeps drift detection + lag alerting but skips the pull — detect-only). Read fresh each broker tick; env override CATALYST_PLUGIN_PULL_OWNER. Flipped to updater only by the supervised catalyst-stack adopt-updater cutover, which first confirms the updater agent is running (fail-closed to broker). A daemonless developer node runs no broker, so the updater is the only thing keeping it fresh. (CTL-1348)
orchestration.phaseAgents.models[phase]
opus
Model per step (opus, sonnet, or haiku). Phases: triage, research, plan, implement, verify, review, pr, monitor-merge, monitor-deploy, teardown
orchestration.phaseAgents.turnCaps[phase]
per-phase
Max Claude turns per step
orchestration.draftPr.enabled
true
Open a draft PR at the first implement commit; phase-pr flips it ready. Set false to create the PR only at the pr phase.
CATALYST_WORKFLOW_GITHUB_TOKEN(env var, never committed)
unset
A GitHub PAT with the workflow OAuth scope. When set, phase-pr automatically routes pushes that touch .github/workflows/ through this token instead of the ambient GITHUB_TOKEN (which lacks workflow scope). When unset and such a push is attempted, phase-pr escalates with an actionable human_question telling the operator to grant the scope or push manually. Provision via the daemon launch environment or ~/.config/catalyst/config-<projectKey>.json. Alternative: gh auth refresh -s workflow re-auths the host token.
orchestration.stalePrRescue.enabled
true
Periodically rescue orphaned PRs that drifted to DIRTY or BEHIND after their workers died.
orchestration.stalePrRescue.intervalSeconds
600
How often the rescue timer ticks (seconds).
orchestration.stalePrRescue.stableSeconds
300
How long a PR must sit DIRTY/BEHIND before a rescue is attempted (avoids reacting to transient states).
orchestration.stalePrRescue.behindThreshold
10
BEHIND-commit count that triggers a rebase rescue (commits-behind below this are skipped).
orchestration.stalePrRescue.maxAttempts
1
Max rescue attempts per ticket. After exhaustion, the ticket is escalated to needs-human.
orchestration.stalePrRescue.maxConflictFiles
5
Max conflicting files before a DIRTY PR is deemed unresolvable and escalated instead of dispatched.
orchestration.orphanPrSweep.enabled
true
Periodically scan all open PRs in the configured repo for ones that have no pipeline worker (orphans). When an orphan has been in a blocker state (DIRTY/BLOCKED/UNSTABLE) for stableSeconds, raises exactly one Needs-You inbox row.
orchestration.orphanPrSweep.intervalSeconds
600
How often the orphan-PR sweep ticks (seconds).
orchestration.orphanPrSweep.stableSeconds
300
How long an orphan must hold a blocker state before a Needs-You row is raised.
orchestration.orphanPrSweep.repo
(auto-detected)
The org/repo slug to pass to gh pr list. Falls back to top-level .catalyst/config.json repo fields, then gh repo view. Set this explicitly when auto-detection is unreliable.
orchestration.stalledPrSweep.enabled
false
Periodically sweep all in-flight worker PRs for review-latency, CI-health, and no-push signals independent of worker liveness (CTL-1608). Default-off — enable only after validating thresholds on the live board. When enabled the timer writes workers/<TICKET>/stalled-pr.json; board-health reads those stamps via getStalledPrState and emits nudge-stalled-pr moves.
orchestration.stalledPrSweep.intervalSeconds
900
How often the stalled-PR sweep ticks (seconds). Configurable per the CATALYST_BH_STALLED_PR_* env thresholds below.
orchestration.githubQuotaSweep.enabled
true
Sample the host’s GitHub core REST quota and atomically publish it to <orchDir>/github-quota.json. Set false to disable the timer; a previous snapshot may remain on disk but becomes stale and cannot arm board-health.
orchestration.githubQuotaSweep.intervalSeconds
300
How often the daemon runs the quota sampler (seconds). The sampler calls the quota-reporting endpoint, which does not consume the core quota it reports.
responder.intervalSeconds
180
How often the daemon-health responder launchd sweep runs (seconds, clamped 60–900). The responder (health-responder.sh, CTL-1509) detects a dead/stale cloud-sync replica writer and issues bounded launchctl kickstarts, escalating after the attempt cap. Baked into the launchd plist at install time (install-health-responder.sh); re-run catalyst-stack install-services after changing it.
CATALYST_CLAUDE_UPDATE_MIN_INTERVAL_MS(env var)
21600000
Minimum gap between claude update invocations (milliseconds; default 6 hours). The claude-selfupdate.sh LaunchAgent installed by catalyst-stack install-services honours this via a durable marker at ~/catalyst/.claude-selfupdate.last. Reduce to 0 to force a run on the next LaunchAgent tick; increase to spread out update checks on resource-constrained hosts. The LaunchAgent itself fires every 6 hours (StartInterval: 21600) and is installed on all node classes. (CTL-2085)
orchestration.reconcile.mode
off
Completion-declaration reconcile timer (CTL-1371). Linear state is driven by explicit completion declarations — the model/pipeline/human says “this is done” via catalyst-linear-reconcile declare <TICKET> — never inferred from PR/merge state (a draft PR opens while work is in progress; a merged PR is not yet Done — the pipeline puts deploy-verification + teardown between merge and Done). The timer drains pending declarations and makes Linear reflect them, retrying any write that didn’t land. off = inert (also the default); notify = compute drift + emit ticket.completion.drift.<ticket> events but never write (safe first-ship); write = write the declared state via the canonical primitive. Runs on the daemon event loop, separate from the dispatch scheduler. Idempotent + CTL-758 backward-write guard (never resurrects a Canceled ticket, never regresses a Done one).
orchestration.reconcile.intervalSeconds
600
How often the drain timer ticks (seconds).
EXECUTION_CORE_RECONCILE_BLIND_ALERT_MS
300000
Time-based total-board-blindness alarm. When every registered team’s reconcile is failing and no team has succeeded inside this window, execution-core raises the fleet admission alert and emits fleet.health.degraded. This complements EXECUTION_CORE_RECONCILE_FAILURE_ALERT_THRESHOLD; it does not replace the count-based per-team latch.
orchestration.reconcile.declarationsDir
~/catalyst/completions
Directory holding the durable per-ticket completion markers (<TICKET>.json). Overridable via CATALYST_COMPLETIONS_DIR.
orchestration.orphanReaper.jobGc.enabled
true
Enable periodic GC of stale ~/.claude/jobs/<id> dirs (CTL-1165 D3). Set false to disable.
orchestration.orphanReaper.jobGc.retentionSeconds
86400
Delete a job dir only if its mtime is older than this many seconds (default 24 h). Env CATALYST_JOB_GC_RETENTION_SECONDS overrides.
orchestration.orphanReaper.jobGc.batchCap
200
Max dirs deleted per sweep tick. Remaining dirs drain on subsequent ticks. Env CATALYST_JOB_GC_BATCH_CAP overrides.
orchestration.orphanReaper.workerGc.enabled
true
Enable periodic GC of stale execution-core/workers/<TICKET>/ dirs (CTL-1205). Set false to disable.
Grace window before a worker dir carrying ZERO phase signals is treated as reclaimable residue rather than phase-agent-dispatch’s mkdir→first-signal window (CAT-24). Below it the scheduler still counts the ticket as started; above it the ticket is re-pulled as new work. A zero-signal dir preserving a non-empty inbox.jsonl is re-pullable immediately regardless. Env CATALYST_EMPTY_WORKER_DIR_GRACE_MS (milliseconds) overrides.
orchestration.orphanReaper.procReaper.mode
shadow
Orphan child-process reaper mode. off disables it; shadow (the default) logs procOrphans.would-reap for each candidate but kills nothing; enforce actually SIGTERM→grace→SIGKILLs them. Candidates are (a) orphaned reparented node/bun grandchildren a dead worker left behind, and (b) since CTL-1531, an orphan of any command that satisfies the ownership conjunction ppid == 1and cwd under worktreeRootand that cwd definitely no longer exists (the sh -c runaway class) — that widened class is gated by its own widenMode knob and is not armed by setting this one to enforce. Ships in shadow so the never-kill allowlist + live-agent process-tree correlation — and the widened class in particular — can be audited on real hosts before any enforce flip.
orchestration.orphanReaper.procReaper.widenMode
shadow
Independent rollout mode for the CTL-1531 widened (any-command) orphan class, deliberately NOT derived from mode. off removes the widened admission entirely (a byte-identical revert of the feature); shadow (the default) classifies and reports widened candidates via procOrphans.would-reap but never signals them, even when mode is already enforce; enforce lets a widened candidate follow mode, so BOTH knobs must be open before an arbitrary PPID-1 command is signalled. An unrecognized value degrades to shadow, never to enforce. This exists because a host that already carries mode: "enforce" — granted for the narrow node/bun class after its shadow bake — must not inherit authority over the widened class on deploy; ADR-023 requires a per-actuator shadow window and an operator-owned flip. Mirrors orphan-sweep.sh’s SWEEP_PROC_WIDEN.
Per-run cap on confirmed terminations from the widened class (0 = uncapped), mirroring orphan-sweep.sh’s SWEEP_PROC_WIDEN_MAX_KILLS and vector 2’s SWEEP_MAX_REMOVALS. Delivered signals carry a second ceiling of widenMaxKills × 2, since a candidate is worth at most SIGTERM + SIGKILL; counting confirmed exits (not attempts) against the first ceiling is what stops a process that ignores SIGTERM from consuming a slot forever. The widened class’s authorizing evidence — “this cwd no longer exists” — is correlated across a host, so a run that wants to kill more than a handful is a root-level event that wants a human. A non-numeric or non-integer value degrades to 5, never to uncapped — a fractional cap such as 0.5 used to floor to 0, which is the documented uncapped value, so a config typo silently removed the ceiling.
orchestration.orphanReaper.procReaper.graceMs
5000
Milliseconds to wait after SIGTERM before re-probing and (only if still alive) SIGKILLing, so node/bun can flush.
orchestration.orphanReaper.procReaper.minEtimeSec
900
A process must have run at least this long (elapsed time) before it is eligible — corroboration only, never the sole gate.
Extra case-insensitive argv substrings to never kill, on top of the built-in allowlist (the daemon, broker/index.mjs, orch-monitor/server.ts, the entire live-agent process tree, Tailscale, pid 1, and any foreign-uid process).
orchestration.fleetHealth.enabled
true
Whether the pre-exhaustion fleet-health probe runs. Set false (or CATALYST_FLEET_HEALTH=0) to disable it entirely.
orchestration.fleetHealth.intervalMs
120000
How often the probe samples the four steady-state signals (milliseconds).
orchestration.fleetHealth.jobsThreshold
500
~/.claude/jobs dir count at or above which the jobs signal trips.
orchestration.fleetHealth.agentsThreshold
12
Live background-agent count at or above which the agents signal trips.
orchestration.fleetHealth.procsThreshold
40
Resident node/bun worker-process count at or above which the procs signal trips.
orchestration.fleetHealth.swapUsedMbThreshold
24576
macOS swap-used MB at or above which the swap signal TRIPS (edge-triggered since CTL-1503). Raised above a 16 GB Mac’s normal-swap ceiling so it stops firing every tick. Env EXECUTION_CORE_FLEET_SWAP_MB_THRESHOLD.
CTL-1503 hysteresis band — the latched swap degradation CLEARS (fires fleet.health.recovered once) only when swap drops strictly below this LOWER threshold, so a value hovering in [clear, trip) can’t re-flap. Clamped below swapUsedMbThreshold if misconfigured. Env EXECUTION_CORE_FLEET_SWAP_MB_CLEAR_THRESHOLD.
orchestration.fleetHealth.selfHealEnabled
false
Whether a sustained breach triggers self-heal (the two orphan-reaper intents plus a bounded ppid==1node/bun child sweep). Default OFF — the first ship is a pure alert. Enable with EXECUTION_CORE_FLEET_SELF_HEAL=1.
orchestration.fleetHealth.sustainedTicks
2
Consecutive degraded ticks required before self-heal fires (once per breach episode; re-armed only after a healthy tick).
orchestration.daemonWatchdog.mode
shadow
Stuck-but-alive daemon watchdog (CTL-1502). off disables; shadow detects + logs would-restart but never restarts; enforce restarts a stuck daemon (otel-forward) once per breach episode. Precedence: CATALYST_DAEMON_WATCHDOG=0 (kill-switch → off) → EXECUTION_CORE_DAEMON_WATCHDOG_MODE → this key → shadow.
orchestration.daemonWatchdog.intervalMs
120000
How often the probe checks the stuck predicates (ms). Env EXECUTION_CORE_DAEMON_WATCHDOG_INTERVAL_MS.
orchestration.daemonWatchdog.dlqMaxBytes
1073741824
DLQ-file size (bytes, via statSync — O(1), robust past 2 GB) at or above which the dlq predicate trips. Default 1 GiB. Env EXECUTION_CORE_DAEMON_WATCHDOG_DLQ_MAX_BYTES.
orchestration.daemonWatchdog.stalenessMs
900000
How long the checkpoint’s lastForwardedTs may stay frozen — while the event log has fresher writes (real backlog) — before the lag predicate trips. Default 15 min. Env EXECUTION_CORE_DAEMON_WATCHDOG_STALENESS_MS.
orchestration.daemonWatchdog.cooldownMs
900000
Minimum time between restarts (ms). Default 15 min — deliberately > the 600s launchd StartInterval so the two supervision layers never race. Env EXECUTION_CORE_DAEMON_WATCHDOG_COOLDOWN_MS.
orchestration.daemonWatchdog.sustainedTicks
2
Consecutive breach ticks required before the watchdog acts (hysteresis). Env EXECUTION_CORE_DAEMON_WATCHDOG_SUSTAINED_TICKS.
orchestration.daemonWatchdog.verifyTicks
2
Post-restart re-check window: if the predicate is still tripped after this many ticks, the watchdog escalates (latched alert + severity:high finding) instead of restarting again. Env EXECUTION_CORE_DAEMON_WATCHDOG_VERIFY_TICKS.
How often the stall-janitor’s git-heavy worktree/stall censuses (J1 orphan-worktree, J3 stall-clear, J4 terminal-signal GC) may run, off the per-tick scheduler hot path. Each fires a git worktree list per repo plus a git status per terminal worktree, so running them every tick on a many-worktree host ages the daemon heartbeat and holds new-work dispatch; this cadence keeps them off the hot path while the cheap J2 ghost-session kill still runs every tick (CTL-1324). Env CATALYST_STALL_JANITOR_INTERVAL_MS (milliseconds) overrides.
orchestration.laneClaim.timelyRecencySeconds
300
CTL-2070. The lane-claim guard (CTL-2068) refuses a pipeline write/dispatch that would regress a ticket a human lane has claimed. It has a timely per-ticket actor source — a durable host-local write-ledger (~/catalyst/lane-claim-write-ledger.json) of the last Linear state this fleet set and when — that covers the ~200 s window after a claim, before the reconcile-only issue_history catches up (CTL-1847: issues.state ~11 s vs issue_history ~201 s), and the 140 fleet tickets that have no history rows. This bound is how long that ledger stays authoritative: past nowMs - issues.updated_at > timelyRecencySeconds, the now-caught-up issue_history ladder governs instead, so the new REFUSE is confined to the exact claim-to-dispatch window. Env CATALYST_LANE_CLAIM_TIMELY_RECENCY_SECONDS overrides. The timely source is on by default, killable with CATALYST_LANE_CLAIM_TIMELY_SOURCE=off (restores exact CTL-2068 behavior — the guard falls back to issue_history); the source identifies the fleet by its own writes, not by an app-actor id set, so it functions independently of botUserIds / CTL-2074.
CATALYST_LANE_CLAIM_DISPATCH_GUARD(env var)
armed
CTL-2068. The lane-claim guard’s operator kill switch for the DISPATCH veto only (refusing to run a phase on a lane-claimed ticket). Any value other than off leaves it armed; off disables the dispatch veto while keeping the write-side protection. Refusing a state write is benign (the write simply retries); refusing a dispatch withholds work, so this switch exists for an operator who needs the fleet to plough through a claim.
The orphan child-process reaper is the corroboration-heavy companion to the session-level reaper:
claude stop deregisters a worker’s claude agent but leaves its reparented node/bun
grandchildren (MCP servers, sub-agent tooling, bun test runners) running — the bulk of the
resident-memory leak. It runs on the same 600-second cadence as the orphan-session sweep and refuses
to act unless every signal corroborates: a successful claude agents read this cycle (a failed read
aborts the whole sweep), the process is reparented and outside the live-agent process tree, its
command and working directory match, and it has persisted across two consecutive sweeps.
CTL-1531 widened which commands can be a candidate without relaxing any of that corroboration. A
non-node/bun orphan — the motivating case was four sh -c "while :; do :; done" loops that
pegged ~4 cores for 16.5 h from a worktree that had been deleted — is admitted only on positive
ownership evidence: strictly ppid == 1, cwd under worktreeRoot, and that cwd path no longer
existing. Both new probes fail closed (an unresolvable cwd, or a cwd-existence check that cannot
answer, spares the process), nothing outside worktreeRoot is ever a candidate regardless of ppid
or command, and the widened row still passes through every pre-existing gate.
That “cwd no longer exists” probe is tri-state, not a boolean: it distinguishes definitely
gone (a stat errno of ENOENT) from cannot tell (EACCES on an unreadable parent, EIO on a
failing disk, ESTALE/ENOTCONN on a dropped network mount), and spares on the latter. A plain
existsSync/[[ -d ]] collapses both into false, which would read an unanswerable probe as
positive evidence the worktree was deleted — the inversion of the fail-closed rule, on the one
conjunct that authorizes killing an arbitrary command.
Because the widened class can signal any command, it is staged behind its own rollout mode —
procReaper.widenMode in the daemon, SWEEP_PROC_WIDEN in the shell sweep — which defaults to
shadowindependently of procReaper.mode. A host already running mode: "enforce" therefore
merely observes the new class until an operator flips the second knob (ADR-023: dark by default, one
knob per actuator, no enable-on-merge). In the daemon, an enforcing widened candidate additionally
has its whole ownership conjunction (ppid, argv, live-agent ownership, cwd still under the root and
still deleted) re-proved from a fresh read immediately before the SIGTERM — candidates are
classified from one snapshot and then signalled serially, so a late candidate would otherwise act on
evidence tens of seconds stale. Neither implementation ever writes a candidate’s full argv to a log
line or event payload: an arbitrary command’s argv routinely carries tokens, passwords and signed
URLs, and both logs are persisted (the daemon’s is shipped to Loki), so only pid, command basename
and reason are recorded.
The same widening lands in the hourly orphan-sweep.sh vector 1 as an additional branch
alongside the legacy bun run|turbo|node branch (which stays path-unrestricted). Both
implementations also carry a widened-class-only command denylist (tmux, screen, sshd,
ssh, mosh-server, login, launchd, init, systemd, nohup — anchored so the progname:
setproctitle form such as tmux: server … and sshd: ryan [priv] is matched): a session
multiplexer is ppid == 1 by construction and inherits its cwd from whatever shell started it, so
one kill would close every pane the operator has open.
The shell branch is staged by these env vars (set them in the LaunchAgent’s EnvironmentVariables;
SWEEP_PROC_WIDEN is baked into the shipped plist template by install-orphan-sweep.sh, which
preserves an existing flip across the plist regeneration that every routine install-services
performs — see docs/orphan-sweep.md → “Flipping the widened branch to enforce”):
Env
Default
Meaning
SWEEP_PROC_WIDEN
shadow
off | shadow | enforce for the widened branch. Dark by default per ADR-023 — the flip to enforce is operator-owned, never enable-on-merge. Any other value logs a warning and falls back to shadow.
SWEEP_PROC_WIDEN_MAX_KILLS
5
Per-run cap on confirmed widened terminations (0 = uncapped), mirroring vector 2’s SWEEP_MAX_REMOVALS. Counted on the enforcing path only, so shadow still reports the full candidate set. Overflow is logged and deferred to the next run. Delivered signals carry a second ceiling of cap × 2.
SWEEP_PROC_WIDEN_MIN_AGE_SECS
900
Minimum process age (elapsed seconds) for a widened kill, matching procReaper.minEtimeSec. An unreadable age fails closed.
SWEEP_PROC_WIDEN_GRACE_SECS
5
Seconds to wait for a confirmed exit after each of SIGTERM and SIGKILL. kill reports delivery, not exit — only a process observed to have actually gone is logged and emitted as reclaimed.
SWEEP_PROC_CWD_TIMEOUT_SECS
5
Deadline for the per-pid lsof cwd probe, so one hung/stale mount cannot wedge the LaunchAgent run. A timed-out probe yields an unknown cwd (spare), never a truncated path.
CATALYST_LSOF_TIMEOUT_MS
5000
The daemon-side sibling: the deadline proc-reaper.mjs puts on its own lsof cwd probe (single-pid and batched). A value outside (0, 600000] degrades to 5000, never to unbounded. Scope is lsof only — the cwd existence probe is unbounded on both sides (declared, not implied).
Both implementations refuse to run the widened branch at all when the worktree root itself is
missing or unreadable (SWEEP_WT_ROOT in the shell, procReaper.worktreeRoot in the daemon): a
renamed or unmounted root makes every cwd beneath it look deleted in the same pass, which is a
root-level fault rather than N independent orphans — and two-sweep persistence is no defense there,
because the same correlated fault answers both sweeps. Both also bound the widened class per run
(widenMaxKills / SWEEP_PROC_WIDEN_MAX_KILLS), bound the lsof cwd probe with a deadline so one
hung mount cannot wedge the sweep, and treat a liveness probe that could not answer as unknown
rather than as a confirmed exit — so a reclamation is only ever recorded for a process observed to
have actually gone.
Because these are two implementations of one policy, they have drifted in both directions across
review rounds. Every shared safety property now carries a PARITY: <slug> marker at its site in
both plugins/dev/scripts/execution-core/proc-reaper.mjs and plugins/dev/scripts/orphan-sweep.sh,
and proc-reaper.test.mjs asserts the two marker sets are identical — so a hardening added to one
side without the other fails CI instead of waiting for a reviewer.
The fleet-health probe is the steady-state guardrail that ties the reapers together: it samples four
degradation signals (the ~/.claude/jobs dir count, the live background-agent count, the resident
node/bun worker-process count, and macOS swap-used MB), each read fail-safe so an unreadable
signal can only cause the probe to under-react. On a threshold breach it emits one
fleet.health.degraded event (the host lives in the OTel resource block, so the monitor composes
fleet.health.degraded.<host>). Self-heal is default OFF — the first ship is pure
observability. When enabled it fires the same two reap intents the 600-second timer emits plus a
capped (25-process) node/bun child sweep, once per sustained breach, re-armed only after the
fleet recovers to healthy.
For execution-core mode, the number of workers comes from a separate committed block,
orchestration.executionCore.maxParallel (default 4). One daemon runs per machine and serves all
your projects.
In execution-core mode, the daemon reads a central registry at
~/catalyst/execution-core/registry.json. Each project there has an eligibleQuery that says which
tickets the daemon should pick up — status: "Todo". The setup tool
setup-execution-core-states.sh writes this for you; you don’t edit it by hand. That mode also
relies on the pipeline states — Research, Plan, Implement, Validate, and PR — which the
same tool creates on top of the Todo and Triage states your team workflow already has.
If the registry is missing (a fresh or headless host), enroll a project with
catalyst-execution-core register --team <TEAM> --repo-root <path> rather than writing the file by
hand — see Remote and unattended hosts.
Each registry entry’s team must exactly equal the catalyst.linear.teamKey declared in that
entry’s <repoRoot>/.catalyst/config.json. The match is case- and whitespace-sensitive, because the
code that consumes it is: the daemon routes Linear events with a strict !== comparison and looks
projects up with strict ===, so a cat-vs-CAT entry silently drops every event and resolves no
repository.
When the two disagree, a team is pointed at another project’s checkout, and worktrees cut from it
inherit that checkout’s Layer-1 catalyst.linear config (teamKey/teamId) and its
project.ticketPrefix. Catalyst reports the drift in two places, both advisory — a violation never
blocks dispatch:
the daemon logs registry entry repoRoot declares a different Linear team on each registry read;
catalyst doctor reports the registry-team-identity check (WARN on mismatch, never FAIL). The
check grades INFO — not PASS — when an entry’s config is absent, unreadable, malformed, or has no
teamKey, since “no mismatch found” is not the same as “contract verified”.
Repair the registry entry, not the checkout’s config:
Preserve a custom eligibleQuery. Re-registering without the eligible-query flags resets the
entry to the default all-Todo query. If the entry filtered by project, label, or priority, pass
those flags again in the same command, and confirm the result in registry.json.
Clean up worktrees already cut from the wrong checkout. Existing worktrees keep the old
checkout’s config and git remote, and because both clones can resolve the same worktree path,
reuse checks that only compare the branch name will happily keep using them — so later phases go
on running in, and pushing from, the wrong repository. Remove (or re-create) any worktree made
from the mismatched checkout after fixing the registry, then confirm each remaining worktree’s
git remote get-url origin and .catalyst/config.json teamKey are the ones you expect.
Catalyst uses a workspace-scoped worker-status Linear label group with four mutually-exclusive
values (queued, blocked, needs-input, needs-human) to surface each worker’s disposition on
the ticket — independently of where the ticket is in the pipeline. The
setup-execution-core-states.sh tool creates this group idempotently and never duplicates it. You
do not configure the label values in config.json — the group is a Linear-side contract that the
setup tool manages. See the
Worker-status labels reference for what each label
means and how the HUD displays them.
Linear app-actor identity (catalyst.linear.bot.{worker,orchestrator,cloud}.botUserId)
Catalyst posts to Linear as a Linear OAuth app actor — the “Linear for Agents” identity that
comments as Catalyst. Linear OAuth apps are account-level (one app serves every team), so the
bot identity and OAuth credentials now live in the global~/.config/catalyst/config.json under
catalyst.linear.bot, split into app actors:
catalyst.linear.bot.worker — the worker app that posts phase-agent mirror comments and mints
tokens via client_credentials.
catalyst.linear.bot.orchestrator — the orchestrator app that posts run-level updates.
catalyst.linear.bot.cloud — the cloud tenant’s app actor whose identity the fleet’s writes
carry when CATALYST_LINEAR_WRITE_PROXY=enforce (CTL-1889/ADR-0031). Recognition-only.
Each carries a botUserId (the Linear user UUID of that app actor). The daemon and orch-monitor
read all the botUserIds into a single set so the self-echo / loop-prevention guard suppresses
comments and issue events from any of those app actors. These UUIDs aren’t secret (they appear on
every comment the app posts), but they are account-specific.
{
"catalyst": {
"linear": {
"bot": {
"worker": {
"clientId": "...",
"clientSecret": "...",
"webhookSecret": "...",
"accessToken": "...",
"botUserId": null
},
"orchestrator": {
"clientId": "...",
"clientSecret": "...",
"accessToken": "...",
"botUserId": null
},
"cloud": {
"botUserId": null
}
}
}
}
}
Key
What it does
catalyst.linear.bot.worker.botUserId
Linear user UUID of the worker app actor. Suppresses self-echo on mirror comments / description updates. Also the read ID for the daemon’s self-echo filter.
catalyst.linear.bot.orchestrator.botUserId
Linear user UUID of the orchestrator app actor. Also drives self-assign on claim (CTL-1011) — the daemon writes this UUID as the Linear assignee when it claims a ticket. When absent, applyAssignee emits a single deduped warn and leaves the ticket unassigned. Daemon reads it only at startup — restart required after changing.
catalyst.linear.bot.cloud.botUserId
Linear user UUID the fleet’s writes carry when CATALYST_LINEAR_WRITE_PROXY=enforce (the cloud tenant’s app actor, ADR-0031/CTL-1889). Recognition-only: enters the self-echo set so proxied comments/updates are suppressed and don’t read as a human reply; does not drive self-assign (readLinearBotWriteId still prefers the orchestrator id). Provisioned from a proxy-confirmed value (CTL-2074). Verify with catalyst doctor (the self-echo-identity-history check).
OAuth app-actor credentials for the worker identity. Secrets — keep in the un-committed global config
Self-assign activation:catalyst.linear.bot.orchestrator.botUserId must be set AND the
app-actor token must carry the app:assignable OAuth scope for the assignee write to succeed. If
the token lacks scope, a deduped warn is emitted once per Linear team with the re-mint remedy.
See Self-assign activation runbook
below.
Every reader prefers the new global path and falls back to the old location, so a running daemon or
webhook receiver keeps working whether the value has been migrated yet:
Bot IDs:catalyst.linear.bot.{worker,orchestrator,cloud}.botUserId (global) → fall back to
catalyst.monitor.linear.botUserId (per-repo .catalyst/config.json, the legacy single-actor
location).
Worker OAuth creds:catalyst.linear.bot.worker.{clientId,clientSecret} (global) → fall back
to catalyst.linear.agent.{clientId,clientSecret} (per-team
~/.config/catalyst/config-{projectKey}.json, the legacy location).
The legacy keys remain readable, so you can migrate the values at any time without coordinating a
restart.
Catalyst’s app identity lets it post comments as the app, and a human reply on a ticket can wake a
parked worker. To make that work, the system must tell the agent’s own comments and description
updates apart from a human’s. Without a botUserId loaded:
The agent’s own mirror comments get written into the worker inbox as if a human had replied —
noise, and a false “human replied” signal.
Bot-authored issue events feed back into the event log as write loops.
So the botUserId set is the self-echo and loop-prevention guard for the whole Linear-for-Agents
channel. Set at least the worker botUserId for any workspace that uses the app-actor comms.
Query viewer.id with each app-actor token. The app OAuth credentials live in the global secrets
file under catalyst.linear.bot.{worker,orchestrator} (legacy: catalyst.linear.agent in the
per-team file):
Write $BOT_ID into ~/.config/catalyst/config.json under catalyst.linear.bot.worker.botUserId
(repeat for the orchestrator actor), then restart both readers — they only load it at startup:
The orch-monitor Inbox’s reply/unblock feature
(lib/linear-comment.mjs)
posts comments as you, not as the Catalyst app — a Linear provenance gate (CTL-1567)
deliberately ignores app-authored comments, so a reply posted as the bot would silently do nothing.
It resolves a candidate token from, in priority order: env LINEAR_API_TOKEN → env LINEAR_API_KEY
→ this file’s linear.apiToken → the nested catalyst.linear.apiToken — and identity-checks EACH
candidate, using the first one that resolves to a real human (skipping, not failing on, any that
resolve to an app actor).
This matters because LINEAR_API_TOKEN/LINEAR_API_KEY are not exclusively a personal-token slot:
lib/linear-app-actor.sh exports the app-actor’s own OAuth token into those same two env vars for
any daemon that needs bot credentials (broker/execution-core/monitor heartbeats). If your monitor
process sources that script, its env will always carry a non-empty (but bot) token — the identity
walk exists precisely so your real linear.apiToken here still gets tried and used instead of being
permanently shadowed. Generate a personal key at Linear → Settings → API → Personal API keys
(lin_api_..., not an OAuth lin_oauth_... value) and put it here.
The catalyst-otel-forward daemon tails the canonical event log (~/catalyst/events/YYYY-MM.jsonl)
and fans events out to OTLP/HTTP, PostHog, and Cloudflare Analytics Engine. Config lives in the
Layer-2 file above under catalyst.observability.forwarders. All forwarders are disabled by
default. This section is the authoritative schema; the developer reference
plugins/dev/references/event-forwarding.md covers architecture, lifecycle, and DLQ operations.
Collector ingest. OTEL_EXPORTER_OTLP_ENDPOINT overrides it; a :4317 port is auto-rewritten to :4318 for HTTP.
batchSize
100
Advisory target only — not currently enforced. Each flush sends the whole accumulated buffer for that destination in one request; a hard per-request cap is a follow-up (CTL-1506).
flushIntervalMs
5000
Per-forwarder flush cadence — each enabled destination runs its own timer at its own interval (CTL-1506). A destination’s flushes are serialized independently: a tick arriving while its previous flush is still in flight is a no-op, so this is a floor.
lokiAcceptWindowMs
3600000 (1 h)
Age cutoff for Loki records (CTL-1506). Records older than this are dropped with a forward_dropped (drop_reason: "aged") event before any send. Note the effective send cutoff is slightly stricter — lokiAcceptWindowMs − min(timeoutMs, lokiAcceptWindowMs/4) — a delivery margin so a near-cutoff record can’t age past the window mid-request. Tune to your Loki reject_old_samples_max_age.
maxRetryElapsedMs
60000 (60 s)
Max elapsed time for HTTP retry backoff on a retryable failure (429/5xx/network) before the batch is dead-lettered (CTL-1506). Terminal 4xx (not 429) is dropped immediately, never retried or DLQ’d.
dropSurface
see below
Thresholds for the host-local drop surface (CTL-1818) — the counter/marker/alarm that makes a discarded event visible. Sub-keys below; each also has an env override, which wins.
Every discard (drop_reason: "aged" or "terminal_4xx") increments a host-local counter and is
written to the marker ~/catalyst/otel-forward-drops.json; a discard rate that stays above
thresholdRecords for sustainMs raises one ERROR line on ~/catalyst/otel-forward.log
(Alloy-shipped, independent of the OTLP egress this daemon is itself responsible for). The alert is
alert-only — it never restarts the forwarder.
Key
Env override
Default
Description
windowMs
CATALYST_FORWARD_DROP_WINDOW_MS
300000 (5 min)
Rolling window the discard rate is measured over. Matches the provisioned Grafana window for forward_failed, so both are reasoned about in one unit.
thresholdRecords
CATALYST_FORWARD_DROP_THRESHOLD_RECORDS
1000
Records discarded within the window that constitutes a breach (inclusive >=). Sits between the two regimes measured on the fleet: a quiet day ≈ 800/window, a loss storm ≈ 6,900/window.
sustainMs
CATALYST_FORWARD_DROP_SUSTAIN_MS
600000 (10 min)
How long the breach must persist before the alert fires, so a single burst self-heals without paging.
A malformed or out-of-range override is ignored (the surface keeps measuring at its previous
value) rather than silently disabling the counter.
PostHog and Cloudflare AE keys mirror the JSON above; see the developer reference for their delivery
semantics.
Linear webhook 401 alarm (CATALYST_LINEAR_WEBHOOK_ALARM, CTL-1841)
The orch-monitor server watches /api/webhook/linear response codes to detect a sustained
authentication-failure window — the failure mode that stopped all new-work dispatch for 7.5 hours
on 2026-08-14. The alarm is alert-only: it writes a durable marker and a console.warn line
that Alloy ships to Loki independently of the broken webhook path. It never restarts the server.
Two properties are worth stating precisely, because an earlier version of this page got both wrong
and the code followed the page:
The two webhook routes do NOT share a tunnel.linearWebhookConfig is “independent of
webhookConfig… so a daemon can run Linear-only or GitHub-only setups”, and its smeeChannel
drives a second smee tunnel (CTL-242). So a GitHub 200 is not a control for Linear, and
requiring one meant a Linear-only install could never raise this alarm, while GitHub merely being
quiet for the window suppressed a real one. The alarm no longer consults GitHub at all — a Linear
401 already proves the Linear pipe delivered, since the request had to arrive to be rejected.
githubOkAgeMs is still reported for observability; it does not gate the decision.
Not every Linear 4xx is an auth failure. The handler verifies the HMAC signature first,
then returns 400 for a missing delivery header or invalid JSON — those are validly signed
requests. Only 401/403 are stamped as evidence here. A 400/404/5xx stamps nothing and leaves
the previous outcome standing, so an authentic-but-malformed delivery cannot latch this alarm and
send an operator to rotate a secret that is fine. (A 5xx storm deserves its own alarm; it is not
this one.)
Set to 0 to disable the alarm entirely (kill-switch).
CATALYST_LINEAR_WEBHOOK_ALARM_SILENT_MS
900000 (15 min)
No Linear 2xx for this long → “silent window”.
CATALYST_LINEAR_WEBHOOK_ALARM_FAIL_RECENCY_MS
1800000 (30 min)
A Linear non-2xx must be seen within this window to be considered “recent”.
CATALYST_LINEAR_WEBHOOK_ALARM_GITHUB_WINDOW_MS
1800000 (30 min)
The GitHub control 2xx must be seen within this window to prove the tunnel is up.
CATALYST_LINEAR_WEBHOOK_ALARM_TICK_MS
60000 (1 min)
Evaluation interval.
The alarm does not depend on Linear event volume — it detects HTTP authentication failure
directly. A genuinely-quiet Linear feed (no deliveries at all, not even 401s) does not alarm.
Wiring the marker to a pager (a Loki alert rule in catalyst-otel) is a separate follow-up.
CATALYST_CLOUD_TOKEN is a single shared service credential — the catalyst-cloud ADMIN_TOKEN
(interim, per CTC-27 / ADR-0006) — that must be identical on every node. It is an optional
extension: provisioning the token does not by itself change Catalyst’s behavior. Catalyst
stays in its normal local-only state unless both the token is set and you have
specifically configured Catalyst to use the cloud (e.g. local replication + cloud-fed read). Nothing
in Catalyst reads the variable; only the opt-in cloud host-sync daemon (out-of-repo:
catalyst-replica / catalyst-cloud) consumes it. So it is safe to provision cluster-wide without
altering default behavior.
Where it lives (shared state): encrypted in the catalyst-cluster repo as
secrets/cluster-cloud.sops.json (a separate SOPS file from cluster-bots, so the cloud token can
rotate / be garbage-collected independently — it is superseded by per-tenant org-scoped keys per
CTC-46):
How it reaches each node’s machine-level environment (no manual per-host step):
cluster-sync (daemon boot) decrypts it to ~/.config/catalyst/cluster-cloud.json (mode
0600), the same path every other cluster-shared secret takes.
cloud-token-env.mjs — run by catalyst-stack start (boot + keep-alive), or on demand via
catalyst-stack sync-cloud-env — projects it:
writes the secret to ~/.config/catalyst/cluster.env (mode 0600), and
ensures a single non-secret guard line in ~/.zshenv that sources cluster.env.
Every login/zsh shell — and any cloud daemon (re)started in a shell context, this fleet’s
convention for env-key pickup — then inherits CATALYST_CLOUD_TOKEN.
Rotation is boot-scoped: after the value changes in the cluster repo, run catalyst cluster sync
(or restart the daemon) to re-decrypt, then catalyst-stack sync-cloud-env, and restart any cloud
daemon so it picks up the new value. catalyst doctor reports an advisory cloud-token WARN if a
token is decrypted but not yet projected to the machine-level env. The operator runbook for
adding/rotating the secret in the catalyst-cluster repo lives in the docs/cluster-onboarding.md
developer guide (“Provisioning the shared cloud token”).
catalyst.checkouts is a Layer-2, machine-local array of absolute paths — extra git
checkouts this host executes code from or keeps current, beyond the ones already declared
elsewhere. Default: unset (an empty list).
It is the third of four sources in the executing-root enumeration
(execution-core/checkout-sync.mjs → resolveExecutingRoots), which is the single answer to
“which checkouts does this host run?”:
registry repoRoots ∪ this checkout ∪ Layer-2 catalyst.checkouts[] ∪ <CATALYST_DIR>/plugin-source
source
where it comes from
registry repoRoots
<CATALYST_DIR>/execution-core/registry.json — the enrolled projects the daemon dispatches into
this checkout
the tree the resolving module was loaded from (nearest ancestor with a .git)
catalyst.checkouts[]
this key — siblings a given host keeps current but is not enrolled to dispatch into
<CATALYST_DIR>/plugin-source
the tree every daemon on a worker node actually runs FROM
Order-stable and de-duplicated (trailing slashes normalised, so one repo cannot enrol twice).
A path that does not exist on this host is dropped — a registry copied between hosts routinely
names a repoRoot that exists on neither (CTL-854). A non-array value is ignored, never
coerced: spreading a bare string would enrol one root per character.
Consumers, and the two views of the one enumeration. The checkout-sync pass acts on
every root — that breadth is the point, since the 2026-08-12 incident that motivated
CTL-1808 was a sibling repo. The catalyst-agent code-currency gauge measures only the
roots whose role is catalyst (classifyExecutingRoots), emitting
catalyst.vcs.commits_behind one series per such root (label catalyst.checkout.root) plus
catalyst.vcs.commits_behind.max — the stalest of them, which is the host’s single currency
number. Deliberately the same enumeration for both: a root one keeps current that the other
cannot see is the same silent gap in a new place.
Two failures shaped that. Before CTL-1825 the gauge measured only the directory the agent’s
own module lived in, so on a host whose agent and daemons run from different trees (every node
on this fleet) it reported a healthy 0 for a tree nobody executes. Measuring the whole
enumeration then overshot the other way: on the laptop, nine of eleven roots were enrolled
product repos, and the stalest of those (personal-os, 58 commits behind its own main)
became the reported maximum — a metric named for Catalyst currency answering for a personal
repository, and firing max by (host_name)(catalyst_vcs_commits_behind) > 20 on it.
A root is catalyst when it is the agent’s own tree or <CATALYST_DIR>/plugin-source (both
by construction, marker or not — excluding either is the original defect), or when it carries
.claude-plugin/marketplace.json. The role therefore follows what a tree is, never which
source named it: this repository is itself an enrolled repoRoot (ADR-028), and
catalyst.checkouts[] is a config namespace rather than a claim about the repo — its own
example above is catalyst-cloud, a sibling.
catalyst.node.class names what kind of machine this is. It is the front door to per-class
packaging — one declarative field that sets sensible defaults for levers that already exist
(cluster-roster membership, boot-drain, which daemons start, where board reads come from). It adds
no new dispatch gate; the scheduler is unchanged.
Class
What it is
developer
A daemonless client you chat on. Not in the cluster roster, boots drained, runs no execution-core daemon or broker — it reads board UI data from a worker’s monitor (agent Linear reads follow the two-mode rule — see the catalyst-dev:linearis skill’s “Reading Linear” section). On catalyst-stack start, the event-mirror daemon fans worker host event logs into the local copy so catalyst-events tail/wait-for see fleet events.
worker
Runs the full stack and picks up work (the default; a laptop that both runs the daemon and is chatted on is a “head-full worker”).
monitor
A dedicated reporting host (CTL-1654). Like developer it carries the observation substrate (broker + monitor + event-mirror) without the execution layer (no heartbeat, no dispatch, no recovery). The event-mirror daemon (event-mirror/index.ts, launchd-supervised) fans each worker host’s event log into the local ~/catalyst/events/YYYY-MM.jsonl via ssh-tail with per-host byte cursors, so catalyst-events tail/wait-for resolve fleet events locally. Verify with catalyst-stack verify-node.
The class is machine-local, so it lives in Layer-2 (~/.config/catalyst/config.json) beside
catalyst.host.name — the same repo is checked out on every machine, so the role is per-machine,
not per-repo:
Resolution mirrors catalyst.host.name (getNodeClass() in execution-core/config.mjs):
Precedence
Source
1
CATALYST_NODE_CLASS env var (test/override)
2
catalyst.node.class in the Layer-2 config
3
default worker
Absent everywhere ⇒ worker — today’s behavior, zero change (the whole fleet is unset until
it is migrated explicitly). A WARN notes that the class was inferred.
An explicit but unrecognized value (a typo’d developr) does not silently become a
work-eligible worker. It is treated as the most restrictive class and catalyst doctorFAILs
until the value is corrected — so a typo can never make a node pick up work.
A missing or malformed Layer-2 file never throws; it falls through to the worker default.
CATALYST_DRAIN_DISABLED=1 is a worker-only, per-node override that makes the node
permanently ignore the drain flag. It exists because an unattributed recurring writer (CTL-1675)
keeps silently setting the drain sentinel fleet-wide, halting all new-work admission; this env lets
a worker neutralize that flag durably without a code change.
Where to set it: in the durable Layer-2~/.config/catalyst/execution-core.env (it is
sourced into the daemon’s environment at launch), then catalyst-stack restart.
Opt-in semantics: strict === "1" (matching CATALYST_BOOT_DRAINED). Any other value —
unset, 0, true, empty — leaves the node honoring the flag. Unset (the fleet default) is a
byte-for-byte no-op; nothing changes until it is set.
What it does: it makes isDraining() return false at the single admission chokepoint, so
every new-work consumer (the scheduler dispatch gate, the monitor triage-dispatch gate, and
the admission-state reporter) admits work again through one seam. The physical drain file is
left in place — the override neutralizes it, it does not delete it.
Tripwire (attribution): whenever the flag is present but ignored, the daemon emits a
once-per-episode node.drain.ignored event carrying the flag’s mtime and a truncated ps
snapshot, to keep gathering evidence for CTL-1675. It re-arms once the flag is removed and
re-created.
The third state is visible:catalyst-execution-core drain --status-read prints
drain flag present but IGNORED (CATALYST_DRAIN_DISABLED=1) (and warns on stderr when you toggle
the flag on a drain-disabled node), catalyst doctor reports an advisory drain-disabled check
(WARN when the flag is present-and-ignored, PASS/INFO otherwise — never a FAIL), and
catalyst-stack verify-node shows a drain-disabled line on the worker profile.
Does not touch boot-drain.CATALYST_BOOT_DRAINED and the boot-drain policy are deliberately
independent and unchanged, so a developer/monitor node keeps booting drained (this override is
worker-only).
Terminal window
# In ~/.config/catalyst/execution-core.env on a WORKER node (e.g. mini, mini-2):
# MUST be `export`ed — this file is sourced (not `set -a`) before the daemon is
# launched, and bash does not pass a bare (unexported) assignment to the child
# process, so a non-exported value would leave the daemon still drained.
Scope — board UI display only.catalyst.readReplica.baseUrl governs the terminal HUD’s board
reads today (pointing the browser/PWA ticket-detail and search flows at the same endpoint is the
forthcoming “split” topology — CTL-1347 / CTL-1354). It is not the agent Linear read path. For
how agents read Linear ticket data, see the catalyst-dev:linearis skill’s “Reading Linear”
section (two-mode rule: standard node → linearis issues read|list|search directly; Catalyst
Cloud node → @catalyst-cloud/sdk-managed local replica first, with linearis as the
evidence-triggered fallback — CTL-1390). Writes always go through linearis in both modes.
Board data lives in a monitor’s filter-state.db replica, which is written only by a node’s
local broker. A daemonless developer node runs no broker, so its local replica is empty — it must
read a worker’s monitor over the network. catalyst.readReplica.baseUrl names that endpoint,
resolved through:
Precedence
Source
1
CATALYST_MONITOR_URL env var (explicit override)
2
catalyst.readReplica.baseUrl in the Layer-2 config (e.g. http://mini:7400)
3
class-aware default — developer/monitor ⇒ no fallback (explicit error; both read a remote replica); worker ⇒ http://127.0.0.1:7400
A developer (or monitor) node with no endpoint configured returns an explicit unset/error
rather than silently reading an empty localhost replica; a worker keeps the 127.0.0.1:7400
default (its own broker fills and serves the replica). This is reads only — writes still require
a host with its own Linear key, preserving per-host rate-limit isolation.
Scope: this resolver currently backs the terminal HUD’s board reads. Pointing the
browser/PWA ticket-detail and search flows, and the catalyst monitor command, at the same remote
endpoint is the “split” deployment topology tracked in CTL-1347 / CTL-1354.
Local Linear replica + cloud-sync writer (catalyst.linearReplica, CTL-1394)
Not the same thing as readReplica.catalyst.readReplica.baseUrl (above) is the HTTP
board endpoint the terminal HUD reads. catalyst.linearReplica is the local SQLite
Linear-read tier — a per-node ~/catalyst/catalyst-replica.db kept fresh from the Catalyst
Cloud change feed by a supervised writer, read by the scheduler’s hot terminal checks
(replica-read.mjs) and the catalyst-linear CLI. It exists to take Linear reads off the
rate-limited linearis path (the 429 unblock), and is opt-in.
The writer is a supervised launchd LaunchAgent (catalyst-stack adopt-cloud-sync) that runs
@catalyst-cloud/sdk’s CatalystReplica with this node’s own cloud token. It runs on every
node class — workers (mini/mini-2) read the replica from the scheduler hot path; developer nodes
(your laptop) read it via catalyst-linear. The token is never placed in the (world-readable)
plist; the launcher sources it from a 0600 file at run time.
The read flag — on makes the scheduler + catalyst-linear trust the local replica; off/unset reads linearis directly. Env (on/1 on, else off) wins over Layer-2 (mode: "on").
off
CATALYST_REPLICA_DB env
Replica file path.
~/catalyst/catalyst-replica.db
CATALYST_CLOUD_TOKEN (the token itself)
The host’s cloud token — read by a standard name on every host (sourced from the 0600cloud-sync.env, or cluster.env). The per-host-ness is the value you provision, not the name — so the writer installs on arbitrary hosts with no code change.
Operator credential step: obtain the cloud credential from the operator who manages the
service. This repository does not contain it. Provision it in the launcher’s 0600 environment
file:
Repeat after resolving any reported failure until every gating check passes. This verifies the
token, replica database and schema, issue rows, fresh writer lock, and non-empty seed cursor; it
also reports writer-agent and read-flag state. For automation, use
catalyst-stack verify-cloud-sync --json --strict.
Enable replica reads through the guarded activation command:
Terminal window
catalyst-stackactivate-replica
The command refuses to change the Layer-2 read flag until the seed checks in step 3 are genuinely
green. Use catalyst-stack activate-replica --dry-run to preview the config merge without
writing.
On a worker node, restart execution-core so the scheduler constructs its replica reader:
Terminal window
catalyst-execution-corerestart
Why the writer can look healthy while doing nothing: the launchd agent uses
KeepAlive={SuccessfulExit:false}. When no token is available, the writer deliberately exits 0,
which tells launchd not to restart it; an installed plist can therefore coexist with an idle writer
and an unseeded database. Each such launch emits catalyst.replica.writer_idle to the unified event
log while preserving the non-crashing exit contract. catalyst-stack verify-cloud-sync is the
authoritative acceptance check—an installed service alone is not proof of a usable replica.
Nothing in the fleet restarts the cloud-sync writer when a dependency changes. plugin-refresh
deliberately stops at restart_needed (“restart stays a gated OPERATOR action”), and the broker’s
stack reload hard-codes three components — monitor, execution-core, otel-forward — of which
cloud-sync is not one. On 2026-08-04 that meant a merged dependency fix installed correctly on both
minis while the running writers kept serving the old modules for days, with nothing red anywhere: a
merged fix that never restarts is indistinguishable from no fix.
The writer therefore records what it actually resolved at boot (module paths, versions, entry
digests, the root it was served from, and a SHA-256 of that root’s bun.lock) to
~/catalyst/cloud-sync.deps.json, re-hashes the lockfile on every heartbeat, and — in enforce —
takes the existing CTL-1508 self-heal exit (self-heal breadcrumb with reason:"dep-skew" +
bounded close() + exit 1) so launchd KeepAlive relaunches it on the new modules. It is not a
new restart mechanism: launchd is the actuator, the heartbeat tick between applied frames is the safe
exit point, and health-responder.sh’s existing no-respawn condition already nets a relaunch that
never came. Never exit 0 — that is the plist’s “clean no-op, stay DOWN” contract, and a clean
exit here would permanently stop the replica writer.
Key / env
Purpose
Default
CATALYST_CLOUD_SYNC_DEP_SKEW
off (fully dormant — no boot record, no comparison), shadow (detect + emit the skew fields and a would-restart warning; mutate nothing), enforce (self-heal restart). Unrecognised values settle at shadow.
shadow
CATALYST_CLOUD_SYNC_DEP_SKEW_TICKS
Consecutive mismatching heartbeats required before acting — bun rewrites bun.lock in place during an install, so one tick can catch it mid-write.
2
CATALYST_CLOUD_SYNC_DEP_SKEW_UPTIME_MS
Uptime floor: a just-relaunched writer never immediately exits again.
120000
CATALYST_CLOUD_SYNC_DEP_SKEW_MAX_RESTARTS
Durable restart budget — the loop terminator for a lockfile being rewritten continuously by a broken install.
1
CATALYST_CLOUD_SYNC_DEP_SKEW_WINDOW_MS
The window that budget is measured over (~/catalyst/cloud-sync.depskew.json).
21600000
The alarm is not catalyst doctor. An on-demand-only check reproduces “invisible until a human
looks” — the very property that made the incident last for days. The skew observation rides the
writer’s existing structured heartbeat line in every mode, so it is alertable in Loki from day
one:
{service_name="catalyst.cloud-sync"} | json
| __error__=""
| `catalyst.cloud_sync.deps.skewed` = "true"
and the one-shot episode warning carries "catalyst.alert": "replica_dep_skew" with both short
digests, the lockfile path, and the reason a restart was or was not taken.
catalyst doctor grades the same boot record as the advisory cloud-sync-skew check, over links
that each fail closed: the record’s pid must still name a live cloud-sync.mjs (a dead or
recycled pid makes the whole comparison stale evidence); the boot lockfile digest is compared against
the current one at the recorded root (the daemons run from ~/catalyst/plugin-source, not your
dev checkout); each package’s boot entry-file digest is re-hashed and compared — a repointed
mutable artifact or a rebuilt workspace output changes the bytes the process holds while the lockfile
text and the package version stay byte-identical, so the digest is the only discriminator that can
see it; and the installed node_modules version is compared against that install location’s own
lockfile entry — bun keys its packages map by install location (<id> hoisted,
<parent>/<id> nested), so a stale root install is never excused by an unrelated nested resolution
carrying the same version. An absent boot record, an unreadable lockfile or entry file, an install
path that cannot be associated with a lock entry, and zero packages compared all report
WARN “skew unknown”, never PASS.
The durable restart budget (~/catalyst/cloud-sync.depskew.json) is read tri-state: genuinely
absent (ENOENT) is a full budget, but a ledger that exists and cannot be read or parsed declines
the restart, and every field is type-checked before any numeric coercion (Number(null) is 0, which
would otherwise read as an expired window and re-arm a full budget). A corrupt ledger therefore holds
restarts until an operator removes it — never destructive, since a skewed-but-running writer is
exactly the pre-CTL-1659 behavior, now named by the check above. When the ledger cannot be persisted
at all, the write is retried on every heartbeat (so a repaired filesystem resumes the self-heal) but
the ERROR line and its dep_skew_would_restart event are edge-triggered per posture, so a
sustained incident announces once instead of flooding the log every 30 s.
Rollout note. A writer that has not restarted since this shipped has no boot record, so the
first catalyst doctor run after deploy legitimately reports cloud-sync-skew WARN “skew
unknown”. It self-clears on the writer’s next boot.
Linear write proxy (CATALYST_LINEAR_WRITE_PROXY, CTL-1889)
Routes a host’s Linear writes through the Catalyst Cloud write proxy under the one Catalyst
Cloud grant, authenticated with the per-host key, instead of writing to Linear directly with
that host’s own “Catalyst Orchestrator” app-actor (ADR-0031 / CTC-486). The host sends nothing
that identifies it — the cloud derives which host wrote from the key it authenticated, so no
hostname, node name, or actor field appears in the request.
Increment 1 wires the two routes the execution-core write chokepoint owns — issue-state and
label (linear-write.mjs: applyPhaseStatus / applyTerminalDone / applyTriageStatus,
applyLabel, removeLabel). The comment route exists in the transport; its call sites
(lib/linear-comment-post.sh and orch-monitor/lib/linear-comment.mjs) are a later increment.
Nothing is retired by this flag: every existing credential, mint path, and LINEAR_* secret is
untouched, and retirement is gated on a shadow window with zero host-originated writes.
Mode resolves from the env var over Layer-2 over the safe default of off. The 0 kill-switch and
any unset/garbage value resolve to off.
Key
Default
Notes
CATALYST_LINEAR_WRITE_PROXY(env var)
off
off / 0 (kill-switch — the transport is never constructed, no key is resolved, no event is emitted; byte-identical to not setting the flag), shadow (emit linear.write.proxy.would-write.<TICKET> and still perform the existing direct write — shadow makes NO cloud call, since for a write “observe by doing it too” would double-write the board), enforce (the proxy IS the write; the direct path is not taken). Garbage values fall back to off. Overrides Layer-2.
catalyst.linearWriteProxy.mode(Layer-2)
off
Same three values; honored when the env var is absent or unrecognised.
catalyst.linearWriteProxy.routes(Layer-2)
(none)
Per-route path override, { "issue-state": "/…", "label": "/…", "comment": "/…" }. Values must be strings beginning /; anything else is dropped. Layer-2 only — a URL path has no business in a daemon env var. Defaults are measured; see below.
The per-host key is resolved through the secret contract’s cloud-token row — the same row and
the same resolveCloudTokenName ladder cloud-sync uses, so a host whose operator set
CATALYST_CLOUD_TOKEN_ENV or Layer-2 catalyst.cloud.tokenEnv is honored. The endpoint base is
CATALYST_CLOUD_BASE_URL over the shipped cloud default, exactly as cloud-sync resolves it.
CATALYST_CLOUD_ACCOUNT, when set, rides the request as ?account=. It is not required — a correct
per-host key resolves its own tenant — but sending it turns a mis-provisioned host’s generic
401 missing account into the specific 403 … no per-host binding, which names the actual defect.
In enforce, a failure never falls back to a direct Linear write. A fall-back would mean the
host keeps writing under its own app-actor precisely when the proxy is broken, so the shadow window
that gates retirement would read “zero host-originated writes” while the host was still writing.
Every failure is a NAMED reason (no-cloud-token, unauthorized, not-found, rate-limited,
server-error, rejected, unreadable-outcome, transport-error, spawn-failed,
body-too-large, unknown-route, proxy-threw, plus cloud:<outcome> for a refusal the cloud
itself named and resolve:<reason> for one the replica resolver named); callers retry on the next
tick. The verdict is read from the response body’s discriminated outcome, not from the HTTP
status alone — a 2xx carrying no parseable outcome is unreadable-outcome and is not counted as
applied, because these routes always answer with one and a bare 2xx means we reached something
other than the route. A host in enforce with no
per-host key fails with no-cloud-token plus an ERROR line naming the env var it looked in — it
never degrades quietly.
⛔ enforce needs a PER-HOST cloud key, and the tenant-wide one is refused by design. These
routes require an org-owned WorkOS API key provisioned per host. The tenant-wide ADMIN_TOKEN
authenticates (it is a machine principal) but carries no WorkOS key id, so the cloud rejects it by
name: credential has no per-host binding. That is a deliberate refusal, not a misconfiguration
to work around — the key’s own id is what attributes a write to a host. Measured 2026-08-17:
mini-2 carries a real per-host key and a proxied write reached Linear end to end; mini
still carries the ADMIN_TOKEN and every write is refused. Probe a host before enabling it:
400 {"outcome":"rejected","reason":"Entity not found: Issue"} = a good key (it reached Linear and
the fake id was rejected there). 403 … no per-host binding = the ADMIN_TOKEN; this host is not
ready. 401 missing account = the ADMIN_TOKEN with no ?account=.
Route paths are measured, against catalyst-cloudorigin/main — /agent/issue-state,
/agent/issue-label, /agent/issue-comment, relative to a base that already ends in /api/v1.
The routes override is kept anyway: the cloud can move a route without a release of this repo,
and the override makes that a config change rather than a release.
The CTL-758 backward-write guard is re-applied to the state --resolve-only reads. The guard
that runs before it is fail-open — when its own read cannot answer it returns null and falls
through — and the resolve step then reads the real state off the fresh replica. On an enforce host
with no linearis that first read is the most likely to fail and the second the most likely to
succeed, so without the second application the proxy could reopen a terminal ticket. The forward
terminal write stays exempt, exactly as in the original guard.
Identifiers are resolved to Linear UUIDs off the local replica, never a live Linear read
(linear-write-proxy-resolve.mjs): ticket identifier → issues.id, label name → labels.id, and
the target state name → workflow_states.id scoped to the issue’s team. The state NAME itself comes
from linear-transition.sh --resolve-only, which runs the one four-rung precedence chain
(per-project stateMap > global stateMap > the registry’s eligibleQuery.triageStatus > a
built-in default) and writes nothing — so this feature does not become a second source of truth for
what inProgress means. That mode deliberately does not require the linearis binary, because
the end state of CTL-1889 is a host with neither linearis nor a Linear credential.
⛔ The resolver is behind the same freshness gate as the read path. “Read the replica” is only
half the rule. Resolution refuses unless the writer heartbeat (<db>.writer.lock mtime) is younger
than CATALYST_LINEAR_REPLICA_STALE_MS (default 300 s) andsync_meta.cursor is non-empty —
the same two conditions lib/linear-read-replica.sh’s replica_fresh applies, except that this
one refuses (replica-stale / replica-reseeding) where the read path falls back to linearis.
The seed check runs in the same read transaction as the resolution query: during a cold reseed
the entity tables repopulate in batches, so a duplicated label can look unique while only one copy
is back, and the ambiguity guard below is precisely a row-count judgement.
⛔ Resolution failure is a refusal, not a fallback. Unlike every other replica reader in the
tree (which is fail-OPEN and falls through to a live read), this resolver fails closed: an
absent ticket, an archived state, an unreadable replica, or a label name matching more than
one label all refuse the write with a named resolve:<reason>. Ambiguity is real — schema,
types, mobile, infra, etl, dbt and api each match four label rows in the live
workspace, and labels carries no team id to disambiguate with. The four labels this increment
writes (needs-human, needs-input, blocked, queued) are each unique today, but a first-hit
resolver would be one new team-scoped label away from putting another team’s label on a ticket.
Verification narrowing in enforce. The CTL-587 label read-back is a live
linearis issues read — the very host credential this feature retires — so it is not
performed on the proxied path; the cloud response is the verdict. A replica-backed read-back is
tracked with the comment-route call sites.
Observable events (all safe for a LogQL filter
{job="catalyst-events"} | json | attributes["event.name"] =~ "linear\\.write\\.proxy\\..*"):
linear.write.proxy.would-write.<TICKET> — shadow hit; the direct write still happened
linear.write.proxy.applied.<TICKET> — enforce hit, the cloud accepted the write
linear.write.proxy.failed.<TICKET> — enforce hit, NOT written (ERROR; reason attached)
Selects the cross-host (ticket, phase) claim mechanism. The default (off) is the
Linear-attachment soft-CAS (cluster-claim.mjs): an unconditional last-writer-wins upsert with
no conditional-write primitive, so any racer can install itself and no racer can refuse —
correctness depends on every participant running the honest read-back-and-yield code. shadow /
enforce route the claim through the cloud lease authority (a per-tenant Durable Object, the
catalyst-repo half of CTC-410), whose POST /lease/claim is a genuine store-side compare-and-swap:
exactly one racer receives a grant and the loser is told it lost ({claimed:false, refusal}),
so a loser backs off silently with no retry and no error.
The client returns the same{won, generation} / CLI-exit contract as the attachment path, so
no dispatch call site (scheduler.mjs / monitor.mjs / recovery.mjs) changes. On a win the
grant is mapped to the numeric generation written into cluster-generation.json, so
fence-guard.mjs keeps gating every external write unchanged (it compares generations by
equality). Mode resolves from the env var over Layer-2 over the safe default off; the 0
kill-switch and any unset/garbage value resolve to off.
Key
Default
Notes
CATALYST_LEASE_AUTHORITY(env var)
off
off / 0 (the attachment soft-CAS runs; no lease module is constructed; byte-identical to not setting the flag), shadow (the lease claim is called and a lease.claim.would-{grant,refuse}.<TICKET> event is emitted, but the attachment verdict stays authoritative for the dispatch decision), enforce (the lease authority is the claim; the attachment CAS is not consulted). Garbage values fall back to off. Overrides Layer-2.
catalyst.leaseAuthority.mode(Layer-2)
off
Same three values; honored when the env var is absent or unrecognised. Not read from Layer-1 — a live-path claim gate in the committable config would let a merge flip the whole fleet into enforce with no per-host rollback (same containment posture as the write proxy and github-feed).
On a single-host deployment (catalyst.deployment.mode) enforce degrades to off — there are
no peers to arbitrate, so a single-host clone that inherited an enforce env can never block its own
dispatch on a cloud round-trip. The claim path is already unreachable single-host (callers gate on a
non-null cluster generation); this makes that guarantee explicit.
The per-host key, endpoint base, and curl-over-stdin credential discipline are the same as the
write proxy above (the cloud-token secret-contract row, CATALYST_CLOUD_BASE_URL, token never in
argv/on disk). The lease routes require the cloud tenant’s admin token today (CTC-418/419 is the
follow-up to open them to per-host service keys); a host whose CATALYST_CLOUD_TOKEN does not
satisfy that gate stays dark in enforce — run node execution-core/lease-authority.mjs probe (the
non-mutating auth spike: 400 ⇒ authorized, 401/403 ⇒ blocked on the cloud-auth dependency).
The node re-entitles itself on the daemon’s cluster-sync cadence when the gate is shadow/enforce
(a not_entitled refusal otherwise self-heals reactively on the next claim). Not in scope
(follow-ups): progress-asserted lease renewal (/lease/renew requires a non-empty progress
assertion, so a phase that runs past workTtlMs ≈ 45 min could see its lease expire) and explicit
lease release on phase-terminal (the lease frees at TTL regardless — releaseViaLease exists as
a tested capability but is not yet wired to a terminal call site, matching the attachment path, whose
emitFenceReleased is likewise unwired). Retiring the attachment CAS happens after an enforce soak.
catalyst.deployment.mode is the ONE declared answer to a question the system otherwise infers from
side effects — whether a webhook tunnel happens to be configured, whether a cluster roster happens
to resolve to more than one host. It is resolved identically by two independently maintained
implementations — lib/deployment-mode.mjs (Node/ESM) and lib/catalyst-deployment-mode.sh (the
Bash mirror, since Bash cannot import a JS leaf) — kept honest by a fixture-matrix cross-stack
parity test (__tests__/deployment-mode-parity.test.sh).
Value
Meaning
single-host (default)
A lone node — no cluster substrate expected.
cluster
A coordinated multi-host fleet; roster/HRW/liveness are graded by catalyst doctor.
cloud
A managed-container node; the smee webhook tunnel must NOT be live and secrets are platform-delivered.
A 4th both value was deliberately rejected (CTL-1617 design §2) — the provider’s own
webhook.delivery.id already makes concurrent smee+cloud ingestion dedup-safe without one.
Deployment mode is genuinely fleet-scoped, so — unlike catalyst.node.class — it lives
primarily in Layer-1 (committed, shared by the whole repo checkout), with a Layer-2 override
as the exception hatch (e.g. a laptop dev-clone of a cluster-declared repo overriding to
single-host):
In .catalyst/config.json (Layer-1, fleet-wide default):
This repository declares cluster (CTL-1617 PR4 — the working installation is a 2-host fleet).
A dev-clone on a machine that runs no Catalyst stack should set the Layer-2 single-host override
above; without it, catalyst doctor on that machine reports a declared-cluster-but-no-roster
deployment-mode WARN (advisory only — nothing else changes).
Absent everywhere ⇒ single-host — zero-config, zero-behavior-change. (Once wired: a WARN
will note the value was inferred — the resolver itself is deliberately log-free; the WARN lives in
the getDeploymentMode() convenience wrapper, and doctor wiring lands in PR2 of the CTL-1617
migration plan.)
An explicit but unrecognized value (a typo) never silently activates cluster/cloud behavior —
it degrades to single-host (the safest direction) at the layer it was found. (Future behavior,
PR2: catalyst doctor will FAIL until the value is corrected.)
A missing/malformed config file, or a present-but-non-string value (true, 123, []), both
settle rather than throw — see lib/deployment-mode.mjs’s classifyCandidate for the full
validity ladder.
ENV-vs-FILE asymmetry: CATALYST_DEPLOYMENT_MODE is captured into a long-lived daemon’s
environment once, at launch. Layer-1/Layer-2 file edits are picked up live, on every call. A
daemon needs restarting for an env change to take effect.
jq-absent degradation (Bash resolver only): when jq is unavailable, a Layer-1/Layer-2 file
that could otherwise decide the mode is treated as absent (falls through) instead of failing the
caller; the resolver exports CATALYST_DEPLOYMENT_MODE_JQ_MISSING=1 as a breadcrumb (reset at the
start of every resolution, so it always reflects the latest call). Grading that breadcrumb is
future doctor work (PR2) — nothing consumes it yet.
PR1 (this file) ships the resolver in isolation — nothing outside its own tests imports it yet;
wiring into webhook ingestion gating, secret-provider selection, and catalyst doctor’s
roster-consistency checks lands in later PRs of the CTL-1617 migration plan.
Fleet membership is two facts, not one. catalyst.entitlement.mode governs the rollout of the
split (see docs/architecture.md → “Host entitlement vs. existence”):
Existence — “is this node in the fleet, observable/monitorable?” — is local, self-declared,
needs no network, and keeps working during a cloud/authority outage (getExistenceHosts()).
Entitlement — “may this node take work?” — is a lease from an external authority, required to
claim, and self-expiring (getEntitledHosts()).
Modes (off | shadow | enforce, default off):
Mode
Effect
off
Default.getEntitledHosts() returns exactly getClusterHosts() — byte-identical to today.
shadow
Emits entitlement.would-shed.<host> for each unentitled rostered host but returns the FULL roster (safe dry-run).
enforce
Actually sheds unentitled hosts from the dispatch/recovery roster (self always admitted; total-outage degrades to the full roster) and revokes this host’s held leases (fence.released.<ticket>) if its own entitlement lapses.
With the default local provider (entitled iff self ∈ roster), enforce is still byte-identical
to today — self is always entitled — so live enforcement only begins once the W12 lease authority
(CTL-1786, unmerged) is injected.
Resolution (resolveEntitlementMode(), the same ladder shape as deployment mode):
Precedence
Source
1
CATALYST_ENTITLEMENT env var
2
catalyst.entitlement.mode in the Layer-2 config
3
catalyst.entitlement.mode in the Layer-1 config
4
constant default off
Absent everywhere ⇒ off — zero-config, zero-behavior-change.
An explicit but unrecognized value (a typo) degrades to off (the safest direction, byte-identical
to today); catalyst doctor’s advisory entitlement-mode check WARNs (never FAILs).
Same ENV-vs-FILE asymmetry as deployment mode: CATALYST_ENTITLEMENT is captured once at daemon
launch; Layer-1/Layer-2 file edits are picked up live per call.
TTL / ordering constraint — the load-bearing invariant is ENTITLEMENT_TTL_MS > work-lease TTL
(otherwise an unentitled node holds work invisible to the reclaim loop — an orphan by construction):
Constant
Default
Notes
ENTITLEMENT_TTL_MS
15 min
> HEARTBEAT_GRACE_MS (10 min) and > the work-lease TTL below.
WORK_LEASE_TTL_MS
5 min
Mirrors the claim-stale window (CLAIM_STALE_MS_DEFAULT); reconciled with W12’s real lease duration when it lands.
assertEntitlementOrdering() throws loudly at module load if the constraint is inverted; catalyst doctor’s advisory entitlement-ordering check reports it (INFO when it holds, WARN if inverted —
never FAIL). Final numbers depend on W12’s lease duration.
Deletion of the inferred-liveness apparatus (cluster.json roster, loki-liveness.mjs, the
heartbeat channels, the deflap) is out of scope here — that is W16 = CTL-1787, safe only once
a live authority exists. This ticket introduces the entitlement seam alongside the existing apparatus.
Every secret Catalyst resolves — the GitHub token, the Linear API token, the OAuth-mint credentials,
the cloud token, the Groq key, the cluster age-key — is a row in one frozen registry,
SECRET_REGISTRY in plugins/dev/scripts/lib/secret-contract.mjs. It models what used to be
independently hand-rolled resolution ladders per secret (the 2026-08-02 fleet 401 outage was four
divergent copies of one chain) — the Linear-token read and the Linear OAuth-mint trio are live
consumers of it, but the github-token/webhook-secret rows and the groq-api-key row are not yet
RESOLVED through the registry: the live GitHub-token/webhook-secret value paths (CTL-1612’s
catalyst-secret-env.sh / github-auth-preflight.mjs) and Groq’s pre-existing
lib/api-key-health.mjs ladder remain their own, unrepointed implementations — but the
github-token/webhook-secret ROWS do have one live production consumer already:
execution-core/cluster-sync.mjs imports SECRET_REGISTRY and derives its boot-captured secret
membership (which changed credentials require daemon-restart signaling) from these rows’ delivery
types, so their fields are load-bearing even before the resolution cutover (docs/architecture.md’s
Secret Contract section has the full per-row cutover status). Bash cannot import a JS leaf, so the
registry has a second, independently-maintained encoding —
plugins/dev/scripts/lib/catalyst-secret-contract.sh — kept honest by a cross-stack three-way
parity test (__tests__/secret-contract-parity.test.sh): bash and JS must each match a
computed-expected value, never merely match each other, and the two registries must enumerate
identical row-id sets. Both files are zero-import leaves (node:fs/node:os/node:path only on
the JS side) so catalyst doctor, which runs under bare Node, can import the engine without pulling
in execution-core/config.mjs’s bun:sqlite-reaching module graph.
A row is a data fact, not code — the ~7-case engine below is what walks it. Every row declares:
Field
Meaning
id
Canonical identity; doubles as the SOPS bare-file basename for file-backed rows.
envNames
Env-var aliases, precedence order (empty for rows with no direct env alias).
delivery
One of the 7 delivery types below.
configJsonPath
Dotted path inside the resolved Layer-2 JSON, for config-json/platform-env rows; null otherwise.
rotation
{ class, trigger? } — see Rotation below.
bootstrapFor
"cluster" | "cloud" | null — the deployment mode this row bootstraps.
Rows with a more specific shape declare additional fields: familyPrefix (the one
bare-file-family row), defaultLocalPath (the one local-only row, resolved relative to HOME),
and — linear-worker-actor only — credentialEnvPair (an env-var pair checked ahead of every
config-file tier) and legacyConfigTiers + requiredObjectFields (§ Linear worker-actor tiers
below).
The 11 seed rows:
id
delivery
rotation
bootstrapFor
notes
github-token
bare-file
re-armable / timer
—
aliases GH_TOKEN, GITHUB_TOKEN
webhook-secret
bare-file
boot-only
—
env alias CATALYST_WEBHOOK_SECRET
linear-webhook-secret
bare-file-family
boot-only
—
familyPrefix: "linear-webhook-secret-"; a predicate, not a scalar
claude-accounts.env
env-file
boot-only
—
presence-only (a whole sourced env file, not one value)
execution-core.env
env-file
boot-only
—
same shape as claude-accounts.env
linear-api-token
env-alias
re-armable / on-401
—
aliases LINEAR_API_TOKEN, LINEAR_API_KEY
linear-orchestrator-actor
config-json
re-armable / on-401
—
catalyst.linear.bot.orchestrator — kept separate from worker-actor
linear-worker-actor
config-json
boot-only
—
catalyst.linear.bot.worker + a legacy fallback chain (below)
groq-api-key
config-json
boot-only
—
env alias GROQ_API_KEY, config path groq.apiKey
cloud-token
platform-env
boot-only
cloud
default env-var CATALYST_CLOUD_TOKEN; the NAME is itself resolvable
age-key
local-only
n/a
cluster
presence-checked only, never value-read; default ~/.config/catalyst/age.key
linear-orchestrator-actor and linear-worker-actor are deliberately separate rows — they mint
identically and differ only in their config path, an easy-to-collapse-wrongly refactor the registry
exists to prevent.
The engine dispatches on exactly one of 7 delivery types — parity cost between the bash and JS
implementations scales per type, not per row:
Delivery
Resolution chain
bare-file
explicit CATALYST_<ID>_FILE override → CATALYST_CONFIG_DIR → the directory holding CATALYST_LAYER2_CONFIG_FILE (or its ~/.config/catalyst/config.json default) → XDG dir → falls back to an inherited env alias if no file is found
bare-file-family
not a resolvable scalar — a membership predicate (isSecretFamilyMember) over an open-ended prefix family
env-file
presence/non-empty check of a whole file at the same bare-file candidate paths (the file is sourced, not read for one value)
env-alias
first non-empty envNames entry, in order — no file search at all
config-json
env alias (if any) → the resolved Layer-2 JSON’s configJsonPath
platform-env
the row’s env-var name is itself resolved (env override → Layer-2 name override → default), then that variable’s value is read
local-only
statSync presence check of a single path only — the value is never read
Bare-file candidate directory ≠ resolveLayer2Path().secretFileCandidates()
(secret-contract.mjs:607-620, mirrored in catalyst-secret-contract.sh:355-380) builds its
Layer-2-directory candidate straight from CATALYST_LAYER2_CONFIG_FILE (or the hardcoded
~/.config/catalyst/config.json default) — it does not call resolveLayer2Path(), so a
CATALYST_MACHINE_CONFIG-only override is never consulted here, unlike every config-json/
platform-env row (below). An operator who sets only CATALYST_MACHINE_CONFIG=/custom/config.json
and drops a sibling /custom/github-token next to it will find that file skipped: both engines
still search the home/XDG locations. Provision bare-file secrets via CATALYST_CONFIG_DIR, the XDG
directory, or CATALYST_LAYER2_CONFIG_FILE itself so both engines look in the same place.
resolveSecret(id, { env, deploymentMode, cwd }) (JS) / catalyst_resolve_secret <id> (bash) never
throws. It returns { value, source, provider, rotation, ...extras } for a known id, or
{ value: null, source: null, provider: null, rotation: null } for an unknown one. provider is
the row’s delivery — a logging breadcrumb only; callers never branch on it. source is one of
shared-file | operator-override | inherited | config-json | legacy-config-json |
platform-env | present | absent | none — with two exceptions, both of them a known id
whose source collapses to null while provider/rotation stay populated (unlike the unknown-id
shape, where every field is null): the one bare-file-family row (linear-webhook-secret) has no
scalar value, so it resolves to
{ value: null, source: null, provider: "bare-file-family", rotation: {...} }; and, in genuine
cloud mode, any row other than the cloud-token bootstrap row itself resolves the identical
{ value: null, source: null, provider, rotation } shape (that row’s own provider/rotation)
when cloud-token fails to resolve — the bootstrap short-circuit, § Cloud guard below. Both are
normal, expected states — not evidence of an unknown id — so callers must not treat a nullsource as impossible or unknown-id-only. The bash mirror echoes the same three fields pipe-joined
(value|source|provider) and additionally exports non-secret
CATALYST_SECRET_LAST_SOURCE/_PROVIDER breadcrumbs for the calling shell — but deliberately
never exports the resolved value (export -n is reasserted on every call, since bash’s export
attribute is sticky across reassignment), so a long-lived daemon shell can’t leak a credential into
every child process it launches.
The cloud provider only ever activates when the full CTL-1617 deployment-mode object satisfies all
three: mode === "cloud", inferred === false, and recognized !== false. Because the guard
lives once in the shared engine, every row gets it for free. When genuinely cloud, resolution
short-circuits to a pure env-alias read of envNames for the secret value — no file search for
the value, ever — with one carve-out: cloud-token itself is platform-env delivery, and even in
genuine cloud mode it first resolves its env-var name via resolveCloudTokenName() (env
override → the Layer-2 file’s catalyst.cloud.tokenEnv → default), so a managed container can still
consult a config file to learn which variable to read before reading that variable’s value.
Anything else (single-host, cluster, or an inferred/unrecognized cloud guess) runs the row’s normal
delivery-type chain unchanged.
A second gate — the bootstrap short-circuit — applies only in genuine cloud mode: if the active
mode’s bootstrapFor row (cloud-token) fails to resolve, every other cloud-mode secret
resolution returns { value: null, source: null, provider, rotation } (that row’s own provider
and rotation, populated) without probing further, so a half-provisioned managed container fails
loudly and coherently instead of limping through partial resolution. The bootstrap row itself is
exempt from this check (it must resolve on its own terms).
The registry’s canonical Layer-2 config path chain, used by every config-json/platform-env row
and exported as resolveLayer2Path(env) (JS) / catalyst_secret_resolve_layer2_path (bash):
This is a distinct chain from lib/deployment-mode.mjs’s own resolveLayer2Path, which
deliberately mirrors execution-core/config.mjs’s legacy homedir-only behavior — the two names
resolve different things on purpose. execution-core/config.mjs’s getLayer2ConfigPath() and
execution-core/lib/node-class.mjs’s own copy now delegate to this canonical chain, dual-read
against their legacy homedir-only chain for one release: the two only disagree on a host that sets
CATALYST_MACHINE_CONFIG or XDG_CONFIG_HOME without also setting CATALYST_LAYER2_CONFIG_FILE,
in which case the canonical (new) path wins and a one-time-per-message WARN is logged.
Captured once (at daemon boot, or per-call for a config-json mint) — a value change requires a restart to take effect.
re-armable / timer
Proactively re-checked on a recurring tick — the row’s declared shape (github-token). Its actual re-arm today still runs through the pre-existing CTL-1612 rearmGithubTokenFromFile (called every daemon cluster-sync tick in execution-core/daemon.mjs), not through this contract’s own registerRearmHook/armSecret seam — see below, github-token remains hookless.
re-armable / on-401
Reactively re-minted on an observed auth failure, not a timer (the Linear OAuth-mint shape).
n/a
The row’s value is never fetched at all (only age-key — rotation isn’t a question the contract can answer for a presence-only row).
armSecret(id, { env, deploymentMode }) never throws and returns
{ armed, rotated, restartRequired }. restartRequired is the literal signal the 2026-08-02 outage
lacked: it is true exactly when a boot-only row’s resolved value has changed since the last
observation. registerRearmHook(id, fn) attaches an in-process rearm implementation to a
re-armable row — it refuses (returns false, never throws) against a boot-only/n/a row, an
unknown id, or a non-function, so a hookless row can never silently claim “no restart needed”. A
re-armable row with no hook registered degrades to exactly the same boot-only-shaped
behavior (resolve fresh, diff against the last-observed value, report restartRequired on change) —
the capability ceiling is honest either way. As shipped, linear-orchestrator-actor has a real hook
(registered by execution-core/linear-remint.mjs against its cooldown-guarded reminter);
github-token and linear-api-token remain hookless, so both currently take the degrade path in
practice.
linear-worker-actor is the one row with a multi-tier fallback chain, folded from
lib/linear-comment-post.sh’s pre-existing four-rung precedence with every rung’s precedence order
preserved. Note the fold GENERALIZED every file-backed tier’s location, not just one: pre-fold, the
primary and global-legacy tiers read a hardcoded $HOME/.config/catalyst/config.json and the
per-team sibling sat beside it — post-fold all three resolve relative to the canonical
resolveLayer2Path(env), so a CATALYST_LAYER2_CONFIG_FILE/CATALYST_MACHINE_CONFIG override
moves all three together. (Deprecating the legacy tiers is an explicit, separate follow-up — not
part of this fold.)
credentialEnvPair — CATALYST_LINEAR_AGENT_CLIENT_ID/CATALYST_LINEAR_AGENT_CLIENT_SECRET,
checked first; both must be non-empty.
The primary configJsonPath tier (catalyst.linear.bot.worker).
legacyConfigTiers’ per-team-legacy tier, tried only once the above both miss: a per-team
legacy file (walking up from cwd for a .catalyst/config.jsonprojectKey, reading a sibling
config-<key>.json) — the generalization: that sibling now resolves relative to the
canonical resolveLayer2Path() directory, not the old script’s hardcoded
$HOME/.config/catalyst, so it moves with
CATALYST_LAYER2_CONFIG_FILE/CATALYST_MACHINE_CONFIG overrides exactly like every other row’s
Layer-2-relative path. An operator relying on the old hardcoded location under a custom Layer-2
path gets a different (but now-correct-for-the-chain) file here than the pre-fold script read.
legacyConfigTiers’ global-legacy tier, tried only once tier 3 also misses: the global
catalyst.linear.agent path in the canonical Layer-2 file.
A row that declares requiredObjectFields (only linear-worker-actor, requiring clientId and
clientSecret) must find every named field present and non-blank in a candidate tier’s raw object
value before that tier is allowed to win — a partially-populated tier (e.g. a credential-free
{webhookSecret, botUserId} object) falls through to the next tier instead of capturing resolution
and failing downstream.
catalyst doctor consults the contract through resolveSecret directly — never a second
hand-rolled presence check — but for most checks it is shadow-only observability: the contract
is resolved and compared against the check’s existing hand-rolled answer, and a disagreement is
reported as its own STATUS.INFO row (<check>-secret-contract-shadow) that never changes the
check’s grade or exit code. checkPeerUniqueness, checkBotCredentials, and checkWorkerLabels
are the one cutover to date: they now read linear-api-token from the contract as their live
answer (their old hand-rolled LINEAR_API_TOKEN ?? LINEAR_API_KEY read is gone, and so is the
shadow comparison that has nothing left to compare against). checkSecretContract itself remains a
standalone INFO-only observation (presence of linear-api-token and groq-api-key) for every node
class. The one exception that does change doctor’s exit code: checkCloudTokenEnv FAILs when the
active deployment mode is declared cloud (recognized, not inferred) and the cloud-token
bootstrap row does not resolve — the one FAIL doctor cannot route around, per the cloud guard’s
bootstrap short-circuit above.
Catalyst can open PRs, fix CI, answer review bots, and merge. But GitHub decides what must pass
before code lands. Those rules live in GitHub branch protection or rulesets, not in
.catalyst/config.json.
For hands-off merging, set your main branch to require pull requests, require status checks to
pass, and require review threads to be resolved. Then Catalyst drives the PR to the finish and
GitHub enforces the gates. To require a human sign-off too, also require one approving review.
Catalyst reads many more keys — for the event broker, the Monitor dashboard, webhooks, deploy
checks, and worktree setup. The setup script writes them, and
plugins/dev/templates/config.template.json lists them all. You only need the keys above to get
started.
The execution-core scheduler protects itself against a single ticket dominating the dispatch loop.
These knobs are env vars on the catalyst-execution-core process:
SCHEDULER_CIRCUIT_BREAKER_THRESHOLD (default 8) — consecutive failed dispatches (no forward
progress) before a ticket is quarantined to terminal stalled + needs-human. A successful
dispatch resets the counter, so a healthy ticket can never trip it.
SCHEDULER_RUNAWAY_THRESHOLD (default 50) — per-ticket phase.*.<ticket> event count within
SCHEDULER_RUNAWAY_WINDOW_MS that fires one phase.dispatch.runaway.<ticket> alert.
Observability only — it surfaces a dominating ticket without quarantining it.
SCHEDULER_RUNAWAY_WINDOW_MS (default 600000, 10 min) — rolling window for the runaway-rate
alert and its once-per-window suppression marker.
The phantom worker-dir validity sweep quarantines a workers/<ticket>/ dir only when all three
hold: the ticket is definitively not-found in Linear (a clean exit-0 not-found body — a nonzero
exit or transient outage classifies as unknown and is never quarantined), it is not in the
eligible set, and it has no live bg worker. This conjunction guarantees a transient Linear
outage can never quarantine a healthy, resolvable, in-flight ticket.
SCHEDULER_CIRCUIT_BREAKER_THRESHOLD is the Linear-independent backstop; the runaway knobs are
observability only.
The disposition/held-label convergers cool down a failed applyLabel (CTL-834/COORD-236). A
deterministic cloud label rejection — an exclusive-conflict add that the write-proxy refuses,
surfaced to the caller as the generic cloud:failed/cloud:rejected and normalized to
cloud:label-rejected — is cool-down-eligible but never provably terminal on the host, so without a
cap it would retry ~once per cool-down window forever (each refused write still spending a cloud
write-budget unit). These env vars on the catalyst-execution-core process bound that:
SCHEDULER_LABEL_RETRY_CAP (default 5) — after this many cool-down cycles for one
(ticket, label) the converger STOPS re-issuing the write and emits one
linear.label.retry-exhausted event (plus a log.error). Attempts are counted per cool-down
cycle, not per tick. The prior behavior was effectively N=∞, so any finite cap is a strict
improvement.
SCHEDULER_LABEL_RETRY_EXHAUSTED_MS (default 1800000, 30 min) — the long back-off held after the
cap is reached. Must be >SCHEDULER_LABEL_COOLDOWN_MS so the cap wins over the ordinary
per-window cool-down. When it elapses, exactly one self-heal probe apply is allowed so a
since-resolved conflict can still land (the label is never permanently abandoned — COORD-236); a
successful apply resets the counter.
These execution-core environment variables bound how long repeated fence suppression can remain
invisible to a human. The count and age must both be reached; the 45-minute floor spans three
15-minute fence cooldowns, preventing a fast tick loop from paging on a transient blip.
Variable
Default
Meaning
CATALYST_FENCE_STANDOFF_CAP
4
Consecutive suppressions required before out-of-band escalation.
CATALYST_FENCE_STANDOFF_MIN_AGE_MS
2700000 (45 min)
Minimum episode age before out-of-band escalation.
CATALYST_FENCE_STANDOFF_COOLDOWN_MS
21600000 (6 h)
After a delivered break-glass, how long the terminal sweep skips this ticket’s probe + fence check + needs-human write.
CATALYST_FENCE_STANDOFF_DELIVERY_RETRY_MAX
5
Consecutive FAILED break-glass deliveries that may bypass the ordinary 15-minute suppression cooldown to retry on the very next tick.
Crossing the bound writes a durable board/push escalation and does not loosen the fence or perform
a Linear write. Two cadence caveats:
The out-of-band escalation fires once per episode. breakGlassAt latches on the first
delivered break-glass and is reset only by clearFenceStandoff (a later fence PASS, or the
terminal branch of the sweep), so CATALYST_FENCE_STANDOFF_COOLDOWN_MS paces the ticket’s
fence re-probe, not a repeat page.
Because that cooldown gates the whole terminal-sweep probe+write block, a standoff that heals
after a break-glass has its needs-human Linear write — and the terminal-clear retraction on a
late Done — deferred by up to one cooldown window (6 h by default, versus the 15 min the
pre-existing .fence-suppressed marker imposed).
A break-glass whose durable-record or event write fails drops the 15-minute marker so delivery
retries next tick. CATALYST_FENCE_STANDOFF_DELIVERY_RETRY_MAX bounds that bypass: past it the
ordinary cooldown is retained, so a persistently unwritable sink degrades to a 15-minute retry
cadence instead of re-running the per-tick Linear probe + fence-check subprocess forever
(the CTL-1329 quota burn).
The broker’s watchdog tick keeps per-session bookkeeping (lastHeartbeat, workerToOrchestrator)
and a wake-dedup cache (_emittedWakeCache). So these stay bounded over a long-lived broker
process, the tick sweeps expired wake-cache entries every pass and evicts a session’s
heartbeat/orchestrator rows once it is definitively finished. This knob is an env var on the
catalyst-broker process:
FILTER_HEARTBEAT_EVICT_MS (default 1800000, 30 min) — horizon past which a stale session’s
lastHeartbeat + workerToOrchestrator rows are evicted even if it never matched an interest. A
session reported dead by claude agents is evicted immediately; an unknown-liveness session is
evicted only once it is both stale (past FILTER_HEARTBEAT_STALE_MS) and older than this
horizon, so eviction can never precede the stale threshold. Generous by default (≈10×
FILTER_HEARTBEAT_STALE_MS) so only unambiguously-finished sessions are dropped — a genuine
revival simply re-checks-in and recreates the rows.
Phase skills write a durable local thoughts doc after each phase (the record
reconstruct-ticket-state.mjs’s thoughts-artifact walk reads to figure out where a ticket left off
after a crash or cross-host takeover — see docs/orchestrator-overview.md). Pushing that doc
off-machine (humanlayer thoughts sync) is what makes it survive the local disk being lost too,
not just the process. catalyst.phaseArtifactSync.mode controls when that push happens:
Mode (resolve_phase_artifact_sync_mode in plugins/dev/scripts/lib/phase-artifact-sync-mode.sh)
Behavior
off (default)
Roster-gated, byte-identical to pre-CTL-1490 behavior: a single-host roster (.catalyst/hosts.json absent or ≤1 entry) skips the sync entirely (exit 0, no push); a multi-host roster syncs, and a sync failure blocks the phase (exit 11, terminal event failed, reason thoughts_sync_failed).
shadow
Always syncs regardless of roster size, but a failure never blocks the phase — it appends a thoughts.sync.failed.<phase>.<ticket> WARN event to the unified event log and continues (exit 0). Use this to observe sync reliability on a single-host node before turning enforce on.
enforce
Always syncs regardless of roster size, and a failure blocks the phase exactly like off’s multi-host path (exit 11, terminal event failed, reason thoughts_sync_failed).
Any unrecognized value fails safe to off — a typo can never silently start (or silently stop)
pushing thoughts docs off-machine.
Precedence (plugins/dev/scripts/lib/phase-artifact-sync-mode.sh): the
CATALYST_PHASE_ARTIFACT_SYNC_MODE env var (case-insensitive, whitespace-stripped) wins over
.catalyst/config.json’s catalyst.phaseArtifactSync.mode, which wins over the off default.
Consumed by plugins/dev/scripts/lib/thoughts-sync-gate.sh, called from every phase skill
immediately before its terminal phase-agent-emit-complete call.
The exhausted recovery-intent sweep can correlate tickets that reached the retry cap for the same
normalized failure signature. Signatures are recorded at attempt time in the host-local intent
ledger, so grouping is per host; legacy or null-signature entries continue to escalate
individually. In enforce mode, one deterministic anchor carries the operator decision and the other
tickets receive pointer briefs rather than separate decisions.
Key
Default
Notes
CATALYST_RECOVERY_CORRELATION
shadow
off preserves independent escalation and emits no correlation events; shadow computes candidate groups and emits recovery.escalation.would-correlate without changing operator-visible behavior; enforce raises one anchored escalation per group and emits recovery.escalation.correlated for each member. Matching is case-insensitive and trims surrounding whitespace. Any unrecognized value (empty, misspelled) falls back to shadow — never to off, so a typo degrades to observe-only rather than silently disabling the telemetry.
CATALYST_RECOVERY_CORRELATION_WINDOW_MIN
60
Maximum span, in minutes, for matching signatures to belong to one correlated group. Must be a finite positive number; 0, a negative value, Infinity, or a non-number falls back to the default.
CATALYST_RECOVERY_CORRELATION_MIN_GROUP
2
Minimum number of matching candidates required for correlation. Must be an integer of at least 2 (a group of one is a singleton by definition); anything smaller, non-integer, or non-numeric falls back to the default.
Both tunables are read once at module load, so a window or group-size change needs a daemon
restart; the mode is re-read per call and takes effect without one.
A missing or null signature never correlates with another missing signature. Changing the mode to
off restores the independent escalation path; there is no on-disk migration to reverse.
In enforce mode a member whose escalation cannot complete (for example a failed needs-human
label write) records a pointer at its anchor under <orchDir>/.escalation-correlation/<TICKET>.json,
so its retry on a later tick stays a pointer instead of becoming a second operator decision. The
pointer expires with CATALYST_RECOVERY_CORRELATION_WINDOW_MIN.
On a low-frequency cadence the scheduler runs a whole-board health scan: a read-only pass that
evaluates board-level invariants the per-item signals never surface — a silently-held dispatch (open
slots + a waiting queue + no recent dispatch), a worker idling far past its phase-normal age, a
ticket blocked by a dead blocker chain, a project gone silent, a rate-limit cliff, a node that owns
work but whose reconcile is failing. It emits one recovery.board-scan event per cadence (the
numbers ride out as chartable OTel attributes via CTL-1291) and proposes tiered remediation moves.
In shadow (the default) it takes no action; in enforce (CTL-1300) a proceeding scan
dispatches one holistic recovery-pass delegate — see the enforce row below.
Mode resolves from the env var (a single operator knob) over Layer-2 over the default. Unlike the
rest of the recovery family (which ships off), the board-health delegate defaults to shadow:
shadow is itself a dark state — it emits the scan and mutates nothing (the no-mutation guarantee is
structural, not configured), so the telemetry that is the feature’s whole point ships on.
Key
Default
Notes
CATALYST_BOARD_HEALTH(env var)
shadow
off / 0 (kill-switch — strict no-op), shadow (scan + emit recovery.board-scan, take no action), enforce (CTL-1300 — on a proceeding scan, dispatch one holistic recovery-pass delegate anchored to a flagged ticket and carrying the whole-board context; reuses the capped + cooldown’d recovery-pass actuator. Operator-gated — never auto-enabled). Garbage values fall back to shadow. Overrides Layer-2.
catalyst.boardHealth.mode(Layer-2)
shadow
Same three values; honored when the env var is unset.
CATALYST_BH_GH_QUOTA(env var)
shadow
GitHub core REST quota invariant mode: off skips the snapshot read, shadow publishes quota state but keeps the invariant unobservable, and enforce lets a fresh low/exhausted snapshot trip Gate 3 with rate-limit-cliff. Garbage values fall back to shadow. Overrides Layer-2.
catalyst.boardHealth.githubQuota(Layer-2)
shadow
Same three values; honored when CATALYST_BH_GH_QUOTA is unset.
CATALYST_BH_GH_CORE_PCT
10
Remaining core REST percentage at or below which the quota state is low.
CATALYST_BH_GH_QUOTA_STALE_MS
900000 (15 min)
Maximum snapshot age. A missing or older snapshot is unknown and unobservable, even in enforce.
CATALYST_BH_PRODUCTIVITY(env var)
shadow
Per-host productivity invariant mode: off skips the peer-heartbeat read, shadow publishes productivity status but keeps the invariant unobservable, and enforce reports a live peer that owns work but has not advanced past a phase boundary within the configured window. The resulting move is escalate-only and cannot dispatch a recovery delegate. Garbage values fall back to shadow. Overrides Layer-2.
catalyst.boardHealth.productivity(Layer-2)
shadow
Same three values; honored when CATALYST_BH_PRODUCTIVITY is unset.
CATALYST_BH_UNPRODUCTIVE_MS
86400000 (24 h)
Maximum age of a live owning peer’s last phase-boundary advance before the productivity invariant reports it as unproductive.
CATALYST_BH_INTERVAL_MS
300000 (5 min)
Cadence floor — the scan runs at most once per interval per host.
CATALYST_BH_DISPATCH_STALL_MS
600000 (10 min)
Dispatch-liveness threshold: free slots + a queue + no dispatch within this window flags a wedge.
Project-silence threshold (no ticket movement in the project past this window).
CATALYST_BH_UNOWNED_INFLIGHT_MS
86400000 (24 h)
Stale-unowned threshold (CTL-1475). A Linear state like Implement is a claim that a worker is on the ticket, not a label — and nothing takes the claim back when the worker dies. Past this age with no live worker signal and no confirmed-open PR, the ticket is flagged unownedInFlight and proposed as a tier2 (anchorable)recover-unowned-in-flight move, so the delegate dispatches a recovery pass rather than merely reporting it. Such tickets are invisible to every other path: admission only pulls Todo, and the recovery census scans worker dirs they have no entry in. Deliberately conservative — any evidence of ownership spares the ticket, since a false negative costs one more scan while a false positive re-dispatches work a human is holding.
CATALYST_BH_STALLED_PR_REVIEW_MS
259200000 (72 h / 3 d)
CTL-1608. How long a PR may sit without a review-request being responded to before board-health emits a nudge-stalled-pr move. Requires orchestration.stalledPrSweep.enabled: true. The stalled-PR timer stamps reviewRequestedAt in workers/<TICKET>/stalled-pr.json; board-health compares now - reviewRequestedAt against this threshold.
CATALYST_BH_STALLED_PR_CI_MS
172800000 (48 h / 2 d)
CTL-1608. How long a PR may have a continuously-failing CI check before it is flagged stalled. The timer stamps ciFirstFailedAt on first CI failure detection.
CATALYST_BH_STALLED_PR_NOPUSH_MS
432000000 (120 h / 5 d)
CTL-1608. How long a PR may go without a push (no new commits) while still open before it is flagged stalled. The timer stamps lastPushAt on each push detected.
The productivity signal requires peers to publish last_advance_at in their heartbeat records.
During a mixed-version fleet rollout, a peer without that field is treated as unknown and is not
flagged; the invariant begins observing that peer only after the upgraded publisher supplies it.
When a ticket is about to be labelled needs-human, the delegate-first seam can route it to the
delegate runner instead. CATALYST_DELEGATE_FIRST=shadow is the safe dry-run: it emits a
delegate.would-route event on the unified event log (with the ticket, site, and reason) and still
labels needs-human — so an operator can prove the seam works before committing to enforce. In
enforce, the label is suppressed and the ticket is enqueued to the delegate runner (with a
fail-safe fallback to labelling if the runner is not enabled).
Mode resolves from the env var over Layer-2 over the safe default of off. The 0 kill-switch and
any unset/garbage value resolve to off.
Key
Default
Notes
CATALYST_DELEGATE_FIRST(env var)
off
off / 0 (kill-switch — strict no-op, byte-identical to not setting the flag), shadow (emit delegate.would-route on the event log; still labels needs-human), enforce (enqueue to the delegate runner instead of labelling; fail-safe: falls back to labelling if the runner is not enabled). Garbage values fall back to off. Overrides Layer-2.
catalyst.delegateFirst.mode(Layer-2)
off
Same three values; honored when the env var is absent or unrecognised.
Fail-safe gate.enforce silences needs-human only when the delegate runner is confirmed
enabled (via CATALYST_BOARD_HEALTH/CATALYST_RECOVERY_PASS). Lighting only
CATALYST_DELEGATE_FIRST=enforce without a live runner would enqueue intents that drain never, with
the label suppressed — a silent black hole. The gate catches this: if no runner is on, it emits
delegate.route-fallback with reason:"runner-disabled" and falls back to labelling immediately.
Observable shadow events. With shadow mode active, every routeStuckTicketToDelegate call emits
one of three delegate.* events:
delegate.would-route — shadow hit (ticket + site + reason)
POST /api/ticket/<ticket>/reply posts operator-authored text to Linear, and the monitor binds
0.0.0.0 with no auth. Its cross-origin guard validates the request’s Origin against an allowlist
that the caller cannot influence. (It previously compared Origin against the request’s own Host
header — under DNS rebinding both are attacker-chosen, so that comparison could not reject the case
it existed for.)
Trusted by default, all qualified with the port the server actually bound:
loopback — on a wildcard bind the whole 127.0.0.0/8 range (so 127.0.0.2 and the
Debian-conventional 127.0.1.1 work), otherwise only when the bind is itself the loopback address
— a LAN-bound monitor does not own <loopback>:<port>) — localhost, plus the literal(s)
matching the bound address family
Residual, and the real fix. Host names (localhost, os.hostname()) are family-ambiguous —
the browser picks. Under a single-family bind, a process squatting the other family’s port can serve
a page whose Origin is one of those names. No allowlist setting closes this; bind dual-stack
and the squat becomes impossible rather than merely untrusted:
Terminal window
MONITOR_HOST=::catalyst-monitorrestart# `start` no-ops when a monitor is already running
Dev-server note. A prefix assignment (MONITOR_HOST=:: catalyst-monitor restart) exists only
for that command. bun run dev:ui is a separate process, so exportMONITOR_HOST (and
MONITOR_PORT) in the shell that runs it, or the Vite proxy will target the default
127.0.0.1:7400 instead of the monitor’s actual bind.
Management-CLI gap (CTL-1599).catalyst-monitor.sh still prints, opens, and health-probes
http://localhost:$PORT regardless of MONITOR_HOST, so a specific non-loopback bind will show a
wrong URL and a failing probe even when the monitor is healthy. Wiring the CLI through is tracked
separately; the allowlist itself honors the bind correctly.
Env var
Default
Notes
MONITOR_HOST
0.0.0.0
Bind address. :: binds dual-stack (accepts IPv4-mapped too) and is the remedy above. A specific address or hostname narrows the allowlist to that socket only — the monitor stops trusting other local interfaces and this host’s own names, since it no longer owns those sockets. (binding 0.0.0.0 is IPv4-only, so [::1] is not trusted: another service can bind [::1] on the same port and its origin would otherwise pass)
this machine’s own names, wildcard binds only (the bare label of an FQDN is included so
http://mini:7400 works when os.hostname() is mini.corp.example; this trusts whatever that
label resolves to, so on a network where a search domain or stale record maps it elsewhere, prefer
a specific bind or an explicit MONITOR_TRUSTED_ORIGINS) (a name resolves to whichever interface
DNS/mDNS picks, which need not be the one a specific bind listens on) — os.hostname() and its
short label, plus the actual mDNS name on macOS (scutil --get LocalHostName). A
<short>.local alias is not synthesized: when it is not the name the system really
advertises, nothing owns it, so any LAN host could claim it over mDNS and pass the guard.
this machine’s own non-loopback addresses (LAN, Tailscale 100.x) — but only for a wildcard
bind (0.0.0.0/::). A server bound to one specific address trusts only that address, since
another service can hold the same port on a different interface
Own names are trusted only on the bound port, and only under the scheme the monitor serves
(http): a bare http://mini would let any other service on the same machine (e.g. something
on :80) drive the reply route. Comparison keys are full origins (scheme://host[:port]), so
http and https on the same host are distinct — a compromised plaintext endpoint cannot drive an
HTTPS route. On macOS the machine’s real Bonjour name (scutil --get LocalHostName) is included,
since it need not share the first label of os.hostname(); it is cached for 5 minutes (the
lookup spawns a subprocess and the allowlist rebuilds on rejected requests, so it must not run
per-request — but a renamed LocalHostName then takes effect without a daemon restart). Only
http/https origins are accepted — a non-special scheme such as chrome-extension:// serializes
to the opaque "null", which would otherwise match every opaque origin.
The allowlist is rebuilt on a 60s TTL and again on any rejection, so an address that appears
later (Tailscale connecting, a DHCP change) is trusted without a restart, and one that is removed
stops being trusted within the TTL rather than lingering until the daemon restarts.
Env var
Default
Notes
MONITOR_TRUSTED_ORIGINS
unset
Comma- or whitespace-separated extra origins for deployments reached by a name that cannot be derived from os.hostname() — a reverse proxy or a full Tailscale MagicDNS alias. Accepts full origins (https://catalyst.example) or bare host:port (mini-2.tail1234.ts.net:7400). Entries are taken exactly as given (not widened to the bound port) and canonicalized the way a browser serializes Origin, so an IDN name may be written in either Unicode or punycode.
MONITOR_TLS_PROXY_PEERS
unset
Peer addresses of the TLS-terminating reverse proxy (comma-separated, e.g. 127.0.0.1,::1). Requests arriving from these peers are classified https when the accounts ?refresh=true guard probes the trusted-origin set — required when MONITOR_TRUSTED_ORIGINS lists an https:// origin served through a local proxy, or every proxied refresh 403s (with it unset, all requests are plaintext). This is a deliberate operator declaration: X-Forwarded-Proto is never trusted (client-spoofable, and appending proxies put the client’s value first). IPv4-mapped spellings are normalized, so 127.0.0.1 also matches a dual-stack bind’s ::ffff:127.0.0.1.
MONITOR_DEV_UI=1 / NODE_ENV=development
unset
Trusts the Vite dev origin (http://localhost:5173 — one spelling, since Vite binds a single address family and the other would be available to any local process). Not needed for the standard bun run dev:ui flow — the Vite proxy sends the monitor’s own origin (ui/vite.config.ts), so proxied replies are already trusted. Accepts 1/true/yes/on. Use only for a dev setup that bypasses that proxy, and note it must be set on the monitor process (dev:ui starts Vite only; the monitor runs out-of-band).
MONITOR_DEV_UI_ORIGINS
unset
Overrides the dev origins above (same format), for a non-default Vite port. Same caveat: set it on the monitor process.
Set this if replies 403. A monitor opened through a proxy/alias not in the default set will
reject every reply until the name is listed here. Addresses are re-derived automatically (see the
TTL above); a name the daemon cannot derive still needs this variable.
Prefer writing a full origin (https://catalyst.example) over a bare host: a full origin pins
the scheme, whereas a bare host[:port] cannot state one and is therefore trusted under both http
and https.
Requests with noOrigin are allowed — browsers always send it on a POST, so only non-browser
clients (curl, tests) omit it, and those are not CSRF vectors. This guard stops a browser being
used as a confused deputy; it is not authentication.
Parking a ticket (CTL-1552). To tell board-health “a human is holding this ticket, stop
re-proposing it,” apply the standalone parked-by-human Linear label to it (set/cleared from Linear
by an operator). Board-health reads the label off the ticket descriptor, so the park applies on
every host — unlike the per-host board-health sanctioned-latch env var it replaced (CTL-1432
B3), which never synced across the cluster. The parked ticket stays visible in the board context’s
frozenNeedsHuman cohort and in the recovery.board-scan event’s details.sanctioned list; it is
suppressed only from proposeMoves/eligibleDeferredAnchors. It is not a member of the
exclusive worker-status label group (a parked ticket keeps its needs-human label).
The broker tails every event, so it is the surviving process that can notice when an upstream
ingestion source has gone silent — the out-of-process check the monitor cannot do for itself (its
own health probe reports up iff it answers, so it can never observe its own death). Each watchdog
tick the broker judges per-source event recency and edge-triggers
catalyst.ingestion.{stale,recovered} (emit-only — it takes no corrective action; CTL-1123 is the
consumer). These knobs are env vars on the catalyst-broker process:
CATALYST_INGESTION_RECENCY (default on; set 0 to disable) — master kill-switch, read at call
time so it toggles without a broker restart.
FILTER_MONITOR_RECENCY_DEGRADED_MS / FILTER_MONITOR_RECENCY_DOWN_MS (defaults 180000 /
600000) — the catalyst.monitor heartbeat thresholds. The monitor beats on a fixed ~30s
cadence, so these are tight (3m degraded / 10m down) and ungated.
FILTER_GITHUB_RECENCY_DEGRADED_MS / FILTER_GITHUB_RECENCY_DOWN_MS (defaults 900000 /
1800000) — the catalyst.github webhook thresholds. GitHub traffic idles organically, so these
are wide (15m / 30m) and activity-gated: github silence only alarms while a worker is
in-flight (a non-terminal worker_state row that has emitted an event within the last 30 min).
With no active worker the source is forced healthy, so an idle fleet never false-alarms.
FILTER_INGESTION_RECENCY_HOLDDOWN_MS (default 600000) — flap guard: minimum gap between a
recovery and the next stale alarm. A sustained outage that begins inside the window is deferred
(re-checked each tick), never dropped.
Linear (catalyst.linear) recency is intentionally not wired: the linear-webhook bot-skip guard
suppresses bot-authored events before they reach the log, so the source goes quiet even during
active work. Its knobs (FILTER_LINEAR_RECENCY_*) are reserved for when a non-flaky threshold is
found.
The detector above only emits low-level catalyst.ingestion.* events. CTL-1123 adds an
alert-policy layer in the broker: it promotes the operator-actionable subset into a stable,
intentional catalyst.alert.{raised,cleared} topic (event.entity=alert, event.label = the
alert kind). Those events flow through the event log → otel-forward → the OTel collector →
fan-out (Loki, dash0), where a downstream alert rule routes them to a channel. Delivery is
deliberately out of scope — the broker emits intent only; no channel or credential lives in the
daemon. These knobs are env vars on the catalyst-broker process:
FILTER_ALERT_ENABLED (default on; set 0 to disable) — master kill-switch for alert emission,
read at call time (toggles without a broker restart).
The system_down alert is promoted from a critical source’s sustained
catalyst.ingestion.stale (currently catalyst.monitor — a dead monitor). It rides that
already-debounced recency edge, so it has no thresholds of its own; raised on stale, cleared on
recovered.
The needs_human_pileup alert is a level signal: how many active, non-terminal tickets
carry a needs-human/needs-input label in the broker’s filter-state.db (Done/Canceled and
removed tickets are excluded so a stale cached label can’t pin the count). Knobs:
FILTER_PILEUP_THRESHOLD (default 3) — minimum labelled-ticket count to alert.
FILTER_PILEUP_PERSISTENCE_MS (default 300000) — the count must stay at/above the threshold
this long before one alert fires (spike guard).
FILTER_PILEUP_COOLDOWN_MS (default 3600000) — minimum gap after a clear before it can
re-fire (flap guard).
The broker’s watchdog can edge-trigger a broker.daemon.degraded / broker.daemon.recovered pair
when its interest table is empty while the fleet is actively working — an empty table on an idle
fleet is the healthy steady state, so the fleet-activity reading is the discriminator. The episode
is edge-triggered with a durable latch (~/catalyst/broker-degraded-latch.json), so a broker
restart mid-episode resumes rather than re-emitting.
The detector is dormant by default and only meaningful on legacy-wave hosts. Under
execution-core dispatch nothing registers interests at all, so interests.size === 0 is
permanently true and the gate carries no information. That is a property of execution-core, not
of every configuration named phase-agents: a legacy-wave host — one driving
/catalyst-legacy:orchestrate, which invokes
plugins/dev/scripts/orchestrate-register-interests.sh — does register interests (pr_lifecycle +
ticket_lifecycle + comms_lifecycle unconditionally, plus a per-ticket phase_lifecycle when
dispatchMode is phase-agents). There an empty interest table IS anomalous, and that is the
deployment where enabling this is appropriate. These knobs are env vars on the catalyst-broker
process:
FILTER_BROKER_DEGRADED_ENABLED (default off; set to exactly 1 to enable) — opt-in
kill-switch. Unset (or any other value) means the detector evaluates nothing and emits nothing.
Changing this requires a broker restart. The value is read at call time, so it takes effect
without a code reload — but a running daemon’s process.env is fixed at launch and there is no
runtime control path that mutates it, so editing the env file (or exporting in a shell) does
not reach a live broker. Restart it, or you will believe the detector is armed while it is
still dormant. Flipping it off discards any in-progress debounce run, so a re-enable re-earns
the full sustained-tick threshold; an already-open episode survives the switch and still emits its
paired recovered.
FILTER_BROKER_DEGRADED_GRACE_MS (default 300000, 5 min) — startup grace. An empty interest
table is not judged at all until the broker has been up this long, so a still-warming process
never trips.
FILTER_BROKER_DEGRADED_SUSTAINED_TICKS (default 5) — consecutive anomalous watchdog ticks
(~60s each) required before the degraded edge fires. The run must be contiguous: any non-anomalous
tick resets it, so a single-tick blip cannot page.
The fleet-activity reading is tri-state — active, proven idle, or unknown (the worker-table
read failed). Only a proven-idle fleet closes an open episode (recovered, reason fleet idle); an
unknown reading neither trips nor clears, so a transient DB failure cannot manufacture a false
recovery followed by a duplicate degraded edge.
This is not a dead-broker detector, and neither is the ingestion-silence detector above: both
run inside the broker process, so a dead broker emits neither. checkSourceRecency detects an
ingestion stall while the broker is alive. Detecting a fully-dead broker requires an
external, absence-based check on the broker’s own heartbeat/log series — e.g. a Loki
absent_over_time alert on broker.daemon.heartbeat or the broker .log stream (absence, because
a fully-dead daemon is a missing series, which count_over_time == 0 cannot assert).
A daemon that is alive but not accepting new work (draining, or a liveness-cold hold) otherwise
looks fully healthy to uptime monitoring while pulling zero work — the recurring “why isn’t work
moving?” blind spot. CTL-1322 makes that state visible in telemetry: every node.heartbeat
event now carries an admission block in body.payload:
accepting mirrors the scheduler’s new-work gate exactly (livenessFresh && !isDraining()), so
the heartbeat can never disagree with what the daemon actually enforces.
holdReason is "drain" (the persistent operator-intent hold, CTL-1095), "liveness-cold" (the
transient snapshot-staleness hold, CTL-731), or null when accepting. Drain takes precedence when
both apply.
effectiveCapacity is the admission ceiling (maxParallel when accepting, 0 when held);
activeWorkers is the live background-worker count.
The orch-monitor surfaces it for the local node: the FleetOps Hosts “Daemon” column renders
holding (<reason>) (amber) instead of a misleading “live”, and the footer health tooltip gains a
<host> holding (<reason>) line — without bumping the health pill (a drain is operator intent, not
a fleet alarm). Remote peers omit the field (the cross-host anchor transport carries no admission
yet) and render “live”.
Alerting is a host-side Loki/Grafana rule (the same out-of-band, delivery-out-of-scope contract
as CTL-1123) — no in-repo change, no secrets in the daemon. Note the two telemetry paths: the
structured admission block rides the node.heartbeatevent (the unified event log), which on
the current stack is not shipped to Loki — it powers the orch-monitor UI (the server reads the
local event log directly) and any on-host event consumer. What is in Loki — via the Alloy
daemon-.log shipper, stream service_name="catalyst.execution-core" — is the scheduler’s
free-text hold line, so alert on that:
with a for: 10m window; scope per node with | host_name="mini.rozich". This one line fires for
both the drain hold (CTL-1095) and the liveness-cold hold (CTL-731). Follow-ups: ship the
catalyst event log to Loki so the structured admission field becomes queryable, plus a
daemon-emitted debounced catalyst.alert.not_accepting edge (cleaner raised/cleared semantics).
The distributed-coordination epic (ADR-022/023) adds a subsystem that durably orders and shares
coordination events across hosts via a coordination-publish background process: it tails the
unified event log, writes the ordered coordination subset to a local-first mirror
(~/catalyst/coordination.jsonl, carrying a monotonic local_seq) synchronously before any network
call, and — in enforce — exchanges those rows with a catalyst-cloud coordination hub (or, until
the hub is wired, an interim Loki-tail transport).
It ships behind the same off→shadow→enforce rollout discipline as the recovery family, but —
unlike the board-health delegate — its floor is off, not shadow: coordination adds an
always-on publisher process and, in enforce, network egress, so the safe default is fully inert
until an operator promotes it.
The publisher is launched by catalyst-stack — after otel-forward, on worker + monitor nodes
(gated catalyst.node.class != developer, following the reporting substrate) — as a self-managed
nohup child
(catalyst-stackstart_coordination). Because the daemon self-exits when the resolved mode is off (the default),
catalyst-stack short-circuits before spawning it, so launching it unconditionally is safe and an
unconfigured node is byte-identical to before: no publisher process, no PID file, no mirror. When
the mode is shadow/enforce, the daemon comes up idempotently (a re-run of catalyst-stack start
— or the keep-alive tick — never double-starts it); catalyst-stack status reports a coordination
line (off (inert), or running … mode=<mode>). A missing bun is non-fatal (a bun-less host
resolves to off and never blocks the rest of the stack). Enforce and the hub transport stay
operator-gated — this wiring only makes shadow/enforce take effect where they were previously
dark.
Mode resolves from the env var (a single operator knob) over Layer-2 over the default. The 0
kill-switch and any unset/garbage value both resolve to off.
Key
Default
Notes
CATALYST_COORDINATION_MODE(env var)
off
off / 0 (kill-switch — strict no-op: no publisher, no mirror, no egress), shadow (run the publisher and write the local ~/catalyst/coordination.jsonl mirror only — no outbound publish, no inbound pull), enforce (also exchange rows with the hub: outbound buffer + inbound merge; operator-gated, never auto-enabled). Unset or garbage falls back to off. Overrides Layer-2.
catalyst.coordination.mode(Layer-2)
off
Same three values; honored when the env var is unset.
CATALYST_COORDINATION_HUB_URL(env var)
(none)
Base URL of the catalyst-cloud coordination changefeed used in enforce. Overrides Layer-2. When empty/unset the publisher uses the interim Loki-tail transport instead.
catalyst.coordination.hubUrl(Layer-2)
null
Same; honored when the env var is unset.
Cloud token (enforce delivery). In enforce mode the publisher sends each batch to
<hubUrl>/coordination/publish with an Authorization: Bearer <token> header. The payload
carries no tenant field — the server derives tenant from the token. The token comes from the
cloud-token secret-contract row (ADR-0008 leg-3 credential): ~/.config/catalyst/cloud-sync.env
must export the resolved var name (default CATALYST_CLOUD_TOKEN; the exact name is
CATALYST_CLOUD_TOKEN_ENV → Layer-2 → CATALYST_CLOUD_TOKEN). catalyst-stack start sources this
file in a subshell when spawning the daemon so the credential is inherited without leaking into the
long-lived shell process.
An enforce host with a valid hubUrl but no token degrades to inbound-only with a one-time
warning — it never hammers the hub with 401s. Provisioning a token and restarting
(catalyst-stack start) bounces the daemon (token presence is part of the restart fingerprint).
Shadow → enforce promotion (ADR-023 rollout gate). Do not flip hosts to enforce until all of the
following are true, verified by bun coordination-publish/parity-check.ts on each candidate host
over a ≥3–5 day window:
Non-zero matched pairs (the mirror contains coordination rows that align with the worker_state
projection — meaning the outbound stream is non-empty and well-formed).
Zero divergences (no terminal-status conflicts between the coordination mirror and the
worker_state projection).
Each candidate has been spot-checked for false positives.
Once criteria are met, flip one host at a time to enforce. Rollback = unset the mode +
catalyst-stack start (the daemon self-exits on off, ADR-023 rule 5). The three-way exit contract
of parity-check.ts: 0 = healthy, 1 = divergent, 2 = inconclusive (mirror empty or no
matching pairs — not an error, just not yet evaluable).