Skip to content

Cluster config mirror contract

This page is the single source of truth for two contracts consumed by cluster setup and monitoring code (M1 mirror tickets, CTL-1192 heartbeat quota):

  1. Config-mirror contract — every config item classified SHARED (copy verbatim to a new node) or PER-NODE (regenerate on each host), with exact file and key locations.
  2. Quota field-name schema — the dotted event-log keys emitted by ratelimit-event.mjs, pinned here so heartbeat and quota consumers (CTL-1192) share one field-name contract.

For the two-layer config model (.catalyst/config.json vs the three Layer-2 siblings: cluster-secrets.json, node.json, config.json) see the configuration reference.


When you provision a second node, copy everything marked SHARED verbatim and regenerate everything marked PER-NODE. The classification is encoded in config.mjs getHostName (PER-NODE), config.mjs resolveClusterHosts (the roster — SHARED, resolved live from the catalyst-cluster repo via readClusterConfig), and config.mjs getLivenessAnchorIssue (SHARED).

Config itemFile / keyClassOn mirror
Bot OAuth orchestrator token~/.config/catalyst/cluster-secrets.json → catalyst.linear.bot.orchestrator.*SHAREDWritten by cluster-sync from cluster-bots.sops.json; one Linear app per workspace, identical across nodes. Falls back to config.json on nodes that haven’t run cluster-sync yet.
Bot OAuth worker token~/.config/catalyst/cluster-secrets.json → catalyst.linear.bot.worker.*SHAREDSame file — worker and orchestrator tokens are workspace-scoped, not host-scoped
Cluster rostercatalyst-cluster repo → cluster.json roster[]SHAREDAdd the new node’s name to cluster.json.roster and push. cluster-sync pulls it and the next scheduler tick honors it — no restart. (The legacy committed .catalyst/hosts.json roster was retired in CTL-1274; the daemon no longer reads it.)
Layer-1 project config.catalyst/config.jsonSHAREDCommitted to git; present after git clone
Liveness anchor issue~/.config/catalyst/cluster-secrets.json → catalyst.cluster.livenessAnchorIssueSHAREDWritten by catalyst-join from the bundle; one Linear ticket identifier per fleet. Falls back to config.json. The configured anchor is infrastructure, not work: do not close, archive, or delete it. catalyst doctor’s liveness-anchor check FAILs if it is archived or missing. To move it, see Moving the liveness anchor.
Cloud token (CATALYST_CLOUD_TOKEN)catalyst-cluster repo → secrets/cluster-cloud.sops.json catalyst.cloud.tokenSHAREDOne shared catalyst-cloud service credential (CTL-1307). Add it once to the cluster repo (SOPS); cluster-sync decrypts it to ~/.config/catalyst/cluster-cloud.json and cloud-token-env.mjs projects it to the machine-level env (cluster.env + ~/.zshenv guard) on every node. Intentionally unread by catalyst core — a prerequisite for the opt-in cloud path, not a switch that turns it on.
Plugin source~/catalyst/plugin-source/SHAREDPull from the same git remote; setup-plugin-source.sh does this
Linear team/state mapLayer-1 catalyst.linear.teamKey / stateMapSHAREDPresent after git clone via .catalyst/config.json
catalyst.host.name~/.config/catalyst/node.json → catalyst.host.namePER-NODEWritten by catalyst-join (non-clobber). Set to the new node’s unique roster entry (must match an entry in the catalyst-cluster repo’s cluster.json.roster; a name that isn’t in the roster owns zero tickets under HRW). Falls back to config.json.
repoRoot~/catalyst/execution-core/registry.json → repoRootPER-NODEThe absolute path on the new host; written by catalyst-execution-core register
Claude Code account loginmacOS Keychain or ~/.claude/.credentials.jsonPER-NODERun claude interactively on the new host; each node uses its own account
OTel endpoints~/.config/catalyst/config.json → OTel keysPER-NODETailscale addresses differ per node; set in Layer-2 on each host
execution-core.env~/catalyst/execution-core/execution-core.envPER-NODEProxy / tuning overrides are host-specific
Event log~/catalyst/events/YYYY-MM.jsonlPER-NODEEach node writes to its own log; nodes never share log files
SQLite databases~/catalyst/*.db (4 files)PER-NODEHost-local state; not replicated
Worktree trust~/.claude.json per worktree pathPER-NODEPaths differ; re-trust on each host
Linear personal token~/.config/catalyst/config-<key>.json → linear.apiKeyPER-NODEPersonal token is user-scoped; each operator provides their own
Webhook secrets~/.config/catalyst/config-<key>.json → webhook keysPER-NODERegenerate or copy securely; not managed by the mirror process

The anchor is the Linear issue every host upserts its catalyst://heartbeat/<host> attachment onto. It is SHARED config: while hosts disagree about which issue it is, they cannot see each other’s heartbeats, so dispatch degrades to the full roster and cross-host failover stops (it does not stop dispatching — see Cross-host ticket ownership in docs/architecture.md).

Move it only when necessary, and move every host in one sitting:

  1. Create or pick the replacement issue. Put “this is the cluster liveness anchor — do not close” in its body, and leave it open.
  2. On every host in the roster: catalyst-cluster set-anchor <TICKET> — this writes catalyst.cluster.livenessAnchorIssue into Layer-2 and reports restartRequired: true.
  3. On every host: catalyst-stack restart. The publisher reads the anchor at arm time, so an un-restarted host keeps publishing to the old issue.
  4. Verify on each host: catalyst doctor → the liveness-anchor check reads PASS, and catalyst-cluster status shows every peer live within one publish interval (EXECUTION_CORE_LIVENESS_PUBLISH_INTERVAL_MS, default 120 s).
  5. Leave the old anchor issue open until step 4 passes everywhere; only then may it be closed.

A host reading liveness from Loki (CATALYST_LIVENESS_READ_SOURCE=loki) is unaffected by the anchor — catalyst doctor grades its liveness-anchor check INFO rather than PASS/FAIL.

Why bot OAuth is SHARED: catalyst.linear.bot.orchestrator and catalyst.linear.bot.worker are credentials for a Linear OAuth application that is registered once per workspace. Every node in the fleet acts on behalf of the same app. The tokens live in machine-global ~/.config/catalyst/cluster-secrets.json (written by cluster-sync from cluster-bots.sops.json; falls back to config.json on nodes that haven’t run cluster-sync yet) so all nodes can share them without per-project duplication.

Why the cloud token is SHARED + machine-level (CTL-1307): CATALYST_CLOUD_TOKEN is a single service credential (the catalyst-cloud ADMIN_TOKEN, interim per CTC-27 / ADR-0006) that must be identical on every node, so it lives once in the catalyst-cluster repo’s secrets/cluster-cloud.sops.json (a separate SOPS file from cluster-bots so its rotation/GC lifecycle is independent — it is superseded by per-tenant org-scoped keys per CTC-46). cluster-sync decrypts it to ~/.config/catalyst/cluster-cloud.json; cloud-token-env.mjs (run by catalyst-stack at boot + keep-alive, or on demand via catalyst-stack sync-cloud-env) projects it into the machine-level environment: the secret is written to a 0600 ~/.config/catalyst/cluster.env, and a single non-secret guard line in ~/.zshenv sources it — so every login shell, and any cloud daemon (re)started in a shell context (this fleet’s convention for env-key pickup), inherits CATALYST_CLOUD_TOKEN. Default behavior is unchanged: nothing in catalyst reads the variable; a node stays fully local-only until the operator separately opts into cloud services.


Single source of truth: ratelimit-event.mjs:63-70 (line 62 emits the account.email identity key, which is not part of the quota schema below).

Event name: account.ratelimit.sampled (severity INFO, emitted every poll tick).

The table below documents the eight dotted attribute keys emitted by buildRatelimitEnvelope. Consumers (orch-monitor, HUD, CTL-1192 heartbeat quota) must reference these names, not the camelCase params used internally by ratelimit-poller.mjs.

Attribute keyTypeMeaning
ratelimit.five_hour_pctnumber5-hour rolling usage as a percentage of the window limit (0–100+)
ratelimit.seven_day_pctnumber7-day rolling usage as a percentage of the window limit
ratelimit.five_hour_resets_atstring (ISO-8601)When the 5-hour window resets
ratelimit.seven_day_resets_atstring (ISO-8601)When the 7-day window resets
ratelimit.seven_day_opus_pctnumber7-day Opus usage as a percentage — the binding limit on Max 20x plans (exhausts before ratelimit.seven_day_pct on Opus-heavy allocations)
ratelimit.seven_day_sonnet_pctnumber7-day Sonnet usage as a percentage
subscription.typestringClaude subscription tier (e.g. "max")
rate_limit.tierstringAPI rate-limit tier identifier

All eight keys are conditional: a key is omitted from the attributes map when its source value is null or undefined. Consumers must treat absent keys as unknown, not as zero.

ratelimit-poller.mjs:257-262 passes values to emitRatelimitEvent using camelCase parameter names (fiveHourPct, sevenDayPct, opusPct, sonnetPct, etc.). These camelCase names are internal-only and must not appear in consumer code or heartbeat schemas. The dotted keys in the table above are the contract; the camelCase params are an implementation detail of the emitter.

CTL-1192 heartbeat quota{} shape (proposed)

Section titled “CTL-1192 heartbeat quota{} shape (proposed)”

When CTL-1192 extends the heartbeat Linear attachment with a quota{} block, it should map the dotted event keys directly:

{
"quota": {
"five_hour_pct": 42,
"seven_day_pct": 18,
"seven_day_opus_pct": 67,
"seven_day_sonnet_pct": 12,
"five_hour_resets_at": "2026-06-16T06:00:00Z",
"seven_day_resets_at": "2026-06-20T00:00:00Z",
"subscription_type": "max",
"rate_limit_tier": "usage_tier_2"
}
}

Use the snake_case field names (strip the ratelimit. prefix) so the heartbeat attachment stays human-readable. The source event keys remain the canonical names — this shape is a derived view.


Account status-transition event (CTL-1653)

Section titled “Account status-transition event (CTL-1653)”

Event name: account.status.changed (v2 OTel envelope, event.entity: "account", severity INFO).

Edge-triggered, not per-sample. Unlike account.ratelimit.sampled (which is emitted every poll tick from the interactive-login /api/oauth/usage probe), account.status.changed is appended by the orch-monitor’s periodic probe only on the ACTIVE account’s ok↔rejected transition — one event per edge, never on a same-status repeat. It is sourced from the CTL-1650 durable-token header probe (claude-accounts-usage.mjs), a different path from the CTL-812 usage sampler; the two coexist. A transport error (sensor failure) is distinct from rejected (account exhausted) and does not trip the transition.

Attribute keyTypeMeaning
account.handlestringThe account label (acctN) that transitioned
account.emailstringThe account email (from the env-file inline comment)
account.statusstringThe new status — "rejected" (exhausted) or "ok" (recovered)
account.binding_windowstringThe binding window driving the status — "five_hour" | "seven_day"
node.namestringThe node whose active account transitioned

The node identity is also carried on the envelope’s resource["host.name"]. The edge is latched durably at ~/catalyst/account-status-latch.json (atomic write, emit-then-advance) so a monitor restart mid-episode does not re-emit. Consumers must treat absent keys as unknown, not as a value.

The same posture is available as a pull surface — the token-free GET /api/accounts endpoint and its GET /api/accounts/stream SSE (see the orch-monitor API).