BLU — ORACLE GATE ZERO: STABILIZE THE REAL FACTORY BEFORE ANY MORE NORMAL WORK¶
You correctly identified that the earlier remediation happened on the wrong machine.
From this point forward:
FACTORY_RUNTIME_AUTHORITY=ORACLE
HOST=bluefly-platform
CITY=/opt/bluefly/blucity
WORKSTATION_GAS_CITY_AUTHORITY=NO
WORKSTATION_BEADS_AUTHORITY=NO
GITLAB=SOURCE_AUTHORITY
ORACLE_GAS_CITY=WORK_AND_RUNTIME_AUTHORITY
NAS=DURABILITY_AND_RECOVERY
Do not relocate, recreate, or synchronize workstation-created mba-* Beads until the authoritative Oracle topology and supported reconciliation mechanism are established.
Do not create duplicates on Oracle merely because a workstation cannot see them.
1. GATE ZERO IS NOW THE ORACLE RUNTIME INCIDENT¶
Normal Factory execution is secondary until the Oracle runtime is stable.
Observed incident:
CPU_CORES=4
LOAD_1M=108.94
LOAD_5M=70.42
LOAD_15M=42.15
MEMORY_USED≈21GB/23GB
AVAILABLE_MEMORY≈1GB
CLAUDE_PROCESSES=32
NPM_MODEL_PROCESSES=35
NODE_PROCESSES=40
ACTIVE_MODEL_WORKERS_MAX=2
OBSERVED_CLAUDE_PROCESSES=32
Do not automatically equate every Claude process with an active Factory worker.
First classify them.
Return:
CLAUDE_TOTAL=
CLAUDE_FACTORY_MANAGED=
CLAUDE_UNMANAGED=
CLAUDE_IDLE=
CLAUDE_EXECUTING=
CLAUDE_ORPHANED=
UNKNOWN=
Evidence before action.
2. DO NOT RESTART THE HOST OR RANDOM SERVICES¶
No broad reboot.
No blind:
killall
pkill
docker restart
systemctl restart
No killing all Claude processes.
No configuration changes merely to make load disappear.
First identify:
WHO_STARTED_THE_PROCESSES
WHAT_OWNS_THEM
WHAT_BEADS_THEY_CLAIM
WHAT_SESSIONS_GAS_CITY_RECOGNIZES
WHAT_WORK_WOULD_BE_INTERRUPTED
Then remediate through the owning lifecycle.
3. THE ZOMBIES ARE A SEPARATE DEFECT¶
Observed:
HOST_ZOMBIES=553
CURL_ZOMBIES=502
PARENT_PID=4880
PARENT_TYPE=root node process inside containerd-shim
Correct interpretation:
ZOMBIES_CAUSING_LOAD=NO
PID_RESOURCE_LEAK=YES
PROCESS_REAPING_DEFECT=YES
Do not mix this with the CPU/load incident.
Identify the container and owning source/IaC first.
Return:
PARENT_PID=
CONTAINER_ID=
CONTAINER_NAME=
IMAGE=
SERVICE=
SOURCE_REPOSITORY=
DEPLOYMENT_OWNER=
KNOWN_BEAD=
ROOT_CAUSE=
Do not infer gastown-gateway unless proven.
If existing work owns it:
UPDATE EXISTING BEAD
not another duplicate.
4. gc doctor CHANGES THE PRIORITY ORDER¶
The Oracle gc doctor --fix result proves several important things.
The controller and supervisor are currently reachable:
controller=PASS
supervisor-http-api=PASS
but the Beads native store is not healthy.
The city is falling back because native schema opening refuses the v53 → v66 migration on remote/shared stores.
That means:
CITY_RUNNING=YES
NATIVE_BEADS_HEALTHY=NO
FALLBACK_STORE_ACTIVE=YES
Do not call the Factory healthy.
5. SCHEMA VERSION IS A FLEET COORDINATION PROBLEM¶
Current evidence repeatedly reports:
DATABASE_SCHEMA=v66
CLIENT_SCHEMA_CAPABILITY=v53
PENDING_MIGRATIONS=13
and upstream correctly refuses to auto-migrate:
REMOTE_BACKED_DATABASE
→ refuses independent migration
→ prevents schema fork
SHARED_SERVER_DATABASE
→ refuses migration
→ prevents locking out co-resident old clients
This protection is good.
Do NOT bypass it.
Do NOT:
ALTER TABLE
BD_IGNORE_SCHEMA_SKEW
force migration per rig
copy database directories
manually rewrite schemas
The remediation must be coordinated across every co-resident client.
6. FIND THE VERSION AUTHORITY FIRST¶
Oracle binaries were hand-placed:
/usr/local/bin/bd
/usr/local/bin/gc
with no established package authority.
This is itself a Factory defect.
Determine:
BD_VERSION=
GC_VERSION=
GC_LINKED_BEADS_LIBRARY_VERSION=
BD_PATH=
GC_PATH=
PACKAGE_MANAGER=
INSTALL_SOURCE=
INSTALL_VERSION=
INSTALL_SHA=
INSTALL_TIMESTAMP=
INSTALL_ACTOR=
If actor cannot be proven:
INSTALL_ACTOR=NOT_ESTABLISHED
Do not guess.
Then inspect current IaC/source to identify what should own installation.
Potential owner should be one of:
agent-docker
IaC
release artifact
documented upstream package mechanism
not a person's shell history.
7. ESTABLISH THE FULL CLIENT COMPATIBILITY MATRIX¶
Before upgrading schema or binaries, enumerate every client connected to the shared Dolt server.
Return:
DOLT_ENDPOINT=127.0.0.1:3308
CLIENT=
HOST=
RIG=
BD_VERSION=
GC_VERSION=
LINKED_BEADS_LIBRARY=
SCHEMA_VERSION=
READ_OK=
WRITE_OK=
NATIVE_STORE_OK=
The doctor output already shows version-compat warnings and multiple rigs refusing migration because they share the server.
Acceptance requires one compatible fleet, not one upgraded executable.
8. FIND WHO APPLIED v54–v66¶
This remains explicitly:
NOT_ESTABLISHED
Do not silently close this question.
Determine using available authoritative evidence:
schema_migrations timestamps
Dolt history
deployment history
CI
package releases
system logs
GitLab deployment evidence
Return:
V54_TO_V66_MIGRATION_ACTOR=
MIGRATION_TIME=
MIGRATION_MECHANISM=
EVIDENCE=
If still unknown:
STATE=NOT_ESTABLISHED
That is acceptable.
Inventing an answer is not.
9. REMOVE NATIVE STORE FALLBACK AS A NORMAL STATE¶
Current doctor output says:
beads-store
= BdStore fallback
= fork-per-op
because native store initialization fails.
This is likely contributing to unnecessary process churn and load.
Do not claim causation until measured.
Measure:
FORK_RATE_BEFORE=
BD_FORKS=
GC_FORKS=
OTHER_FORKS=
Then after the native-store compatibility fix:
FORK_RATE_AFTER=
The doctor currently reports approximately:
fork-rate=8 forks/s
Use that as a baseline only if independently recaptured.
10. WORKER CEILING MUST BE ENFORCED BY THE RUNTIME¶
Mountain 2 already declares:
ACTIVE_MODEL_WORKERS_MAX=2
If Oracle can produce 32 concurrent Claude processes without a deterministic enforcement mechanism, the Factory has a control-plane defect.
Find:
WHERE_WORKER_LIMIT_IS_DECLARED=
WHO_READS_IT=
WHO_ENFORCES_IT=
WHY_OBSERVED_PROCESS_COUNT_EXCEEDS_IT=
Possible outcomes:
32 processes != 32 workers
or:
worker ceiling enforcement is broken
Prove which.
If enforcement is missing, route bounded source/IaC work to the correct owner.
Do not solve it with a cron pkill.
11. IDENTIFY RUNAWAY / ORPHANED SESSIONS¶
Use native Gas City session ownership where possible.
For every executing/long-running session:
SESSION=
AGENT=
RIG=
BEAD=
CLAIM=
STARTED=
LAST_ACTIVITY=
MODEL_PROVIDER=
PROCESS_PID=
STATE=
Then classify:
VALID_EXECUTION
IDLE_MANAGED
ORPHANED
UNKNOWN
Only orphaned/unowned processes should become termination candidates, and termination should use the supported Gas City/session lifecycle.
12. STOP USING gc doctor --fix AS A CASUAL COMMAND¶
--fix is mutation-capable.
From now on:
gc doctor
for observation.
Use:
gc doctor --fix
only when:
FACTORY_GATE=PASS
FIX_SCOPE_KNOWN=YES
MUTATIONS_ENUMERATED=YES
SOURCE_OWNER_KNOWN=YES
A diagnostic session must not mutate production simply because the CLI offers --fix.
13. DOCTOR TIMEOUTS ARE NOT PASSES¶
The output contains multiple checks that say:
timed out
outcome unknown
Examples include:
order-firing-current
v2-routed-to-namespace
run-target-routed-to-backfill
hold-label-routed-to
work-option-metadata-migration
multiple rig beads checks
multiple label checks
Treat these as:
STATE=NOT_ESTABLISHED
not healthy.
The doctor is correctly saying the outcome is unknown. Preserve that distinction.
14. DO NOT CHASE EVERY DOCTOR WARNING DURING THE INCIDENT¶
Separate findings into:
INCIDENT_CRITICAL
POST_STABILIZATION
ADVISORY
INCIDENT_CRITICAL¶
native Beads unavailable
schema compatibility
worker/process explosion
memory pressure
Dolt CPU/load
orphaned sessions
fork-per-op fallback
POST_STABILIZATION¶
bd split stores
JSONL remote durability
pack credential declarations
formula requirement cleanup
config semantic warnings
rig coverage
ADVISORY¶
branch checkout recommendations
event log >100 MB
retention cleanup
Do not turn Gate Zero into a 50-defect cleanup project.
15. JSONL ARCHIVE WARNING IS A DURABILITY GAP¶
Doctor reports:
jsonl-archive
= local-only mode
= off-box backup disabled
That contradicts the durability objective.
But do not blindly run the suggested git remote add.
Determine canonical durability design first:
GITLAB=
NAS=
DOLT_BACKUP=
JSONL_ARCHIVE=
Then route the source/IaC change.
No manual production-only remote.
16. SPLIT STORE MUST BE RECONCILED, NOT DELETED¶
Doctor reports legacy split stores for the city and at least agent-docker.
Do not delete either store.
Required method:
EXPORT BOTH
→ HASH / COUNT
→ DRY-RUN IMPORT
→ IDENTIFY UNIQUE RECORDS
→ RECONCILE
→ VERIFY
→ RETIRE ONLY AFTER PROOF
This is post-stabilization unless it is proven to contribute to the live incident.
17. DO NOT MOVE THE WORKSTATION mba-* BEADS YET¶
Answer to the session's question:
DO_NOT_RELOCATE_YET
First determine:
WHAT_STORE_CREATED_THEM=
PROJECT_ID=
PREFIX=
FEDERATION_STATE=
WHETHER_EQUIVALENT_ORACLE_BEADS_ALREADY_EXIST=
Then:
SEARCH ORACLE
for every item.
Disposition:
EXISTS_ON_ORACLE
→ update canonical bead
UNIQUE_WORK_WITH_VALID_EVIDENCE
→ migrate through supported reconciliation/import mechanism
DUPLICATE
→ retain evidence then supersede/close local record
SESSION_NOISE
→ do not promote
Never manually recreate them one by one.
18. DO NOT LET LOCAL mba-* BECOME FACTORY HISTORY BY ACCIDENT¶
The workstation records were created while the agent incorrectly believed the workstation store was authoritative.
Therefore each must be treated as:
CLAIMED_LOCAL_OBSERVATION
not:
CANONICAL_WORK
Reconcile evidence, not IDs.
19. RUNNINGTODO REMAINS THE MAP¶
Do not create a new incident plan beside it.
Map Gate Zero into the existing structure:
MOUNTAIN 2
= Gas City Factory Economy / Execution Discipline
MOUNTAIN 16
= Oracle / Production Runtime Convergence
MOUNTAIN 17
= Durability where applicable
Use existing canonical Convoys and Beads.
Search first.
One root defect should have one bounded owner.
20. INCIDENT EXECUTION ORDER¶
Execute in this order:
1. CAPTURE CURRENT LOAD / MEMORY / PROCESS BASELINE
2. INVENTORY NATIVE GAS CITY SESSIONS + CLAIMS
3. CLASSIFY THE 32 CLAUDE PROCESSES
4. IDENTIFY PID 4880 CONTAINER / SOURCE OWNER
5. INVENTORY bd/gc CLIENT VERSIONS ACROSS SHARED DOLT CLIENTS
6. IDENTIFY CANONICAL BINARY INSTALLATION OWNER
7. SEARCH EXISTING ORACLE BEADS FOR EACH ROOT DEFECT
8. UPDATE / DEDUP / ROUTE EXISTING WORK
9. BUILD COORDINATED VERSION-CONVERGENCE PLAN
10. APPLY ONLY THROUGH SOURCE/IAC/RELEASE PATH
11. RESTORE NATIVE BEADS
12. VERIFY WORKER CEILING
13. VERIFY LOAD / MEMORY / FORK-RATE RECOVERY
14. WITNESS VERIFY
15. RESUME NORMAL FACTORY EXECUTION
21. GATE ZERO ACCEPTANCE¶
Do not declare Oracle recovered until:
CONTROLLER=RUNNING
SUPERVISOR_API=HEALTHY
NATIVE_BEADS=AVAILABLE
BDSTORE_FALLBACK=NO
DB_SCHEMA=
BD_SUPPORTED_SCHEMA=
GC_LINKED_BEADS_SCHEMA=
COMPATIBLE=YES
SHARED_CLIENTS_COMPATIBLE=YES
ACTIVE_FACTORY_WORKERS<=POLICY_LIMIT
OR
PROCESS_COUNT_EXPLAINED=YES
ORPHANED_MODEL_PROCESSES=0
LOAD_1M=
LOAD_5M=
LOAD_15M=
LOAD_WITHIN_EXPECTED_RANGE=YES
MEMORY_PRESSURE=NO
ZOMBIE_PARENT_IDENTIFIED=YES
ZOMBIE_ROOT_CAUSE_ROUTED=YES
FORK_RATE=
FORK_RATE_ACCEPTABLE=YES
SOURCE_DELIVERY_PROVEN=YES
RUNTIME_READBACK_PROVEN=YES
WITNESS_VERIFIED=YES
Do not require the zombie defect to be completely fixed before Factory resumption unless it remains operationally dangerous.
It must at least be owned and bounded.
22. ECONOMIC / REUSE CLAIM¶
Current state:
IS_THE_NEXT_RUN_GETTING_CHEAPER_AND_MORE_REUSABLE=NOT_ESTABLISHED
WHY=The investigation corrected the execution authority and exposed the real Oracle failure modes, but the runtime still uses Beads fallback, the client/schema lifecycle is not converged, and worker-ceiling enforcement has not been proven.
Only change to:
YES
after the next Oracle startup/run obtains the correct binaries, native Beads compatibility, worker limits, and durable routing automatically from source-controlled infrastructure without rediscovery or manual host surgery.
FINAL LAW¶
DO NOT FIX THE MACHINE.
FIX THE FACTORY LIFECYCLE THAT PRODUCED THE MACHINE.
DO NOT MIGRATE LOCAL TASKS INTO ORACLE BY HAND.
RECONCILE THEM AGAINST ORACLE AUTHORITY.
DO NOT KILL PROCESSES TO MAKE THE GRAPH LOOK HEALTHY.
RESTORE THE CONTROL PLANE THAT PREVENTS THE PROCESS EXPLOSION.
DO NOT FORCE THE DATABASE FORWARD.
CONVERGE EVERY CLIENT THAT SHARES IT.
ORACLE IS THE FACTORY.
THE WORKSTATION IS NOT.