Skip to content

BLU — ORACLE GATE ZERO: STABILIZE THE REAL FACTORY BEFORE ANY MORE NORMAL WORK

You correctly identified that the earlier remediation happened on the wrong machine.

From this point forward:

FACTORY_RUNTIME_AUTHORITY=ORACLE
HOST=bluefly-platform
CITY=/opt/bluefly/blucity

WORKSTATION_GAS_CITY_AUTHORITY=NO
WORKSTATION_BEADS_AUTHORITY=NO

GITLAB=SOURCE_AUTHORITY
ORACLE_GAS_CITY=WORK_AND_RUNTIME_AUTHORITY
NAS=DURABILITY_AND_RECOVERY

Do not relocate, recreate, or synchronize workstation-created mba-* Beads until the authoritative Oracle topology and supported reconciliation mechanism are established.

Do not create duplicates on Oracle merely because a workstation cannot see them.


1. GATE ZERO IS NOW THE ORACLE RUNTIME INCIDENT

Normal Factory execution is secondary until the Oracle runtime is stable.

Observed incident:

CPU_CORES=4
LOAD_1M=108.94
LOAD_5M=70.42
LOAD_15M=42.15

MEMORY_USED≈21GB/23GB
AVAILABLE_MEMORY≈1GB

CLAUDE_PROCESSES=32
NPM_MODEL_PROCESSES=35
NODE_PROCESSES=40

ACTIVE_MODEL_WORKERS_MAX=2
OBSERVED_CLAUDE_PROCESSES=32

Do not automatically equate every Claude process with an active Factory worker.

First classify them.

Return:

CLAUDE_TOTAL=
CLAUDE_FACTORY_MANAGED=
CLAUDE_UNMANAGED=
CLAUDE_IDLE=
CLAUDE_EXECUTING=
CLAUDE_ORPHANED=
UNKNOWN=

Evidence before action.


2. DO NOT RESTART THE HOST OR RANDOM SERVICES

No broad reboot.

No blind:

killall
pkill
docker restart
systemctl restart

No killing all Claude processes.

No configuration changes merely to make load disappear.

First identify:

WHO_STARTED_THE_PROCESSES
WHAT_OWNS_THEM
WHAT_BEADS_THEY_CLAIM
WHAT_SESSIONS_GAS_CITY_RECOGNIZES
WHAT_WORK_WOULD_BE_INTERRUPTED

Then remediate through the owning lifecycle.


3. THE ZOMBIES ARE A SEPARATE DEFECT

Observed:

HOST_ZOMBIES=553
CURL_ZOMBIES=502
PARENT_PID=4880
PARENT_TYPE=root node process inside containerd-shim

Correct interpretation:

ZOMBIES_CAUSING_LOAD=NO
PID_RESOURCE_LEAK=YES
PROCESS_REAPING_DEFECT=YES

Do not mix this with the CPU/load incident.

Identify the container and owning source/IaC first.

Return:

PARENT_PID=
CONTAINER_ID=
CONTAINER_NAME=
IMAGE=
SERVICE=
SOURCE_REPOSITORY=
DEPLOYMENT_OWNER=
KNOWN_BEAD=
ROOT_CAUSE=

Do not infer gastown-gateway unless proven.

If existing work owns it:

UPDATE EXISTING BEAD

not another duplicate.


4. gc doctor CHANGES THE PRIORITY ORDER

The Oracle gc doctor --fix result proves several important things.

The controller and supervisor are currently reachable:

controller=PASS
supervisor-http-api=PASS

but the Beads native store is not healthy.

The city is falling back because native schema opening refuses the v53 → v66 migration on remote/shared stores.

That means:

CITY_RUNNING=YES
NATIVE_BEADS_HEALTHY=NO
FALLBACK_STORE_ACTIVE=YES

Do not call the Factory healthy.


5. SCHEMA VERSION IS A FLEET COORDINATION PROBLEM

Current evidence repeatedly reports:

DATABASE_SCHEMA=v66
CLIENT_SCHEMA_CAPABILITY=v53
PENDING_MIGRATIONS=13

and upstream correctly refuses to auto-migrate:

REMOTE_BACKED_DATABASE
→ refuses independent migration
→ prevents schema fork

SHARED_SERVER_DATABASE
→ refuses migration
→ prevents locking out co-resident old clients

This protection is good.

Do NOT bypass it.

Do NOT:

ALTER TABLE
BD_IGNORE_SCHEMA_SKEW
force migration per rig
copy database directories
manually rewrite schemas

The remediation must be coordinated across every co-resident client.


6. FIND THE VERSION AUTHORITY FIRST

Oracle binaries were hand-placed:

/usr/local/bin/bd
/usr/local/bin/gc

with no established package authority.

This is itself a Factory defect.

Determine:

BD_VERSION=
GC_VERSION=
GC_LINKED_BEADS_LIBRARY_VERSION=

BD_PATH=
GC_PATH=

PACKAGE_MANAGER=
INSTALL_SOURCE=
INSTALL_VERSION=
INSTALL_SHA=
INSTALL_TIMESTAMP=
INSTALL_ACTOR=

If actor cannot be proven:

INSTALL_ACTOR=NOT_ESTABLISHED

Do not guess.

Then inspect current IaC/source to identify what should own installation.

Potential owner should be one of:

agent-docker
IaC
release artifact
documented upstream package mechanism

not a person's shell history.


7. ESTABLISH THE FULL CLIENT COMPATIBILITY MATRIX

Before upgrading schema or binaries, enumerate every client connected to the shared Dolt server.

Return:

DOLT_ENDPOINT=127.0.0.1:3308

CLIENT=
HOST=
RIG=
BD_VERSION=
GC_VERSION=
LINKED_BEADS_LIBRARY=
SCHEMA_VERSION=
READ_OK=
WRITE_OK=
NATIVE_STORE_OK=

The doctor output already shows version-compat warnings and multiple rigs refusing migration because they share the server.

Acceptance requires one compatible fleet, not one upgraded executable.


8. FIND WHO APPLIED v54–v66

This remains explicitly:

NOT_ESTABLISHED

Do not silently close this question.

Determine using available authoritative evidence:

schema_migrations timestamps
Dolt history
deployment history
CI
package releases
system logs
GitLab deployment evidence

Return:

V54_TO_V66_MIGRATION_ACTOR=
MIGRATION_TIME=
MIGRATION_MECHANISM=
EVIDENCE=

If still unknown:

STATE=NOT_ESTABLISHED

That is acceptable.

Inventing an answer is not.


9. REMOVE NATIVE STORE FALLBACK AS A NORMAL STATE

Current doctor output says:

beads-store
= BdStore fallback
= fork-per-op

because native store initialization fails.

This is likely contributing to unnecessary process churn and load.

Do not claim causation until measured.

Measure:

FORK_RATE_BEFORE=
BD_FORKS=
GC_FORKS=
OTHER_FORKS=

Then after the native-store compatibility fix:

FORK_RATE_AFTER=

The doctor currently reports approximately:

fork-rate=8 forks/s

Use that as a baseline only if independently recaptured.


10. WORKER CEILING MUST BE ENFORCED BY THE RUNTIME

Mountain 2 already declares:

ACTIVE_MODEL_WORKERS_MAX=2

If Oracle can produce 32 concurrent Claude processes without a deterministic enforcement mechanism, the Factory has a control-plane defect.

Find:

WHERE_WORKER_LIMIT_IS_DECLARED=
WHO_READS_IT=
WHO_ENFORCES_IT=
WHY_OBSERVED_PROCESS_COUNT_EXCEEDS_IT=

Possible outcomes:

32 processes != 32 workers

or:

worker ceiling enforcement is broken

Prove which.

If enforcement is missing, route bounded source/IaC work to the correct owner.

Do not solve it with a cron pkill.


11. IDENTIFY RUNAWAY / ORPHANED SESSIONS

Use native Gas City session ownership where possible.

For every executing/long-running session:

SESSION=
AGENT=
RIG=
BEAD=
CLAIM=
STARTED=
LAST_ACTIVITY=
MODEL_PROVIDER=
PROCESS_PID=
STATE=

Then classify:

VALID_EXECUTION
IDLE_MANAGED
ORPHANED
UNKNOWN

Only orphaned/unowned processes should become termination candidates, and termination should use the supported Gas City/session lifecycle.


12. STOP USING gc doctor --fix AS A CASUAL COMMAND

--fix is mutation-capable.

From now on:

gc doctor

for observation.

Use:

gc doctor --fix

only when:

FACTORY_GATE=PASS
FIX_SCOPE_KNOWN=YES
MUTATIONS_ENUMERATED=YES
SOURCE_OWNER_KNOWN=YES

A diagnostic session must not mutate production simply because the CLI offers --fix.


13. DOCTOR TIMEOUTS ARE NOT PASSES

The output contains multiple checks that say:

timed out
outcome unknown

Examples include:

order-firing-current
v2-routed-to-namespace
run-target-routed-to-backfill
hold-label-routed-to
work-option-metadata-migration
multiple rig beads checks
multiple label checks

Treat these as:

STATE=NOT_ESTABLISHED

not healthy.

The doctor is correctly saying the outcome is unknown. Preserve that distinction.


14. DO NOT CHASE EVERY DOCTOR WARNING DURING THE INCIDENT

Separate findings into:

INCIDENT_CRITICAL
POST_STABILIZATION
ADVISORY

INCIDENT_CRITICAL

native Beads unavailable
schema compatibility
worker/process explosion
memory pressure
Dolt CPU/load
orphaned sessions
fork-per-op fallback

POST_STABILIZATION

bd split stores
JSONL remote durability
pack credential declarations
formula requirement cleanup
config semantic warnings
rig coverage

ADVISORY

branch checkout recommendations
event log >100 MB
retention cleanup

Do not turn Gate Zero into a 50-defect cleanup project.


15. JSONL ARCHIVE WARNING IS A DURABILITY GAP

Doctor reports:

jsonl-archive
= local-only mode
= off-box backup disabled

That contradicts the durability objective.

But do not blindly run the suggested git remote add.

Determine canonical durability design first:

GITLAB=
NAS=
DOLT_BACKUP=
JSONL_ARCHIVE=

Then route the source/IaC change.

No manual production-only remote.


16. SPLIT STORE MUST BE RECONCILED, NOT DELETED

Doctor reports legacy split stores for the city and at least agent-docker.

Do not delete either store.

Required method:

EXPORT BOTH
→ HASH / COUNT
→ DRY-RUN IMPORT
→ IDENTIFY UNIQUE RECORDS
→ RECONCILE
→ VERIFY
→ RETIRE ONLY AFTER PROOF

This is post-stabilization unless it is proven to contribute to the live incident.


17. DO NOT MOVE THE WORKSTATION mba-* BEADS YET

Answer to the session's question:

DO_NOT_RELOCATE_YET

First determine:

WHAT_STORE_CREATED_THEM=
PROJECT_ID=
PREFIX=
FEDERATION_STATE=
WHETHER_EQUIVALENT_ORACLE_BEADS_ALREADY_EXIST=

Then:

SEARCH ORACLE

for every item.

Disposition:

EXISTS_ON_ORACLE
→ update canonical bead

UNIQUE_WORK_WITH_VALID_EVIDENCE
→ migrate through supported reconciliation/import mechanism

DUPLICATE
→ retain evidence then supersede/close local record

SESSION_NOISE
→ do not promote

Never manually recreate them one by one.


18. DO NOT LET LOCAL mba-* BECOME FACTORY HISTORY BY ACCIDENT

The workstation records were created while the agent incorrectly believed the workstation store was authoritative.

Therefore each must be treated as:

CLAIMED_LOCAL_OBSERVATION

not:

CANONICAL_WORK

Reconcile evidence, not IDs.


19. RUNNINGTODO REMAINS THE MAP

Do not create a new incident plan beside it.

Map Gate Zero into the existing structure:

MOUNTAIN 2
= Gas City Factory Economy / Execution Discipline

MOUNTAIN 16
= Oracle / Production Runtime Convergence

MOUNTAIN 17
= Durability where applicable

Use existing canonical Convoys and Beads.

Search first.

One root defect should have one bounded owner.


20. INCIDENT EXECUTION ORDER

Execute in this order:

1. CAPTURE CURRENT LOAD / MEMORY / PROCESS BASELINE

2. INVENTORY NATIVE GAS CITY SESSIONS + CLAIMS

3. CLASSIFY THE 32 CLAUDE PROCESSES

4. IDENTIFY PID 4880 CONTAINER / SOURCE OWNER

5. INVENTORY bd/gc CLIENT VERSIONS ACROSS SHARED DOLT CLIENTS

6. IDENTIFY CANONICAL BINARY INSTALLATION OWNER

7. SEARCH EXISTING ORACLE BEADS FOR EACH ROOT DEFECT

8. UPDATE / DEDUP / ROUTE EXISTING WORK

9. BUILD COORDINATED VERSION-CONVERGENCE PLAN

10. APPLY ONLY THROUGH SOURCE/IAC/RELEASE PATH

11. RESTORE NATIVE BEADS

12. VERIFY WORKER CEILING

13. VERIFY LOAD / MEMORY / FORK-RATE RECOVERY

14. WITNESS VERIFY

15. RESUME NORMAL FACTORY EXECUTION

21. GATE ZERO ACCEPTANCE

Do not declare Oracle recovered until:

CONTROLLER=RUNNING
SUPERVISOR_API=HEALTHY

NATIVE_BEADS=AVAILABLE
BDSTORE_FALLBACK=NO

DB_SCHEMA=
BD_SUPPORTED_SCHEMA=
GC_LINKED_BEADS_SCHEMA=
COMPATIBLE=YES

SHARED_CLIENTS_COMPATIBLE=YES

ACTIVE_FACTORY_WORKERS<=POLICY_LIMIT
OR
PROCESS_COUNT_EXPLAINED=YES

ORPHANED_MODEL_PROCESSES=0

LOAD_1M=
LOAD_5M=
LOAD_15M=
LOAD_WITHIN_EXPECTED_RANGE=YES

MEMORY_PRESSURE=NO

ZOMBIE_PARENT_IDENTIFIED=YES
ZOMBIE_ROOT_CAUSE_ROUTED=YES

FORK_RATE=
FORK_RATE_ACCEPTABLE=YES

SOURCE_DELIVERY_PROVEN=YES
RUNTIME_READBACK_PROVEN=YES
WITNESS_VERIFIED=YES

Do not require the zombie defect to be completely fixed before Factory resumption unless it remains operationally dangerous.

It must at least be owned and bounded.


22. ECONOMIC / REUSE CLAIM

Current state:

IS_THE_NEXT_RUN_GETTING_CHEAPER_AND_MORE_REUSABLE=NOT_ESTABLISHED
WHY=The investigation corrected the execution authority and exposed the real Oracle failure modes, but the runtime still uses Beads fallback, the client/schema lifecycle is not converged, and worker-ceiling enforcement has not been proven.

Only change to:

YES

after the next Oracle startup/run obtains the correct binaries, native Beads compatibility, worker limits, and durable routing automatically from source-controlled infrastructure without rediscovery or manual host surgery.


FINAL LAW

DO NOT FIX THE MACHINE.

FIX THE FACTORY LIFECYCLE THAT PRODUCED THE MACHINE.

DO NOT MIGRATE LOCAL TASKS INTO ORACLE BY HAND.

RECONCILE THEM AGAINST ORACLE AUTHORITY.

DO NOT KILL PROCESSES TO MAKE THE GRAPH LOOK HEALTHY.

RESTORE THE CONTROL PLANE THAT PREVENTS THE PROCESS EXPLOSION.

DO NOT FORCE THE DATABASE FORWARD.

CONVERGE EVERY CLIENT THAT SHARES IT.

ORACLE IS THE FACTORY.
THE WORKSTATION IS NOT.