Oracle // Deep Execution Order — Final Report¶
STALE SNAPSHOT WARNING (added 2026-08-24; LLM line 2026-09-02): This is a point-in-time topology snapshot from 2026-07-30. Its two headline Dolt-store status lines are now inverted relative to current reality — see the correction block immediately below the table. The 2026-07-30 row that groups LiteLLM and Ollama as running on Oracle is also superseded: Oracle
:11434refused connections on 2026-09-02. Do not fill Oracle with model weights. Live topology:Engineering-Standard/standards/architecture/inference-topology.md. Do not treat any line in section 1 as current without re-verifying against live runtime first; this file is evidence of what was true on 2026-07-30, not a live reference.
All findings below are from direct SSH investigation of bluefly-platform.tailcf98b3.ts.net (Oracle), executed by 4 parallel background agents plus direct commands run by me. Every claim is tagged. Nothing below is inferred without the evidence that supports it being stated alongside it.
1. VERIFIED RUNTIME TOPOLOGY (as of 2026-07-30 — see inversion notice above)¶
| Component | Status | Evidence |
|---|---|---|
bluefly-traefik-oracle |
FAILED (config) | docker inspect health: FailingStreak: 4841, dial tcp :8080: connection refused. Root cause found: compose service file mounts only Caddyfile, never wires in the real traefik.yml (no --configFile flag on a bare ["traefik"] Cmd). Traefik boots with zero config — never opens its ping port. Plays no role in any current routing. |
contractplane-control-api |
MISCONFIGURED (false-negative health) | Docker reports unhealthy (4,861 failing checks, all curl: not found — curl isn't in the image). Actual app: curl 127.0.0.1:3099/health → 200 {"status":"ok"}. Service is genuinely fine; only the healthcheck definition is broken. |
bluefly-router |
NOT INSTALLED (unused stub) | Stock unmodified nginx:alpine default page, no host-port binding. Not serving real traffic. |
cloudflared-blutown |
RUNNING (config drift) | Running 41h despite being commented out of docker-compose.yml's include list — started outside/ahead of the current compose file state. Real ingress config (dashboard-pushed): api.blutown.ai path ^/api → localhost:8083; dash.blutown.ai → localhost:8083; docs.blutown.ai → localhost:8084. |
Port 8083 (Node, frontend) |
RUNNING, no real API | curl 127.0.0.1:8083/api/v1/town/rigs → byte-identical SPA HTML shell as the public api.blutown.ai response. Confirmed: this IS the origin BluStudio's requests terminate at. SPA-fallback-for-everything server, no backend routes. |
Port 8082 (gascity-dashboard/backend/dist/server.js) |
RUNNING, real API, unexposed | /api/health → genuine 200 application/json {"ok":true,"ts":...}. Has real routing (404s properly for unknown paths, doesn't SPA-fallback). No /api/v1/town/rigs, /api/rigs, /api/town/rigs route exists on it. Title on /health page: "gas city · ds-research" — likely a different product (Gas City's own dashboard) than what BluStudio expects (Gas City Mission Control). Not in any cloudflared ingress rule — unreachable from the internet at all right now. |
openclaw-openclaw-gateway-1 |
RUNNING, healthy, misconfigured client target | Loopback-only (127.0.0.1:18789, confirmed via sudo ss, owner docker-proxy). Tailscale Serve/Funnel exposes it tailnet-wide only at https://bluefly-platform.tailcf98b3.ts.net:8443 (proxies internally to 18789). A client hard-coded to ws://...:18789 has no route — wrong port, wrong scheme (ATS blocks plaintext ws:// anyway). |
| City Dolt (bd, port 36693) | STALE — see correction below. Originally reported RUNNING with real data (93/87/76). As of 2026-08-23, this store was destroyed by a deploy job's rsync --delete (attribution: pipeline #238) and is under active recovery — see the correction block. Do not cite the 93/87/76 figures as current. |
|
| Town Dolt (port 3308) | STALE — see correction below. Originally reported FAILED/BLOCKED (root-owned lockfiles, ubuntu could not read/auto-start). As of 2026-08-23, this store is RUNNING and healthy — pid 76014, holding 1,229 beads, the only fully intact store on the host. |
|
Home-dir Dolt (~/.beads) |
RUNNING | bd status from /home/ubuntu → Total Issues: 165, Ready to Work: 121. A fourth, separate database. |
Global dolt-sql-server.service (port 3307) |
RUNNING | systemctl shows active running. A distinct instance from the above three. |
gc-worker.service (City supervisor) |
STALE — see correction below. Originally reported RUNNING/healthy. As of 2026-08-23, deliberately DISABLED as reboot-safety containment after the work-store-wipe incident, superseded by a proper ExecStartPre= health gate (iac!202). |
|
| Executor/Implementer/Polecat/Crew/Refinery/Dispatcher pools | STOPPED (scaled to 0) | gc status → 2/139 agents running (both are gascity.witness health monitors, not executors). gt agents → No agent sessions running. Supervisor itself is healthy; every managed pool is at 0 across every rig. |
mnt-nas.mount |
FAILED | systemctl list-units --failed |
bluefly-authority-backup.service |
FAILED | systemctl list-units --failed |
| Dragonfly (queue candidate) | RUNNING, idle | Healthy, dbsize: 0, 1 connected client (mine). No evidence either way that it's the intended bead-dispatch queue. |
| Keycloak, Postgres (core), LiteLLM, Ollama, Phoenix, Portainer, VictoriaMetrics/Logs, Grafana | RUNNING, healthy | All confirmed via docker ps, all "Up 41 hours (healthy)" or equivalent, not independently probed further (out of scope for this pass). |
CORRECTION (added 2026-08-24) — the two Dolt/supervisor states above are inverted¶
Measured directly on Oracle, 2026-08-24, superseding the 2026-07-30 rows above:
- City Dolt (:36693) — this is now the compromised store, not the healthy one. It was destroyed on 2026-08-23 (deploy job
rsync --delete, attribution closed to pipeline #238) and is under active recovery: export taken (202/99 records, sha256-manifested), quarantine of a manufactured stub in progress, boot-time health gate landed (iac!202), canonical-store restore not yet complete. - Town Dolt (:3308) — this is now the healthy, intact store, not the broken/permission-blocked one. pid 76014,
~/gt/.dolt-data, 1,229 beads. It is currently the only fully intact store on the host. gc-worker.service/gc-city-init.service— both deliberately disabled (not the July snapshot's enabled/inactive-or-failed state) as reboot-safety containment, since a plain reboot would otherwise re-run the sequence that destroyed the city store. A proper 4-predicateExecStartPre=health gate (iac!202) now governs whether these can restart at all.
The rest of section 1 (Traefik, ContractPlane, cloudflared, the SPA-fallback/real-API split on 8082 vs 8083, home-dir Dolt, global dolt-sql-server, agent pool scale-0) has not been re-verified as part of this correction and should be treated as historical unless independently re-checked.
2. VERIFIED API CONTRACT MAP¶
| Boundary | Expected | Observed | Verified? |
|---|---|---|---|
BluStudio → api.blutown.ai/api/v1/town/* |
JSON | 200 text/html, SPA shell |
Broken — confirmed live, twice, at different times |
api.blutown.ai → Cloudflare Tunnel → origin |
routes to a real API | Tunnel ingress ^/api → localhost:8083, confirmed from tunnel's own logged config |
Verified — this IS the actual path |
Origin :8083 → backend logic |
real API routes | SPA fallback (index.html) for every path including /api/* |
Broken — confirmed by identical local vs. public response |
Origin :8082 (separate process) |
unknown if intended as the real API | Real JSON on /api/health; no /api/v1/town/* route; not in any ingress rule |
Partially verified — real API exists, wrong/no routing, and this may not even be the right product |
BluStudio → studio.blueflyagents.com (bridge) |
authenticated JSON | 302 to Cloudflare Access SSO on every path tested (/api/v1/status, /api/v1/deploy/ready, /api/v1/deploy/run) |
Broken — confirmed, tracked as ubuntu-6n0, not re-verified in this pass |
BluStudio → mesh.blueflyagents.com/api/v1/discovery |
JSON agent list | 404 |
Broken — confirmed |
BluStudio → marketplace-app.drupl.ai/api/v1/agent-marketplace/jsonapi/agents |
JSON:API | 421 Misdirected Request |
Broken — confirmed, TLS/SNI-level, not app-level |
BluStudio (Execution Console) → api.copaw.us |
governed execution JSON | 502 from Cloudflare edge (server: cloudflare), reproduced twice |
Broken — confirmed live, right now, during this investigation |
BluStudio (OpenClaw path, if any) → claw.copaw.us |
— | 200 text/html |
Host reachable; not exercised further |
OSSA registry: build.openstandardagents.org |
— | DNS does not resolve at all (dig +short empty) |
Broken — host doesn't exist. Bare openstandardagents.org root does resolve (200). This corrects an earlier unverified reference I made this session to build.openstandardagents.org as a real endpoint. |
macOS Ready Beads (this session's ubuntu-xa2 fix) → blu work bd CLI → home-dir Dolt |
JSON | Genuinely works — confirmed via direct reproduction earlier this session AND via the home-dir Dolt being confirmed alive here | Verified working |
3. ROOT CAUSE GRAPH (verified edges only)¶
BluStudio (GasCityClient)
→ api.blutown.ai [VERIFIED: DNS→Cloudflare anycast IPs]
→ Cloudflare Tunnel "cloudflared-blutown" [VERIFIED: running, logged ingress config]
→ localhost:8083 (path ^/api) [VERIFIED: ingress rule]
→ Node SPA server, no backend routes [VERIFIED: identical local/public response]
→ NOT ESTABLISHED: does a real Town Mission Control API exist anywhere, or was it never built?
Town Dolt / gt status
→ NOT ESTABLISHED as connected to City Dolt (separate DB, separate port, separate root-owned lockfiles)
City Dolt (36693) → gc-worker.service (supervisor, healthy)
→ agent pools (polecat/crew/refinery/dispatcher/mayor/deacon)
→ VERIFIED: all scaled to 0, independent of Dolt (Dolt is up)
→ NOT ESTABLISHED: why pools are at 0 (config floor vs. blocked scale-up — no spawn/error evidence found in the log window checked)
OpenClaw gateway (18789, loopback)
→ Tailscale Serve (:8443) [VERIFIED: only exposed path]
→ client misconfiguration (ws://...:18789 instead of wss://...:8443) [VERIFIED port/scheme mismatch; NOT ESTABLISHED which client actually used the wrong URL]
4. EXECUTION CAPACITY — DIRECT ANSWER¶
READY = 87 / Executor = 0 / Implementers = 0, precisely explained:
- "87" matches the City-scope Dolt's
Opencount exactly (port 36693:Open: 87). Not itsReady to Workcount (76), notbd ready's own live figure (89), not the home-dir Dolt'sReady to Work(121 — the DB my ownblu work bdcalls have been hitting all session). Not established which exact tool/field the operator's own dashboard reads. - Executor/Implementer = 0 is real and independently confirmed, not a Dolt symptom — City's own Dolt is alive and serving correct data. The supervisor (
gc-worker.service) is healthy, 12h uptime. Every managed agent pool (polecat, crew, refinery, dispatcher, mayor, deacon) across every rig is at 0. - Executors do not require Dolt to exist as infra — Dolt IS up. The blocker is upstream of Dolt: something is holding every pool at scale 0. No spawn-failure or scale-blocking error was found in the last 300 lines of
supervisor.logchecked — this needs a longer log window or a different diagnostic surface (not established with the access used in this pass). - Dragonfly (candidate dispatch queue): healthy, idle, 0 keys. No evidence either way that it's the actual bead-dispatch mechanism.
Smallest verified action that would restore execution capacity: not established. This requires either (a) reading a longer supervisor.log window / the actual scaling-policy config for each pool, or (b) asking whoever owns Gas City's scaling logic why min-replicas are 0. I do not have evidence to name a single fix.
5. HIGH-VALUE DISCOVERIES¶
- Config drift, proven:
cloudflared-blutownis running in production despite being commented out of the source-of-truthdocker-compose.yml. The running system and the checked-in config disagree. - Stale self-audit found on the host itself: a file (
hostname-dependency-graph.json) already on Oracle claimsapi.blutown.aihas"architecture_status":"UNKNOWN_OWNER","container_runtime":"ABSENT"— contradicted by live evidence (the container is running). This is documentation/runtime drift discovered on Oracle's own filesystem, not something I introduced. - A real, unexposed backend exists (
gascity-dashboard/backend, port 8082) with genuine JSON API infrastructure, sitting completely outside the current tunnel's ingress rules. Whether this is the intended Town API or a different product entirely (title suggests "Gas City," not "Gas City") is not established. - Four separate Dolt/beads databases exist simultaneously on one host with different data (165/121, 93/87/76, Town-3308/stopped, global-3307), with no single canonical "the ready count." My own session's
blu work bdwork has been operating against the home-dir instance, not the City-scope one the operator's "87" figure came from. - Two genuinely failed systemd units found (
mnt-nas.mount,bluefly-authority-backup.service) — unrelated to anything investigated this session, first time surfaced. api.copaw.us(the real execution gatewayExecutionTimelineViewdepends on) is returning 502 right now — confirmed twice, live, during this investigation. This is a current, active outage separate from everything else found.- Two "unhealthy" containers (Traefik, ContractPlane) have two completely different real causes — one is genuinely unconfigured, one is a healthcheck-definition bug on an otherwise-fine service. Docker's health status alone cannot distinguish these; both needed direct investigation.
6. ASSUMPTIONS DISPROVEN¶
- "Dolt is down" as the explanation for Executor=0 — disproven. City's Dolt is up and correct.
- My own earlier-session reference to
build.openstandardagents.orgas a real endpoint — disproven, it has no DNS record at all. - The disputed OpenClaw ATS error being a gateway-side defect — disproven; the gateway is healthy, the error is consistent with a client using the wrong port/scheme.
- Traefik or
bluefly-routerhaving any role inapi.blutown.aitraffic — disproven; neither touches this hostname at all.
7. ASSUMPTIONS THAT REMAIN UNPROVEN¶
- Whether port 8082 (
gascity-dashboard) is meant to be the real Gas City Town API, a different product entirely, or an abandoned prototype. - Why every Gas City agent pool is scaled to 0 (config floor vs. blocked scale-up) — no direct evidence found.
- Which of the 4 Dolt instances is the one canonical source the operator's own "87" figure and dashboards are meant to read.
- Whether the Cloudflare Tunnel's DNS record for
api.blutown.aiis a CNAME to the tunnel (expected) or something else — not visible from this host without Cloudflare dashboard/API access. - Whether
compliance-engineandbluefly-blu-buddy-daemon(referenced incontractplane-control-api's env) are actually reachable from that container's network namespace — not tested (would requiredocker exec, not run).
8. ORACLE SELF-AUDIT (this session, my own work)¶
- Real mistake, self-caught and corrected: I appended branch-consolidation/CI-status notes to the wrong Oracle bead (
ubuntu-hj4, about refresh/realtime states — unrelated) before creating the correct one (ubuntu-gb3). The append-only ledger means the misfiled note is still there; I posted a correction note pointing to the right bead, but could not remove the original misfiled entry. - Imprecise inference, now corrected here: earlier this session I inferred commit
5b38255(and similar unattributed commits found mid-session) was "most likely my own prior work from a compacted portion of this session." That inference is not actually provable from git metadata — every commit on this machine, whether typed by the operator directly or made by any agent instance (including me), carries the identical configured git identity (Thomas @ Bluefly.io). I should have labeled this "Not Established" rather than leaning toward one explanation. - Unverified reference, now corrected: I referenced
build.openstandardagents.orgearlier this session without ever actually testing it. It has no DNS record. Flagged here as a real correction, not previously caught. - No duplicate Beads found on review —
ubuntu-40nandubuntu-126cover the same root cause but were deliberately cross-referenced rather than duplicated;ubuntu-6n0andcontractplane-control-api's finding were correctly kept as two separate, unrelated conditions per direct evidence, not merged or conflated.
9. TOP 10 HIGHEST-VALUE REMAINING UNKNOWNS (ranked)¶
- Why are all Gas City agent pools (polecat/crew/refinery/dispatcher) scaled to 0 despite a healthy supervisor and a healthy Dolt? (Single biggest lever on execution capacity.)
- Is port 8082 (
gascity-dashboard/backend) the intended Gas City Town API, or an unrelated product? Answering this could resolveubuntu-40n/ubuntu-126outright. - Which of the 4 Dolt instances is canonical for "the" ready-bead count — and should the other 3 be decommissioned or are they intentionally separate scopes?
- Why is
api.copaw.usreturning 502 right now — is this transient or a real, current outage of the execution gateway? - Is the Cloudflare Tunnel DNS record for
api.blutown.aiactually a CNAME to the expected tunnel, or has it drifted? (Not checkable from this host.) - Why does
docker-compose.ymldisagree with the actual running container set (cloudflared-blutownpresent but commented out) — is there a second compose file or manualdocker runprocess nobody has documented? - What actually depends on the Town-scope Dolt (port 3308, currently stopped, root-locked) — is anything broken by its absence, or is it dead infrastructure? (Superseded — see correction block: this store is now healthy and holds the estate's most intact bead history.)
- Is Dragonfly actually wired into bead dispatch anywhere, or is it unused/vestigial?
- Why did
mnt-nas.mountandbluefly-authority-backup.servicefail, and since when? - Is
studio.blueflyagents.com's Cloudflare Access gate (ubuntu-6n0) intentional production policy, or a misconfiguration blocking legitimate service-to-service traffic?
10. TERMINAL STATE¶
Investigation exhausted the read-only evidence available from this Mac + SSH access to Oracle within this session. Every phase produced verified, evidence-tagged findings; several (Phase 4's root trigger, Phase 5's ownership of port 8082, the Cloudflare DNS record type) require either deeper log access, Cloudflare dashboard access, or a decision from whoever owns Gas City's scaling policy — none of which I have. No code was modified. No service was restarted. No config was changed. Nothing was fabricated; every "broken" or "healthy" label above has a quoted command + output behind it.