Failure Signature Catalog¶
Authority: Bluefly Engineering (hand-maintained). Scope: Recurring, recognizable failure patterns — named so the next agent that hits one recognizes it faster than the last one did, instead of re-diagnosing it from scratch as a fresh mystery. Provenance: entries below are
SOURCE_VERIFIED(confirmed by inspecting the actual failing request/URL/process, not inferred from the symptom alone) unless marked otherwise.
How to add an entry¶
A failure signature earns a row when it has recurred, or when two failures produced the same symptom but had different root causes and fixes (the whole point of naming it is telling those apart faster next time). A single one-off incident belongs in a dated receipt, not here — see Receipts Standard.
Catalog¶
GitLab job-token-scope inbound allowlist missing¶
- Symptom: Composer reports
Invalid credentials(or an equivalent 401/403) against a project-level registry URL (/api/v4/projects/:id/packages/composer/packages.json). - Root cause: the producer project's CI/CD → Token Access → job-token allowlist does not include the consumer project.
- Fix: add the consumer project to the producer project's job-token inbound allowlist.
- Distinguishing test: the failing URL is project-scoped
(
/projects/:id/...).
Group-level Composer facade doesn't support CI_JOB_TOKEN¶
- Symptom: the identical
Invalid credentialsmessage, but against a group-level Composer facade URL (/api/v4/group/:id/-/packages/composer/packages.jsonor similar). - Root cause: GitLab's group-level Composer package facade does not
accept
CI_JOB_TOKENat all, regardless of allowlist configuration — this is an upstream GitLab capability boundary, not a misconfiguration. - Fix: use the masked
COMPOSER_REGISTRY_TOKENgroup CI/CD variable via HTTP Basic auth instead. Never introduce agitlab-token/COMPOSER_AUTHoverride as a workaround — that reintroduces per-project credential plumbing thatgitlab_componentsis supposed to own centrally. - Distinguishing test: the failing URL is group-scoped
(
/group/:id/...), not project-scoped. - Note: these two signatures were genuinely confused with each other during tonight's convergence work (2026-09-06) — the error message alone is identical; the fix only becomes obvious once the actual failing URL is inspected. Always inspect the URL before choosing a fix.
Duplicate City/Dolt authority¶
- Symptom: dashboards, tools, or agents report inconsistent Beads/Dolt
state depending on which host or session queried them;
bd doctorfingerprint checks disagree between hosts that should be identical. - Root cause: someone ran
gc initon a second host, or cloned a second independent City, instead of connecting to the existing one as a client — see the One-City Model in Gas City Adoption (MAC_CITY_AUTHORITY=NO,LOCAL_DOLT_AUTHORITY=NO). - Fix: retire the duplicate City/Dolt instance; reconnect the affected host as a remote client of the one Oracle-hosted City.
- Distinguishing test:
gc context liston the affected host shows a local/independent City rather than a remote binding to the Oracle endpoint.
Two systems share a name, port, or path segment¶
- Symptom: a finding gathered against one host/store/process is applied
to a different one that happens to share a name, port number, or path
fragment — e.g. a
city.tomlon the workstation and one on Oracle both namedcity, or two repositories both containing a directory literally namedblucity-docs. - Root cause: evidence was recorded without naming the exact host/process/store it came from, so it silently generalized across a boundary it never actually crossed.
- Fix: every finding names its exact source (
SOURCE_STATE: <host, path, or process>vsRUNTIME_STATE: <host, path, or process>) before it is used to justify an action elsewhere. Re-verify immediately before acting if any ambiguity exists about which instance is being discussed. - Distinguishing test: ask "which literal host/process/checkout did this observation come from?" — if the answer is unclear, treat the finding as unverified for any other instance of the same name.
Shared local git checkout used by two concurrent agent sessions¶
- Symptom: a stale-or-fresh
.git/index.lockappears in a checkout another agent session is also actively operating in; one session's uncommitted work risks silent collision with the other's commit/push. - Root cause: two agent sessions were pointed at the same local working-tree path for write access, rather than each having its own clone/worktree off the same remote.
- Fix: one dedicated local clone or worktree per active writing agent session; a shared checkout is read-only-safe for inspection but never a write target for more than one concurrent session.
- Distinguishing test: before any commit in a shared-looking local
path, check
ps aux/ process list for another agent session rooted in the same tree, and re-rungit statusimmediately before writing — don't trust a status check from even a few minutes earlier. - Provenance:
OPERATOR_RELAY— observed directly by the docs-writer session during this same 2026-09-06 directive, against[NAS-ROOT]/BluCity-Docs.
Detector that cannot detect¶
- Symptom: a healthcheck, guard, or exclude-list reports green/pass while the exact defect it exists to catch is actively present.
- Root cause: the check was written or amended so the failure mode it targets literally cannot register — the check measures something adjacent to the real condition instead of the condition itself.
- Instances (2026-09-06):
litellm-proxyhealthcheck: a bare TCP connect, which cannot fail while the port stays bound — green for 3 days at 100% error rate underneath.gc-site-bind/cloud-init fstab guard: greps for mount-point presence only, so it can never detect wrong options — which was the only thing ever actually wrong.- Composer
packages.drupal.orgexclude list: 8 entries where doctrine specifies 5. The 3 extras were deliberately added to an existing list, each addition making the detector blind to one real vendor violation. This variant is the harder one to catch in review, because it reads as ordinary list maintenance in the diff rather than a weakened-by-omission check built wrong from the start — the first two instances were never built to catch the failure; this one was actively extended to stop catching it. - Misleading-remediation subtype, more severe than the three above:
the doctor's own remediation text for a missing blucity-docs rig
binding was
gc rig add /opt/bluefly/BluTown-docs --name blutown-docs— three stacked defects: (1) wrong path (only lowercaseblutown-docsexists on Oracle in any casing), (2)gc rig addvalidates the entire city before writing, so per hq-ztdb it always fails on whichever other rig is unbound at the time and writes zero bytes (4 of 4 invocations confirmed failing this way on Oracle 2026-08-26) — and (3), the dangerous one: even if the command somehow ran,--name blutown-docswould rename the rig away from the canonicalblucity-docslogical name, actively reverting an operator ruling rather than fixing anything. This doesn't merely fail to detect a problem — it hands the operator a command structurally guaranteed to fail, and would cause real regression if it ever succeeded. Fix all three defects together or none: fixing only the path while leaving--namewould convert a harmless broken command into a working one that reverts the ruling. - Fix: verify a check by first forcing the exact failure it claims to catch and confirming it actually goes red. A check that has never been seen to fail is unverified, not proven reliable.
- Distinguishing test: ask "what change would make this check fail?" If the answer is narrower than the actual failure mode it's named for, it's this pattern.
Self-reported state accepted as measured¶
- Symptom: a system or session's own status report is treated as verified fact and drives escalation or investigation, without anyone asking "how do you know?"
- Root cause: a self-declared constraint or capability claim is carried forward — often across a context-compaction/summary boundary — and never re-tested before being acted on. A false report can sound more confident than the truth it's hiding.
- Instance (2026-09-06): a session identifying as "FOUNDRY" reported
capability-zero (no Bash, no git, no glab, no alternate GitLab path),
classified
AGENT_ROLE_PROVIDER_CAPABILITY_MISMATCH/HIGH, triggering a multi-agent investigation into Gas City provider/agent.toml config. The same session then retracted it: it had never actually tested Bash directly, and was carrying forward a stale self-declared constraint from before a context summary. Once tested, Bash, git, glab (viaop run), node, npm, pnpm, andEnterWorktreewere all real and working — confirmed by then using them to fix a real defect (duplicate Commander.jsexportcommand collision inopenstandardagents), get a greenopenstandardagents!890pipeline, and a verified merge. - Fix: treat any self-reported capability, completion, or state claim as unverified until confirmed by an independent check (a real command run, an API call, a second party) — especially immediately after a context-compaction or session-resume boundary, where stale constraints are most likely to be carried forward unexamined.
- Distinguishing test: was the claim ever actually tested in this session, or only asserted/inherited? If untested, it is not measured state.
Absence proven by a single mechanism is not absence¶
- Symptom: a
NOT_FOUND/"doesn't exist" conclusion is drawn from one probe returning empty — and that empty result reads as more confident evidence than a populated one would have. - Root cause: the probe mechanism itself silently failed to reach the real data (a blobless/shallow clone, a wrong CLI subcommand, a truncated default page size) and returned empty rather than an error, so the empty result was accepted as a measurement instead of a non-answer.
- Instances (2026-09-06):
- A blobless clone's first probe of a repo returned an empty
composer.json, nearly filed as "the BOM metapackage doesn't exist" — caught only by re-verifying withls-treeacross all three refs, which showed the file was real. gc agents(wrong subcommand) returned zero results, nearly filed as "FOUNDRY is not a Gas City agent" — caught only by retrying with the correctgc agent.- Instance (2026-09-09):
qmd's search index and GitLab's own blob search were both confirmed stale during the same upstream-doc reconciliation pass.qmdwas 5 days stale (71% orphaned embedding chunks). GitLab blob search returned a hit for "Health Patrol" text inside an ADR file that, read directly, contained no such text at all -- a false positive, not just a missed true positive, from the same staleness cause. A search-index HIT still needs the live file read to confirm it's current; a search-index MISS never proves absence on its own. - Fix: confirm any absence/NOT_FOUND claim with a second, differently-shaped probe before filing it. A single mechanism's silence is not proof of absence.
- Distinguishing test: could the probe mechanism itself produce an empty result on a real, populated target (wrong subcommand, shallow clone, default pagination, wrong scope)? If yes, the NOT_FOUND is unverified until a second mechanism agrees.
- Note: a live instance of this same class occurred in this channel
during this same session (2026-09-06) — an MR's
state: mergedplus a green pipeline was accepted as proof of current branch content formcp_gateway'scomposer.json, without re-checking the live branch tip. A hotfix commit 26 minutes after the merge had already reverted the change; the true current state only surfaced when a second party re-measured against fresh clones. "Merged" is a historical fact about one commit, not a live-content claim — always re-read the current branch tip before asserting what a repo's content is now.
Closed work as false authority¶
- Symptom: a closed bead, MR, or ticket reads as settled and is treated as more reliable than live content — a closed record is almost never re-verified, which makes a wrong closed record more dangerous than an open one. Nobody argues with a closed ticket.
- Root cause: the record's outcome was true at close time (or was never re-checked at all), but the underlying source was later reverted or changed, and nothing reopened or annotated the closed record to reflect that. This inverts "chat is not the work graph" — the graph is supposed to outrank memory, but a wrong closed record turns it into the least-questioned source of a wrong answer.
- Instance (2026-09-06): bead
bc-1xn(Mac-local store, not Oracle canonical) is closed on the premise thatmcp_registryshould be renamed tobluefly/mcp_gateway. That rename shipped (mcp_gateway!136), then was reverted 25 minutes later as wrong — composer name must match the actual Drupal module machine name (mcp_registry.info.yml), which was never renamed. The bead's recorded outcome is the opposite of the real end state. Two independent agents tonight nearly re-proposed the identical breaking rename on that bead's authority before either checked live branch content. - Fix: a closed bead asserting a source change should cite the exact SHA that proves it, so a later reader can check rather than trust. Reverting a commit named in a closed bead should reopen or annotate that bead. Closed beads asserting source outcomes are worth a proactive audit rather than waiting to trip over more of them.
- Distinguishing test: does the closed record cite a checkable SHA, or only a description of the intended outcome? If only the latter, treat it as unverified until the live branch tip is read directly.
Fault blocks its own remediation path¶
- Symptom: the tooling needed to record, broadcast, or verify a fix is
itself degraded by the fault being fixed — the bead note that cannot be
written because writes are failing, the
gc session listthat times out so the "stop retrying" rule cannot be nudged out, thegc mailthat fails while trying to coordinate the mail outage. - Root cause: remediation depends on the same substrate as the fault (one shared Dolt server behind Beads, mail, hooks, and session listing).
- Instances (2026-09-20): Oracle Dolt write starvation — see
ledger/evidence/incidents/2026-09-20-oracle-dolt-write-starvation.md. Three in one night: bead append, session nudge, mail send. - Fix: keep a durable record outside the failing substrate (a file on the host, a cross-session message) and say so explicitly; verify every write by read-back; do not retry into a store whose failures may commit.
- Distinguishing test: ask "does recording this finding require the thing that is broken?" If yes, write it somewhere else first.
Stale record manufactures work¶
- Symptom: several independent sessions escalate, for hours, for an action that was already executed and proven ineffective.
- Root cause: a retraction did not reach every projection the original claim reached. A shared agent-memory file kept a retracted attribution and the words "containment never applied"; each new session read it as established, added urgency, and never re-tested.
- Instance (2026-09-20): four sessions between 07:24Z and 11:16Z asked
for
systemctl stop gascity-dashboard.service, which had been executed 05:43:56Z, measured with no effect, and rolled back 05:56:27Z. TheExecMainStartTimestampof the rollback was read as proof the stop never happened. Corrected at source with the original preserved. - Fix: a retraction is complete only when it has overwritten or
annotated every copy — bead, ledger, shared memory, session handoff.
Memory entries carry an evidence class and a supersession pointer;
unverified memory is not runtime truth. Defect:
STALE_AGENT_MEMORY_CAUSES_REPEATED_FALSE_REMEDIATION. - Distinguishing test: is the "urgent" action's timestamp evidence consistent with it never having run, or with it having run and been reverted? Check the runtime before escalating.
Config export on a host-mounted checkout silently drops files¶
- Symptom:
drush config:exportintoconfig/syncreports success anddrush config:statusimmediately reports "no differences", yet minutes later some.ymlfiles are gone from the sync directory;git statusshows deletions nobody made. Losses vary run to run (198, 4, 12, 0, 0, 0 files across six exports on 2026-09-20; earlier 129, 242 and 18). - Root cause (INFERRED, not yet proven on a clean host): the export
target is a virtiofs mount of the Mac checkout inside the DDEV web
container (OrbStack,
performance_mode: none). Drush deletes every file in the target (FileStorage::deleteAll) and rewrites ~1,500 files; a host-side process concurrently rewrites.git/indexwhile no git runs in the container; a 250 ms readdir+open poll during the export listed 622 entries with 99 unopenable (ENOENT), and transient ENOENT persisted ~2 s after drush exited before settling with files permanently missing. The same export into 17 other directories on the same mount lost nothing. Ruled out: DB reads (drush throws on a failed read; every lost object was present in the config table),config_split(4 splits, all inactive), monitoring churn (fixed separately), coreDatabaseStorage. - Instance: bluefly.io, 2026-09-20, Drupal 11.4.7, drush 13.8; bead
bluefly-9yp.3.7. - Fix / mitigation: DDEV
performance_mode: mutagenso writes land on a native volume; identify the host watcher touchingconfig/sync; if it persists on a clean host, report upstream to OrbStack with this reproduction. No wrapper scripts around drush. - Distinguishing test: after
drush cex, wait several seconds, thengit statusmust show only the intended changes before committing.drush cstalone is insufficient: it reads the files before they vanish.
Transitional agent creates a second checkout of an authority repo¶
- Symptom: two clones of the same authority repository on one
workstation (
$ESTATE_ROOT/BluCity-Docsand$ESTATE_ROOT/DEMOs/blucity-docs); the second has an HTTPSgitlab.comremote instead of the estate'sgit@gitlab-bluefly:alias; commits land under the operator's identity from a checkout nobody governs. - Root cause: a transitional (non-authority) agent session cloned relative to its own working directory instead of resolving the estate's canonical checkout, and the tool's bootstrap did not name that checkout or the ssh remote.
- Instance (2026-09-20): Antigravity session cloned blucity-docs at
00:19 -0400, committed and pushed
feature/agent-capability-estate(5-line governance edit, later merged). Clone trashed at 15:03Z after proof it was clean and fully pushed. - Fix: transitional-agent bootstraps (Antigravity, Gemini, Codex) must point at the canonical checkout and use the ssh remote; a clone of an authority repo anywhere else on the workstation is a defect, not a convenience. Preserve unique commits (push or patch) before trashing.
- Distinguishing test:
find ~/Sites -maxdepth 3 -type d -name .git -path '*blucity-docs*'returns more than one path, orgit -C <clone> remote get-url originstarts withhttps://.
Idempotency check does not describe the rule it guards¶
- Symptom: a "check-then-insert" guard (
iptables -C … || iptables -I …,grep -q … || echo … >>,kubectl get … || kubectl create …) runs on every unit start, timer tick, or reconcile loop and the guarded resource grows monotonically: thousands of identical firewall rules, repeated config lines, duplicate CRDs. Cost shows up elsewhere first —iptables-savetakes seconds, every caller that walks the chain (k3s, kube-router, docker) burns CPU, load rises with no single hot process. - Root cause: the check's spec is not byte-identical to the insert's
spec.
iptables -Cmatches the whole rule, so a check that omits-m comment --comment <tag>never matches a rule inserted with one; the guard is a no-op and the loop is unbounded. The defect is invisible at one execution and only compounds under restart or reconcile churn. - Instance (2026-09-20):
gastown-gateway.serviceExecStartPreon Oracle — NRestarts 43,301 × 2 CIDRs = 84,507 duplicate INPUT rules, INPUT chain 84,930 rules,iptables-save9.6 s at 64 % CPU, ≈1.5 of 4 cores spent in k3s iptables walks. The repo template (agent-dockerdeployments/gastown/systemd/) carried the same guard. Same class, separate instance: 417 duplicatekube-router netpolrules. - Fix: the check must describe exactly the thing the insert creates —
same match modules, same comment, same target. Where the tool offers an
atomic idempotent form (
iptables-restoreof a managed chain,kubectl apply,ensure-style APIs), prefer it over check-then-insert. Never pair an unboundedRestart=with a non-idempotentExecStartPre. Source fix: agent-docker MR !399; host flush is a separate authorized step. - Distinguishing test:
iptables -S | sort | uniq -d | wc -lis non-zero, or NRestarts × inserts-per-start equals the rule count. More generally: run the guard twice and count the resource — any growth on the second run is this signature.
discover.duadp.org 502 Bad Gateway¶
- Symptom: discover.duadp.org returns 502 Bad Gateway via Cloudflare.
- Root cause: The origin daemon on port 4201 is not running (localhost:4201 is down).
- Fix: Start or restart the discover daemon on the origin host on port 4201.
- Distinguishing test: Confirm Cloudflare tunnel is healthy but
curl -I localhost:4201on the origin fails.
Qdrant Multi-Instance Violation¶
- Symptom: Qdrant data is missing or out of sync between agents, or lock file conflicts occur.
- Root cause: Multiple Qdrant instances are running instead of a single authoritative agent-brain instance. Qdrant ownership belongs to agent-brain, NOT agent-studio.
- Fix: Terminate rogue Qdrant instances and route all vector operations through the canonical agent-brain service.
- Distinguishing test: Check for multiple qdrant processes or containers on the host.
World-Writable .env File (P0 Security)¶
- Symptom: .env files have 777 or 666 permissions.
- Root cause: Misconfigured deployment script or manual permission change.
- Fix: chmod 600 or 640 on the .env file.
- Distinguishing test:
ls -l .envshows world-writable permissions.