Skip to content

Failure Signature Catalog

Authority: Bluefly Engineering (hand-maintained). Scope: Recurring, recognizable failure patterns — named so the next agent that hits one recognizes it faster than the last one did, instead of re-diagnosing it from scratch as a fresh mystery. Provenance: entries below are SOURCE_VERIFIED (confirmed by inspecting the actual failing request/URL/process, not inferred from the symptom alone) unless marked otherwise.

How to add an entry

A failure signature earns a row when it has recurred, or when two failures produced the same symptom but had different root causes and fixes (the whole point of naming it is telling those apart faster next time). A single one-off incident belongs in a dated receipt, not here — see Receipts Standard.

Catalog

GitLab job-token-scope inbound allowlist missing

  • Symptom: Composer reports Invalid credentials (or an equivalent 401/403) against a project-level registry URL (/api/v4/projects/:id/packages/composer/packages.json).
  • Root cause: the producer project's CI/CD → Token Access → job-token allowlist does not include the consumer project.
  • Fix: add the consumer project to the producer project's job-token inbound allowlist.
  • Distinguishing test: the failing URL is project-scoped (/projects/:id/...).

Group-level Composer facade doesn't support CI_JOB_TOKEN

  • Symptom: the identical Invalid credentials message, but against a group-level Composer facade URL (/api/v4/group/:id/-/packages/composer/packages.json or similar).
  • Root cause: GitLab's group-level Composer package facade does not accept CI_JOB_TOKEN at all, regardless of allowlist configuration — this is an upstream GitLab capability boundary, not a misconfiguration.
  • Fix: use the masked COMPOSER_REGISTRY_TOKEN group CI/CD variable via HTTP Basic auth instead. Never introduce a gitlab-token/COMPOSER_AUTH override as a workaround — that reintroduces per-project credential plumbing that gitlab_components is supposed to own centrally.
  • Distinguishing test: the failing URL is group-scoped (/group/:id/...), not project-scoped.
  • Note: these two signatures were genuinely confused with each other during tonight's convergence work (2026-09-06) — the error message alone is identical; the fix only becomes obvious once the actual failing URL is inspected. Always inspect the URL before choosing a fix.

Duplicate City/Dolt authority

  • Symptom: dashboards, tools, or agents report inconsistent Beads/Dolt state depending on which host or session queried them; bd doctor fingerprint checks disagree between hosts that should be identical.
  • Root cause: someone ran gc init on a second host, or cloned a second independent City, instead of connecting to the existing one as a client — see the One-City Model in Gas City Adoption (MAC_CITY_AUTHORITY=NO, LOCAL_DOLT_AUTHORITY=NO).
  • Fix: retire the duplicate City/Dolt instance; reconnect the affected host as a remote client of the one Oracle-hosted City.
  • Distinguishing test: gc context list on the affected host shows a local/independent City rather than a remote binding to the Oracle endpoint.

Two systems share a name, port, or path segment

  • Symptom: a finding gathered against one host/store/process is applied to a different one that happens to share a name, port number, or path fragment — e.g. a city.toml on the workstation and one on Oracle both named city, or two repositories both containing a directory literally named blucity-docs.
  • Root cause: evidence was recorded without naming the exact host/process/store it came from, so it silently generalized across a boundary it never actually crossed.
  • Fix: every finding names its exact source (SOURCE_STATE: <host, path, or process> vs RUNTIME_STATE: <host, path, or process>) before it is used to justify an action elsewhere. Re-verify immediately before acting if any ambiguity exists about which instance is being discussed.
  • Distinguishing test: ask "which literal host/process/checkout did this observation come from?" — if the answer is unclear, treat the finding as unverified for any other instance of the same name.

Shared local git checkout used by two concurrent agent sessions

  • Symptom: a stale-or-fresh .git/index.lock appears in a checkout another agent session is also actively operating in; one session's uncommitted work risks silent collision with the other's commit/push.
  • Root cause: two agent sessions were pointed at the same local working-tree path for write access, rather than each having its own clone/worktree off the same remote.
  • Fix: one dedicated local clone or worktree per active writing agent session; a shared checkout is read-only-safe for inspection but never a write target for more than one concurrent session.
  • Distinguishing test: before any commit in a shared-looking local path, check ps aux / process list for another agent session rooted in the same tree, and re-run git status immediately before writing — don't trust a status check from even a few minutes earlier.
  • Provenance: OPERATOR_RELAY — observed directly by the docs-writer session during this same 2026-09-06 directive, against [NAS-ROOT]/BluCity-Docs.

Detector that cannot detect

  • Symptom: a healthcheck, guard, or exclude-list reports green/pass while the exact defect it exists to catch is actively present.
  • Root cause: the check was written or amended so the failure mode it targets literally cannot register — the check measures something adjacent to the real condition instead of the condition itself.
  • Instances (2026-09-06):
  • litellm-proxy healthcheck: a bare TCP connect, which cannot fail while the port stays bound — green for 3 days at 100% error rate underneath.
  • gc-site-bind/cloud-init fstab guard: greps for mount-point presence only, so it can never detect wrong options — which was the only thing ever actually wrong.
  • Composer packages.drupal.org exclude list: 8 entries where doctrine specifies 5. The 3 extras were deliberately added to an existing list, each addition making the detector blind to one real vendor violation. This variant is the harder one to catch in review, because it reads as ordinary list maintenance in the diff rather than a weakened-by-omission check built wrong from the start — the first two instances were never built to catch the failure; this one was actively extended to stop catching it.
  • Misleading-remediation subtype, more severe than the three above: the doctor's own remediation text for a missing blucity-docs rig binding was gc rig add /opt/bluefly/BluTown-docs --name blutown-docs — three stacked defects: (1) wrong path (only lowercase blutown-docs exists on Oracle in any casing), (2) gc rig add validates the entire city before writing, so per hq-ztdb it always fails on whichever other rig is unbound at the time and writes zero bytes (4 of 4 invocations confirmed failing this way on Oracle 2026-08-26) — and (3), the dangerous one: even if the command somehow ran, --name blutown-docs would rename the rig away from the canonical blucity-docs logical name, actively reverting an operator ruling rather than fixing anything. This doesn't merely fail to detect a problem — it hands the operator a command structurally guaranteed to fail, and would cause real regression if it ever succeeded. Fix all three defects together or none: fixing only the path while leaving --name would convert a harmless broken command into a working one that reverts the ruling.
  • Fix: verify a check by first forcing the exact failure it claims to catch and confirming it actually goes red. A check that has never been seen to fail is unverified, not proven reliable.
  • Distinguishing test: ask "what change would make this check fail?" If the answer is narrower than the actual failure mode it's named for, it's this pattern.

Self-reported state accepted as measured

  • Symptom: a system or session's own status report is treated as verified fact and drives escalation or investigation, without anyone asking "how do you know?"
  • Root cause: a self-declared constraint or capability claim is carried forward — often across a context-compaction/summary boundary — and never re-tested before being acted on. A false report can sound more confident than the truth it's hiding.
  • Instance (2026-09-06): a session identifying as "FOUNDRY" reported capability-zero (no Bash, no git, no glab, no alternate GitLab path), classified AGENT_ROLE_PROVIDER_CAPABILITY_MISMATCH/HIGH, triggering a multi-agent investigation into Gas City provider/agent.toml config. The same session then retracted it: it had never actually tested Bash directly, and was carrying forward a stale self-declared constraint from before a context summary. Once tested, Bash, git, glab (via op run), node, npm, pnpm, and EnterWorktree were all real and working — confirmed by then using them to fix a real defect (duplicate Commander.js export command collision in openstandardagents), get a green openstandardagents!890 pipeline, and a verified merge.
  • Fix: treat any self-reported capability, completion, or state claim as unverified until confirmed by an independent check (a real command run, an API call, a second party) — especially immediately after a context-compaction or session-resume boundary, where stale constraints are most likely to be carried forward unexamined.
  • Distinguishing test: was the claim ever actually tested in this session, or only asserted/inherited? If untested, it is not measured state.

Absence proven by a single mechanism is not absence

  • Symptom: a NOT_FOUND/"doesn't exist" conclusion is drawn from one probe returning empty — and that empty result reads as more confident evidence than a populated one would have.
  • Root cause: the probe mechanism itself silently failed to reach the real data (a blobless/shallow clone, a wrong CLI subcommand, a truncated default page size) and returned empty rather than an error, so the empty result was accepted as a measurement instead of a non-answer.
  • Instances (2026-09-06):
  • A blobless clone's first probe of a repo returned an empty composer.json, nearly filed as "the BOM metapackage doesn't exist" — caught only by re-verifying with ls-tree across all three refs, which showed the file was real.
  • gc agents (wrong subcommand) returned zero results, nearly filed as "FOUNDRY is not a Gas City agent" — caught only by retrying with the correct gc agent.
  • Instance (2026-09-09): qmd's search index and GitLab's own blob search were both confirmed stale during the same upstream-doc reconciliation pass. qmd was 5 days stale (71% orphaned embedding chunks). GitLab blob search returned a hit for "Health Patrol" text inside an ADR file that, read directly, contained no such text at all -- a false positive, not just a missed true positive, from the same staleness cause. A search-index HIT still needs the live file read to confirm it's current; a search-index MISS never proves absence on its own.
  • Fix: confirm any absence/NOT_FOUND claim with a second, differently-shaped probe before filing it. A single mechanism's silence is not proof of absence.
  • Distinguishing test: could the probe mechanism itself produce an empty result on a real, populated target (wrong subcommand, shallow clone, default pagination, wrong scope)? If yes, the NOT_FOUND is unverified until a second mechanism agrees.
  • Note: a live instance of this same class occurred in this channel during this same session (2026-09-06) — an MR's state: merged plus a green pipeline was accepted as proof of current branch content for mcp_gateway's composer.json, without re-checking the live branch tip. A hotfix commit 26 minutes after the merge had already reverted the change; the true current state only surfaced when a second party re-measured against fresh clones. "Merged" is a historical fact about one commit, not a live-content claim — always re-read the current branch tip before asserting what a repo's content is now.

Closed work as false authority

  • Symptom: a closed bead, MR, or ticket reads as settled and is treated as more reliable than live content — a closed record is almost never re-verified, which makes a wrong closed record more dangerous than an open one. Nobody argues with a closed ticket.
  • Root cause: the record's outcome was true at close time (or was never re-checked at all), but the underlying source was later reverted or changed, and nothing reopened or annotated the closed record to reflect that. This inverts "chat is not the work graph" — the graph is supposed to outrank memory, but a wrong closed record turns it into the least-questioned source of a wrong answer.
  • Instance (2026-09-06): bead bc-1xn (Mac-local store, not Oracle canonical) is closed on the premise that mcp_registry should be renamed to bluefly/mcp_gateway. That rename shipped (mcp_gateway!136), then was reverted 25 minutes later as wrong — composer name must match the actual Drupal module machine name (mcp_registry.info.yml), which was never renamed. The bead's recorded outcome is the opposite of the real end state. Two independent agents tonight nearly re-proposed the identical breaking rename on that bead's authority before either checked live branch content.
  • Fix: a closed bead asserting a source change should cite the exact SHA that proves it, so a later reader can check rather than trust. Reverting a commit named in a closed bead should reopen or annotate that bead. Closed beads asserting source outcomes are worth a proactive audit rather than waiting to trip over more of them.
  • Distinguishing test: does the closed record cite a checkable SHA, or only a description of the intended outcome? If only the latter, treat it as unverified until the live branch tip is read directly.

Fault blocks its own remediation path

  • Symptom: the tooling needed to record, broadcast, or verify a fix is itself degraded by the fault being fixed — the bead note that cannot be written because writes are failing, the gc session list that times out so the "stop retrying" rule cannot be nudged out, the gc mail that fails while trying to coordinate the mail outage.
  • Root cause: remediation depends on the same substrate as the fault (one shared Dolt server behind Beads, mail, hooks, and session listing).
  • Instances (2026-09-20): Oracle Dolt write starvation — see ledger/evidence/incidents/2026-09-20-oracle-dolt-write-starvation.md. Three in one night: bead append, session nudge, mail send.
  • Fix: keep a durable record outside the failing substrate (a file on the host, a cross-session message) and say so explicitly; verify every write by read-back; do not retry into a store whose failures may commit.
  • Distinguishing test: ask "does recording this finding require the thing that is broken?" If yes, write it somewhere else first.

Stale record manufactures work

  • Symptom: several independent sessions escalate, for hours, for an action that was already executed and proven ineffective.
  • Root cause: a retraction did not reach every projection the original claim reached. A shared agent-memory file kept a retracted attribution and the words "containment never applied"; each new session read it as established, added urgency, and never re-tested.
  • Instance (2026-09-20): four sessions between 07:24Z and 11:16Z asked for systemctl stop gascity-dashboard.service, which had been executed 05:43:56Z, measured with no effect, and rolled back 05:56:27Z. The ExecMainStartTimestamp of the rollback was read as proof the stop never happened. Corrected at source with the original preserved.
  • Fix: a retraction is complete only when it has overwritten or annotated every copy — bead, ledger, shared memory, session handoff. Memory entries carry an evidence class and a supersession pointer; unverified memory is not runtime truth. Defect: STALE_AGENT_MEMORY_CAUSES_REPEATED_FALSE_REMEDIATION.
  • Distinguishing test: is the "urgent" action's timestamp evidence consistent with it never having run, or with it having run and been reverted? Check the runtime before escalating.

Config export on a host-mounted checkout silently drops files

  • Symptom: drush config:export into config/sync reports success and drush config:status immediately reports "no differences", yet minutes later some .yml files are gone from the sync directory; git status shows deletions nobody made. Losses vary run to run (198, 4, 12, 0, 0, 0 files across six exports on 2026-09-20; earlier 129, 242 and 18).
  • Root cause (INFERRED, not yet proven on a clean host): the export target is a virtiofs mount of the Mac checkout inside the DDEV web container (OrbStack, performance_mode: none). Drush deletes every file in the target (FileStorage::deleteAll) and rewrites ~1,500 files; a host-side process concurrently rewrites .git/index while no git runs in the container; a 250 ms readdir+open poll during the export listed 622 entries with 99 unopenable (ENOENT), and transient ENOENT persisted ~2 s after drush exited before settling with files permanently missing. The same export into 17 other directories on the same mount lost nothing. Ruled out: DB reads (drush throws on a failed read; every lost object was present in the config table), config_split (4 splits, all inactive), monitoring churn (fixed separately), core DatabaseStorage.
  • Instance: bluefly.io, 2026-09-20, Drupal 11.4.7, drush 13.8; bead bluefly-9yp.3.7.
  • Fix / mitigation: DDEV performance_mode: mutagen so writes land on a native volume; identify the host watcher touching config/sync; if it persists on a clean host, report upstream to OrbStack with this reproduction. No wrapper scripts around drush.
  • Distinguishing test: after drush cex, wait several seconds, then git status must show only the intended changes before committing. drush cst alone is insufficient: it reads the files before they vanish.

Transitional agent creates a second checkout of an authority repo

  • Symptom: two clones of the same authority repository on one workstation ($ESTATE_ROOT/BluCity-Docs and $ESTATE_ROOT/DEMOs/blucity-docs); the second has an HTTPS gitlab.com remote instead of the estate's git@gitlab-bluefly: alias; commits land under the operator's identity from a checkout nobody governs.
  • Root cause: a transitional (non-authority) agent session cloned relative to its own working directory instead of resolving the estate's canonical checkout, and the tool's bootstrap did not name that checkout or the ssh remote.
  • Instance (2026-09-20): Antigravity session cloned blucity-docs at 00:19 -0400, committed and pushed feature/agent-capability-estate (5-line governance edit, later merged). Clone trashed at 15:03Z after proof it was clean and fully pushed.
  • Fix: transitional-agent bootstraps (Antigravity, Gemini, Codex) must point at the canonical checkout and use the ssh remote; a clone of an authority repo anywhere else on the workstation is a defect, not a convenience. Preserve unique commits (push or patch) before trashing.
  • Distinguishing test: find ~/Sites -maxdepth 3 -type d -name .git -path '*blucity-docs*' returns more than one path, or git -C <clone> remote get-url origin starts with https://.

Idempotency check does not describe the rule it guards

  • Symptom: a "check-then-insert" guard (iptables -C … || iptables -I …, grep -q … || echo … >>, kubectl get … || kubectl create …) runs on every unit start, timer tick, or reconcile loop and the guarded resource grows monotonically: thousands of identical firewall rules, repeated config lines, duplicate CRDs. Cost shows up elsewhere first — iptables-save takes seconds, every caller that walks the chain (k3s, kube-router, docker) burns CPU, load rises with no single hot process.
  • Root cause: the check's spec is not byte-identical to the insert's spec. iptables -C matches the whole rule, so a check that omits -m comment --comment <tag> never matches a rule inserted with one; the guard is a no-op and the loop is unbounded. The defect is invisible at one execution and only compounds under restart or reconcile churn.
  • Instance (2026-09-20): gastown-gateway.service ExecStartPre on Oracle — NRestarts 43,301 × 2 CIDRs = 84,507 duplicate INPUT rules, INPUT chain 84,930 rules, iptables-save 9.6 s at 64 % CPU, ≈1.5 of 4 cores spent in k3s iptables walks. The repo template (agent-docker deployments/gastown/systemd/) carried the same guard. Same class, separate instance: 417 duplicate kube-router netpol rules.
  • Fix: the check must describe exactly the thing the insert creates — same match modules, same comment, same target. Where the tool offers an atomic idempotent form (iptables-restore of a managed chain, kubectl apply, ensure-style APIs), prefer it over check-then-insert. Never pair an unbounded Restart= with a non-idempotent ExecStartPre. Source fix: agent-docker MR !399; host flush is a separate authorized step.
  • Distinguishing test: iptables -S | sort | uniq -d | wc -l is non-zero, or NRestarts × inserts-per-start equals the rule count. More generally: run the guard twice and count the resource — any growth on the second run is this signature.

discover.duadp.org 502 Bad Gateway

  • Symptom: discover.duadp.org returns 502 Bad Gateway via Cloudflare.
  • Root cause: The origin daemon on port 4201 is not running (localhost:4201 is down).
  • Fix: Start or restart the discover daemon on the origin host on port 4201.
  • Distinguishing test: Confirm Cloudflare tunnel is healthy but curl -I localhost:4201 on the origin fails.

Qdrant Multi-Instance Violation

  • Symptom: Qdrant data is missing or out of sync between agents, or lock file conflicts occur.
  • Root cause: Multiple Qdrant instances are running instead of a single authoritative agent-brain instance. Qdrant ownership belongs to agent-brain, NOT agent-studio.
  • Fix: Terminate rogue Qdrant instances and route all vector operations through the canonical agent-brain service.
  • Distinguishing test: Check for multiple qdrant processes or containers on the host.

World-Writable .env File (P0 Security)

  • Symptom: .env files have 777 or 666 permissions.
  • Root cause: Misconfigured deployment script or manual permission change.
  • Fix: chmod 600 or 640 on the .env file.
  • Distinguishing test: ls -l .env shows world-writable permissions.