Investigation Methodology¶
Purpose¶
Defines the operating sequence for any investigation or engineering change, so that discovery, verification, and mutation never collapse into a single unreviewed step.
The gate order¶
Bead → Authority Check → canonical worktree → minimal edit → repository verification
→ commit → push → verify SHA → merge request to release/v0.1.x → CI pass
→ merge → verify containment → destroy worktree → reconcile bead
Authority Check runs before the worktree, not after the edit. Before touching anything:
what capability is changing, what is the smallest independently ownable unit of it, who is the
current authority, what evidence supports that, does an authoritative implementation already
exist, and is this a build, a validation, a convergence, or an amendment (see
capability-convergence.md).
Discovery Check¶
Discovery changes work. New evidence immediately invalidates an obsolete plan. If evidence proves a thing already exists, the work item changes from building it to validating, converging, or amending it. Construction never remains the default once existence is discovered. Work superseded this way is closed as superseded by discovery — not as wrong.
Worked example: standards/architecture/inference-topology.md began as a config-fix task
("point LiteLLM somewhere that works") and, once the owning repository was identified, became
a decision-matrix document instead — because discovery proved there was no replacement target
to converge toward, only a runtime state to converge away from. The plan changed; nothing was
built prematurely to fill the gap.
Phased execution (for changes with real data at stake)¶
- DISCOVER — gather evidence only. Full status, full branch/remote/worktree listing. Untracked files are listed, never ignored.
- REVIEW — inspect every change. Classify (Intentional / Generated artifact / Build output / Temporary / Unknown). Default is Intentional unless there is specific evidence otherwise.
- VERIFY — confirm: correct branch, upstream exists, not detached, no merge/rebase in progress. Any failure here is a hard stop, reported with the exact command and output.
- COMMIT — stage everything classified intentional. One real commit, real message.
- PUSH — to tracked upstream only. Never force. Never rewrite history. If it fails, stop and report the exact error.
- VERIFY (post) — prove the push landed: local HEAD equals the remote ref. A push is not a delivery until this is confirmed.
Escalation before repair¶
On any failure (corrupted repo, missing infrastructure, broken tool): observe → classify (verified / not established) → decide the smallest safe next step.
- Read-only inspection is always available, even while a system is quarantined.
- Repair, deletion, force operations, and history rewrites require explicit authorization, scoped to exactly what was authorized.
- A repository's own governance gate (deletion guard, branch-name rule, merge protection) is respected, not routed around via override flags or restructured changes.
- Verify one instance before generalizing a fix across many similar ones.
Blocked-pivot rule¶
A genuine external blocker (missing credentials, no write access, interactive-auth requirement) does not produce idling. Report the exact blocker precisely, then pivot to the highest-value adjacent work that doesn't depend on the same blocker. Never retry an identical failing call in a loop; never invent a workaround that bypasses the blocker's purpose.
Hypothesis discrimination¶
When more than one explanation fits the evidence, do not pick the likeliest — identify the observation that would tell them apart, then go and make it.
When competing hypotheses predict different observations, prioritize evidence capable of distinguishing them over additional evidence merely consistent with the current explanation. If no discriminating observation exists, record that the available evidence cannot distinguish the hypotheses. That is a legitimate terminal state — an unresolved discrimination recorded honestly is worth more than a confident selection between explanations the evidence cannot separate.
The sequence:
- State the competing explanations explicitly. Two explanations left implicit will be collapsed into one, and it will be the convenient one.
- Ask what each predicts differently. If two explanations predict the same observations everywhere, no amount of further looking will separate them — say so and stop.
- Verify the chosen observation is actually discriminating. Distinct from step 2. Step 2 asks whether the hypotheses differ anywhere; this asks whether this specific test is one of the places they differ. A test may be selected from a real set of differences and still not be one of them:
H1 predicts X.
H2 predicts X.
-> This observation cannot distinguish them, however expensive it is to make.
Skipping this gate produces data that raises confidence without reducing uncertainty — the most expensive kind of nothing. 4. Make the observation. Prefer the cheapest test that discriminates over the most thorough test that does not. 5. Revise, and record the retraction. A hypothesis eliminated by evidence is a result. It is recorded, not deleted — see execution-receipt-specification.md § Hypotheses.
The rule matters most when the explanations imply opposite remediations. Worked example, 2026-07-27:
Four requested catalog fields —
owner,status,consumers,lifecycle— were absent from everypack.tomlinblucity-packs. Two explanations fit equally well:H1 — Bluefly is not populating fields the schema supports. Remediation: populate the packs. H2 — the upstream schema does not define these fields. Remediation: find another authority, or extend the spec upstream.
Both predicted the same repository state, so no further inspection of
blucity-packscould separate them. The discriminating observation lay outside the repository entirely: the upstream Pack spec's field list. It definesname,schema,version,requires_gc,description,requires— and no ownership or lifecycle field.H2. Under H1 the fix was to add
owner = "..."; under H2 that same edit produces a non-conformant manifest. The obvious remediation was the wrong one, and only a test capable of falsifying H1 revealed it.
Two failure modes this guards against:
- Confirmation drift — accumulating evidence consistent with the working explanation while never testing one that would refute it.
- Convenience collapse — silently adopting whichever explanation implies work already within reach, because the discriminating observation lies outside the current context.
An investigation that never attempted to falsify its own explanation has not established it, however much evidence it collected.
Measurement normalization¶
Instrumentation must enumerate every supported representation before producing a count. A measurement that recognises one encoding of a thing will silently report zero for every other encoding, and a confident zero is worse than a missing value — it closes the question.
Observed failures of this rule, all from a single investigation on 2026-07-27:
| Measurement | Assumed one representation | Reality | Wrong conclusion produced |
|---|---|---|---|
.beads presence |
~/.beads and ./.beads are distinct paths |
cwd was $HOME; both resolved to the same directory |
A pre-existing directory reported as newly created |
allow_failure: true in CI |
every match is a job | one match sat under an include block |
3 non-blocking jobs reported; the true count was 2 |
| Formula step count | steps are always [[steps]] array-of-tables |
a second syntax exists — [steps.<name>] tables with on_complete |
Nine 5–7 step formulas reported as single-step |
The class is identical in each: the instrument assumed one valid representation while the system supported several. None was a reasoning error; each was a counting error that then produced confident, wrong reasoning downstream.
Before a count becomes a finding:
- Enumerate representations. Does this concept have more than one legal encoding — a second syntax, an alias, a symlink, a nested form? Grep for the concept, not for one spelling of it.
- Normalize, then count. Resolve paths before comparing them. Parse structure rather than matching a line prefix, where a parser is available.
- Sanity-check zero and total. A zero, or a total equal to the whole population, is more often an instrument defect than a discovery. Verify one positive case by hand.
- State the instrument in the receipt. Record what was matched, so a reader can see what the measurement would have missed.
A count is an inference about the world, not an observation of it. It inherits the confidence of the instrument that produced it — see evidence-contract.md.
Discovering the second representation is often the finding. Two of the three errors above surfaced defects — a duplicate-name collision set and a dual formula syntax — that a correct count would have rendered invisible.
Mechanisms beyond representation¶
Representation is one mechanism. The rest of this section catalogues the others, observed during a single investigation on 2026-09-06 to 09-12 by every agent working it. None involved a second encoding of anything. Add rows as they are found — the prose below deliberately depends on no count.
| Mechanism | Instrument | What went wrong silently | Reported | True |
|---|---|---|---|---|
| Scope truncation | grep -rlI <token> |
every .git/config in the tree |
4 credential locations | 8 |
GitLab search API, scope=blobs |
everything past the default per_page=20 |
20 hits | 27 files / 31 lines | |
groups/<g>/projects |
subgroups — no include_subgroups=true |
26 projects | subgroup-inclusive count never established | |
git worktree list in one repo |
worktrees belonging to other repos, registered to their own gitdirs | 17 unregistered | 104 worktrees, 103 live, 1 broken, all registered | |
find killed at the 120 s tool timeout |
the unwalked remainder | 370 files | 2001 | |
| Stale referent | git grep -l <s> HEAD |
a checkout 140 commits behind origin |
string absent | present (1 and 27) |
| same, across 7 rigs | checkouts 0–140 behind, 2 with no upstream | identity hits 6 / 3 / 1 | 17 / 28 / 6 | |
| reading a doc from a local branch | 92 lines of newer content | doctrine gap | already covered on origin |
|
| Basis mismatch | date -u printed beside stat local time |
that the two use different clocks | mtime 1.5 h in the future | 18.5 h in the past |
| Match width | ps aux \| grep 'dolt\|gc\|beads' |
nothing — it over-matched agent prompt text carried in command lines | 143 KB of noise | 4 processes (pgrep -x) |
strings <bin> \| grep -cx <term> |
every line where the term is not alone on it | 0 for every term, controls included | doctor 1203, issue-prefix 5 |
|
| Non-persistence | unset -f glab; type glab in one call |
that tool-driven shells do not keep state between calls | fixed | wrapped again on the next call |
| Query corruption | git show "$REF:templates" in zsh |
the :t tail modifier rewrote the ref before the tool saw it — origin/release/v0.1.x:templates became v0.1.xemplates |
two findings "nothing found" | both present; ${REF}: is the fix |
Scope truncation drops part of the domain. A stale referent measures a copy instead of the authority. A basis mismatch compares two instruments calibrated differently. Match width misses by being too broad or too narrow. Non-persistence verifies under conditions that do not survive the check. Query corruption alters the question before the instrument ever receives it, so the tool answers correctly — a different question. Each is orthogonal to representation, and to the others.
The single shared property. Every row above returned a confident value rather than an error. None raised, timed out, or reported that it could not measure. The instrument had no way to say I could not answer that, so it answered something else.
Read the Reported column and note how few are zeros. Most are undercounts, and those are the more
dangerous half. A zero at least invites the question — nobody expects a credential scan or an
estate sweep to legitimately find nothing. 4 where the answer is 8, or 20 where it is 27,
reads as a result and gets acted on.
This paragraph deliberately carries no tally. An earlier revision hardcoded one — "six returned a zero" against a ten-row table — and it was wrong on the day it was written and would have been wrong again the first time a row was added. A count stated in prose beside a table it does not own is a claim with no owner and no way to stay true. State the property; let the table carry the arithmetic.
Two consequences follow, and the second is the dangerous one:
- A false zero closes a question that should have stayed open.
- A false zero can also manufacture a refutation.
strings <gc> | grep -cxreturning 0 for every term would have been published as the binary contains none of these strings — active evidence against a correct hypothesis. The control that caught it was running the same query fordoctorandbeads, words that must be present.
When the tool itself cannot report failure¶
The same disease appears below the operator. Observed in bd on 2026-09-12:
bd show bc-2s5 -> ✓ bc-wisp-2s5aw · order:nudge-on-route [P2 CLOSED] exit 0
bd show bc-zzzzz -> Error: no issue found matching "bc-zzzzz" exit 1
This is a lie about identity: a near-match is preferred over an exact-match failure, and the substitution is reported with a success marker and no warning that the returned id differs from the one requested. The error path itself is correct, as the second line shows — the defect is specifically that a substring match outranks an honest miss.
Treat an id echoed back by a tool as a value to verify, not as confirmation of what you asked for.
The operator-side mirror of the same defect is a classifier keyed on one named failure string.
Testing bd show <id> for the substring no issue found and treating its absence as success
classifies every unanticipated failure as a pass: with the backend down, all twelve candidates
returned failed to open database: dolt circuit breaker is open, so all twelve were recorded as
real issues — including blu-book, iac-promo and bc-vendor, which are project names.
if not <known error>: it worked inverts the burden of proof. Test for the success condition, not
for the absence of one remembered failure.
A cautionary note on how this section nearly grew a second, false entry. The first draft also
claimed bd exits 0 on backend failure. It does not. That reading came from
bd list 2>&1 | head -5; echo $?, which reports head's exit status, not bd's. Measured
with clean redirection, bd list exits 1 and writes to stderr; with set -o pipefail, the
pipeline exits 1 as well. The tool was correct and the instrument was not — while cataloguing
instrument failures, in the section about instrument failures.
$? after a pipeline belongs to the last command in it. Redirect, or set pipefail, before
attributing an exit status to the program you care about.
Claim-method alignment¶
State what the method actually proves before reporting what it seems to prove. A method that answers a narrow question is routinely reported as having answered the broader claim it was run to support. The two are not the same sentence, and the gap between them is where wrong conclusions enter a receipt looking exactly like right ones.
Before reporting any conclusion, write out three lines:
CLAIM=<what is about to be reported>
METHOD=<the exact command or check that was run>
WHAT_METHOD_ACTUALLY_PROVES=<the narrowest true statement the method supports>
If CLAIM is broader than WHAT_METHOD_ACTUALLY_PROVES, do not report CLAIM. Either narrow it to what was actually shown, or run the additional check that closes the gap.
Negative claims need a second independent proof; positive findings generally do not. A positive result -- a file exists, a pattern matched, a value equals X -- leaves an artifact that can be inspected and, if wrong, contradicted. A negative result -- does not exist, no remote, no unique work, no secrets found, not reachable -- leaves nothing to inspect. A broken method and a true absence return the identical empty result, at the identical confidence. That asymmetry is the reason negatives carry a higher evidentiary bar, not caution for its own sake.
Observed failures of this rule, across a single verification session on 2026-09-02:
| Claim reported | Method run | What the method actually proved | Correction |
|---|---|---|---|
| Branch is local-only; work is at risk | git branch -r, git rev-parse --verify refs/remotes/origin/branch | What the last fetch wrote into this clone's remote-tracking cache | Single-branch clone (remote.origin.fetch scoped to one ref) meant the cache could never hold that ref regardless of fetch count. git ls-remote --heads origin branch queries the server directly; the branch existed, local HEAD was a strict ancestor. Three of four "local-only" verdicts inverted. |
| Branch has 3 unique commits, will be lost | git log branch --not base | Which commits are unreachable from base by ancestry alone | The identical change had been authored twice (same commit message, two SHAs). Two-dot git diff base branch was empty -- zero unique content. Ancestry and content answer different questions; duplicate authorship makes them disagree. |
| 0 literal credential assignments in scope | grep for secret patterns piped through exclusion filters | That zero lines survived both the pattern and the exclusion filter | Exclusion terms matched against grep's file:line:content output -- including the path segment, not just the matched content. Any file with an excluded substring anywhere in its path was silently dropped. A control run with exclusions off found 469 raw matches, ~370 real candidates after correct content-only filtering. |
| A path is protected by a path-scoped access policy | Observed a tool denial on a compound command targeting the path | That one specific command string was denied | The compound happened to end in a subcommand caught by a broader, unrelated guard (git stash *, which cannot distinguish read-only stash list from a mutation). No rule anywhere referenced that path; the directory did not exist at all. A denial proves the command was denied -- it never proves why until the matching rule is read. |
| Two doc lines need semantic secret-reference remediation | grep -c for the pattern, without printing the matched lines (correct secret-hygiene move) | That the pattern's text is present on those lines | Both lines were prose prohibiting the pattern, not instances of it. Presence of text proves presence of text; semantic classification requires reading context, which the correct hygiene choice (not printing secrets) had foreclosed -- the fix is to route that read to a channel that can see it safely, not to classify blind. |
GitLab's private-resource 404 is the sharpest member of this family -- a 401 or 403 at least looks like an authentication failure; a 404 on a resource the caller cannot see is indistinguishable from the resource not existing, by design. Never read a private-resource 404 as non-existence without independently establishing AUTHENTICATED=YES first -- see the GitLab Private-Resource 404 Rule in authentication-secrets-constitution.md section 5, which this section cross-references rather than duplicates.
A seventh case is a different failure shape: the method was right, and the answer was true when measured, then carried forward past the moment it stopped being true. Time-sensitive evidence expires. A branch equality, remote tip, clean-merge state, authorization result, or configuration observation proves the state at the time it was measured -- not at the time the dependent action actually executes. This is invisible in review: the evidence is genuinely correct and the reasoning is valid, so only the elapsed time between measurement and action makes the conclusion false. It is the one case in this set that two agents hit independently within the same hour, on two unrelated repositories -- one of them while actively writing this section. That recurrence is why the rule belongs here as a mechanical field, not a caution to keep in mind.
| Claim reported | Method run | What the method actually proved | Correction |
|---|---|---|---|
| Retargeting a branch to release/v0.1.x is a safe merge vehicle | main and release/v0.1.x were equal at the moment the branch was cut | State at branch-cut time only | Checked directly at retarget time: conflicts=true, pipeline failed, 376 files / ~239K lines of drift -- release/v0.1.x had moved on in the interim. Historical equality said nothing about current-state safety. |
| A rebase from a diagnosed remote SHA rewrites no published commits and needs no force-push | git ls-remote at diagnosis time | The remote tip at that instant | The remote can advance between diagnosis and execution; the same rebase can silently become a force-push over someone else's published work if it does. Flagged in review before being committed. |
Board notation that makes the expiry visible instead of implied:
REMOTE_TIP_AT_DIAGNOSIS=<sha> MEASURED_AT=<utc>
REMOTE_TIP_AT_ACTION_TIME=MUST_REVERIFY
Any consequential action that depends on a time-sensitive observation -- branch equality, remote tip, clean-merge state, an authorization result, a configuration read -- re-verifies it immediately before execution. The gap between diagnosis and action is exactly where the observation can go stale.
Mandatory pairs, each closing one gap above:
- Remote branch/ref existence -> git ls-remote --heads origin ref. Never git branch -r, git rev-parse --verify refs/remotes/..., or git remote show -- all three report the local cache, which a single-branch clone or a stale fetch can make lie.
- Unique-work-on-a-branch claims -> ancestry (git log/rev-list branch --not base) and content (git diff base branch), never either alone.
- A scanner or filter's zero-result -> a positive control proving the scanner fires on a known match before the zero is trusted. Report ZERO_RESULT= and CONTROL_MATCH=; a failed control makes the result ZERO_RESULT_INVALID, not zero.
- A tool/command denial -> the matching entry in the actual deny configuration, read directly. A denial is evidence the command matched a rule, not evidence of which rule or why.
- Presence of a sensitive-looking string -> does not by itself establish its semantic class (live secret vs. prohibition-text vs. stale example). Classify from context, or route the classification to a channel that can safely read it.
- A time-sensitive observation (branch equality, remote tip, merge cleanliness, auth result, config state) -> re-verify immediately before the dependent action executes, not at whatever point it was originally diagnosed.
- An API-derived count -> exhaust pagination, or read the total from a response header, before it is trusted. A bare per_page default (commonly 20) silently truncates a larger true result set with no visible marker. Real instance: a "20" result was GitLab's default page size; the true figure was 27 files / 31 lines.
- Content measurement against a local checkout -> an explicit
origin/<branch>, not the local clone as-is.git grep HEAD(or equivalent) against a clone that may be stale produces confident, false results with no error. If a local clone must be used, first prove currency withgit rev-list --count local..origin. Real instance: a 3x identity-count undercount and a false zero-hit claim, both against clones 56-140 commits stale. - Before a hygiene/identity edit, a duplicate-risk check -> the current content of the target ref, not only a search for open MRs on the topic. A fix that already landed directly (no MR, or an MR already merged) is invisible to an open-MR-only search. Real instance: two sessions independently fixed the identical line in the same file the same night -- one via MR, one already merged upstream, neither aware of the other.
For preservation decisions specifically, an unproven negative defaults to preserve. Being wrong about "this is unique" costs disk space. Being wrong about "this is already duplicated elsewhere" costs the work itself. The two errors are not symmetric, and the default should not be either.
Closeout discipline¶
An investigation closes when repository/system integrity is verified, every open question is either answered or explicitly recorded as unresolved, no cleanup was silently performed, and the remaining queue is separated from the investigation record as ordinary next work.