Skip to content

Evidence Contract

Purpose

Every engineering report SHALL separate what was directly observed from what is inferred, decided, or unresolved. A report that blurs these is not evidence — it is narrative.

The six sections

Every non-trivial finding, receipt, or investigation report uses exactly these:

  • OBSERVED — the artifacts actually inspected, anchored to the artifact examined, never to implied history. ("The version examined before commit contained X" — not "X previously contained…", which implies a history no one looked at.)
  • REPORTED — a claim supplied by another agent, person, pipeline, artifact, or prior receipt, but not independently inspected by the agent making the current report. A reported claim MUST preserve attribution (who/what reported it) and MUST NOT be silently promoted to OBSERVED. Confirmed 2026-08-24 (ubuntu-o4fl): a session relayed "another session is authoring this doctrine" as settled fact without independently checking; the receiving session was corrected by the third party and verified the correction directly — the assignment had been refused, not accepted. The failure was REPORTED treated as OBSERVED, one hop of attribution lost in a relay. Attribution chains longer than one hop compound this risk and should be re-verified at the point of use, not carried forward from memory.
  • INFERRED — conclusions bounded strictly by the observed artifacts. An inference MUST NOT drift into a recommendation, and MUST NOT generalize beyond the operations actually exercised. ("The exercised write operation failed" — not "the write subsystem is broken.")
  • DECISION — the resulting engineering action, stated as an action, not restated evidence.
  • NOT VERIFIED — distinguishes lack of evidence from lack of access. State which. NOT VERIFIED ≡ NOT_ESTABLISHED ≡ UNKNOWN (rule 3, below) ≡ UNVERIFIED (rule 6, below) — one semantic label with four names in use across this contract and receipts, not four distinct states. All four mean: evidence missing, contradictory, stale, incomplete, or drawn from a scope that cannot support the requested conclusion. This is a correct terminal state, not a placeholder for later resolution. Use whichever name the surrounding context already uses; do not treat picking a different one of the four as a different claim.
  • BLOCKING — stated only at the contract boundary (what interface is unavailable), never drifting into speculation about why it's unavailable.

Measurement boundary and verification boundary

Every promoted conclusion SHALL declare: - Measurement boundary — what population was actually examined (which repos, which files, which time window). - Verification boundary — what remains unknown even after the measurement.

Never report a reduction, a fix, or a completion until both sides of a claimed delta have been measured within a declared boundary.

Core rules

  1. A declaration is not an execution. A CI/config definition states intended behavior; a log or API response states what actually happened. Reading a definition and reporting it as an event is a category error.
  2. A missing observation is not an observation of absence. Absence from one queryable system is not evidence of absence in a different, unqueried system.
  3. Truncation is not evidence of absence. A head, a grep, a partial search, a crashed sweep, or a timed-out query may only produce UNKNOWN. Absence may be asserted only after completeness is verified (a full listing, a full ref enumeration, a complete inventory).
  4. An unexpected metric is a hypothesis, not a finding. Re-derive it from raw data before acting on it; the gate belongs inside the command that acts, not in the reasoning that produced the number.
  5. A failed diagnostic is evidence about the diagnostic, not automatically about the subsystem it measures. Reconcile a failed check against independent evidence before concluding the thing being checked is actually broken.
  6. Conversation is not authority. Reports are not authority. Only directly observed evidence or authoritative systems (Git, GitLab, Beads, a live filesystem/API check) are authority. If a claimed finding cannot be reproduced from the visible record or an authoritative system, it is UNVERIFIED until reproduced — never adopted, never refined, never treated as continuity.
  7. Evidence unavailable (real, but out of current scope): rerun the observation.
  8. Evidence absent (no record anywhere checkable): withdraw the claim.
  9. Don't let a true narrative absorb one unsupported statement. A closeout table, receipt, or summary is audited row by row.
  10. Neutral framing on "not established." State that a finding was not established by this investigation — never imply the underlying condition doesn't exist.
  11. An empty result requires a proof chain before it means zero. A query, scan, or search returning no rows is not itself evidence of absence. Before reporting OBSERVED=0, establish in order: the reader actually ran (not silently skipped or short-circuited), it targeted the correct target (the intended ref, store, project, or population — not a default, cached, or stale one), it held sufficient authority/permissions to see everything in scope (an access-denied result and a genuinely-empty result are indistinguishable without this check), the scope queried was the scope claimed (no silent filter, pagination truncation, or narrower population than stated) — only then may the result be reported as OBSERVED=0. Skipping any link in this chain converts a plausible absence into an asserted one; report NOT VERIFIED (see equivalent names, above) instead until the chain is complete — this includes the partial-search and pagination-truncation cases in rule 3, which this rule does not override or narrow.

See also: standards/architecture/capability-convergence.md (ownership decisions, upstream integration gate), standards/Decision-Vocabulary.md (terminal-state definitions).