Skip to content

Evidence Reporting Standard

Parent contract: Bluefly Engineering Execution Contract Version: 1.0 Applies to: All agents operating in the Evidence role.


1. Confidence Classification

Every claim in an evidence report carries exactly one confidence level.

Level Label Meaning Source Required
1 VERIFIED_RUNTIME Direct runtime query or observation confirms the claim Runtime output, query result, health-check response
2 VERIFIED_SOURCE Source code or configuration artifact in the inspected workspace confirms the claim File path, line number, content excerpt
3 DOCUMENTED Maintained upstream documentation or governance documents assert the claim Document URL or file path
4 INFERRED Logical conclusion from verified/documented evidence; not directly observed Explicit reasoning chain from cited evidence
5 UNPROVEN Plausible but not demonstrated in any inspected artifact Statement of what was searched and not found

Rules:

  • Never promote a finding to a higher confidence level without additional evidence.
  • INFERRED findings must cite the specific VERIFIED or DOCUMENTED evidence they derive from.
  • UNPROVEN is always preferred over an unsupported INFERRED claim.

2. Three Evidence States

These states are distinct. Never conflate them.

State Definition Example
Configuration exists An artifact declares intent — a TOML key, a YAML manifest, a systemd unit file, a schema reference before_sling = "mol-bluefly-context-preflight" exists in pack.toml
Runtime exercised The configuration has been loaded and executed at least once The orchestrator has started and evaluated at least one order cycle
Runtime verified Execution has been observed, validated, and its output confirmed correct A bead was created, transitioned through its lifecycle, and produced the expected output

Rules:

  • A configuration that exists does NOT imply it has been exercised.
  • A configuration that has been exercised does NOT imply its output is correct.
  • Evidence reports must state which of the three states has been established for each finding.
  • When only state (1) is established, use language like "declares," "is configured to," or "provides a configuration point that could be used for." Never use "uses," "runs," or "executes."

2.1 Capability-Claim Ladder (Hypothesis → Observed → Verified → Proven)

For claims about a capability (an upstream feature, a replacement, a migration outcome), the three states extend into a four-level ladder used by ADRs, migration ledgers, and receipts:

Level Meaning Maps to
Hypothesis Design assumption; not yet seen anywhere real pre-state (DECLARED / POSSIBLE)
Observed Seen in a real runtime at least once by direct inspection states 1–2 (exists / exercised, witnessed)
Verified Deliberately exercised end-to-end with the outcome checked state 3 (runtime verified)
Proven Multiple successful production executions over time longitudinal extension of state 3

Rules:

  • Observed ≠ Verified. "I saw it" and "I proved it" are different statements; never promote on sight.
  • A level only moves forward on new evidence, recorded where the claim lives.
  • Retirement and replacement decisions gate on Verified, not Observed (see Convergence Doctrine §7).

3. Investigation Boundary

Every evidence report must include an Investigation Boundary section near the top. This section prevents readers from misusing the report.

Required contents:

  • What the report establishes: Which repositories, directories, and systems were inspected. What types of evidence were collected.
  • What the report does NOT establish: Explicit list of claims the report cannot support. At minimum: current runtime health, deployment state, successful execution, operational correctness.
  • Workspace absence ≠ system absence: When the report states "not found," it means the artifact was not located in the inspected scope. It does not mean the artifact does not exist elsewhere.

Template:

## Investigation Boundary

This report establishes:
- [what was inspected]
- [what types of evidence were collected]

This report does NOT establish:
- [explicit exclusions]

Workspace absence ≠ system absence. "Not found in inspected workspace"
means the artifact was not located in the repositories and directories
surveyed. It does not assert absence from [other locations].

4. Negative Evidence Qualification

When an artifact is searched for and not found:

  • Always qualify the scope: "Not found in the inspected workspace" — never "does not exist."
  • State what was searched: List the directories, repositories, or systems surveyed.
  • Assign confidence: Negative evidence is VERIFIED_SOURCE (the search was conducted and returned no results within scope).
  • Note plausible alternative locations: If the artifact could reasonably exist elsewhere (e.g., Oracle runtime filesystem, private repositories), state this.

5. Report Structure

Evidence reports should follow this structure:

  1. Document header (type, scope, date, status, source artifacts, constraint)
  2. Confidence Classification reference
  3. Investigation Boundary
  4. Findings (organized by domain)
  5. Negative Evidence Summary (qualified)
  6. Source Artifact Inventory
  7. Conclusions (using evidence-appropriate language)

6. Language Rules

Instead of Write
"The system uses X" "The configuration declares X" (state 1)
"X runs on Oracle" "X is declared as a systemd unit on Oracle" (state 1)
"X authorizes Y" "X provides a configuration point that could be used for authorization" (state 1)
"X does not exist" "X was not found in the inspected workspace"
"Owner: BLUEFLY" "Source Authority: BLUEFLY"
"Therefore X happens" "If exercised, this configuration would result in X" (state 1 → INFERRED)

7. Observations vs conclusions

Evidence reports and audits must separate facts from interpretation.

Output Required tag Rule
Command output, counts, file presence, runtime response OBSERVED No interpretation in the same sentence
Quoted upstream doc or committed artifact at SHA RETRIEVED Cite path or URL
Root cause, sprawl diagnosis, “should delete” INFERRED Cite OBSERVED / RETRIEVED inputs in the same paragraph
Searched in scope, not found NOT_FOUND State scope
Not yet checked UNKNOWN Preferred over guessing

Examples:

  • OBSERVED: iac cloud-init installs gascity, gascity, and beads binaries.
  • OBSERVED: Workspace doctrine places Oracle SoR on ~/gt + gt/bd.
  • INFERRED: Duplicate execution models may cause agents to work in two places — if both paths are active on the same host without a declared migration boundary.

Removal hypotheses belong in Infra Convergence Roadmap as candidate for removal with preconditions — never as bare “delete X” and never as scheduled implementation work until approved for removal with a receipt.

Investigation termination: Engineering Authority Exhaustion (NOT_FOUND).

8. Evidence supersession

Evidence outranks previous conclusions. When new evidence invalidates an earlier conclusion, the conclusion must be updated — not defended.

  • A conclusion is only as durable as the evidence that produced it. New OBSERVED / RETRIEVED facts that contradict a prior INFERRED conclusion supersede it.
  • When a conclusion is revised, record the correction in the same ledger (Beads) that holds the original — as an explicit retraction, not a silent overwrite. State what the evidence now supports and what it does not prove.
  • Do not preserve a conclusion because it was previously reported, because effort was spent on it, or because it is convenient. Consistency with evidence outranks consistency with earlier statements.
  • Distinguish contributing factor from proven cause. A plausible mechanism that fits the evidence but is not established by it is INFERRED and must be phrased as "likely contributed to," never "traces to" or "caused by."

Example (from hq-aqn/hq-cwp, 2026-07-12): the claim "every token leak traces to broken 1Password auth" was retracted — the evidence supports "plaintext secrets exist in IaC" (OBSERVED) and "inconsistent 1Password integration likely contributed to credential sprawl" (INFERRED), but does not prove causation.


9. Future Commitments vs Verified Facts

A verified fact is something observed or reproduced. A future commitment is a stated intention. They are not interchangeable.

Conflating the two inflates confidence and undermines the evidence standard. A sincere declaration is not evidence that the declared behavior will occur.

Status Vocabulary

Status Meaning
VERIFIED Supported by observable, reproducible evidence right now.
DECLARED Stated intention or future commitment. Evidence only at the moment it is exercised.
NOT FOUND The search did not locate the item within the stated scope. Does not prove absence.
POSSIBLE Plausible but unconfirmed. Requires further investigation.
UNKNOWN Insufficient information to form a conclusion.
FALSE Contradicted by evidence.

Correct Receipt Format

Separate every claim by its status. Do not group VERIFIED and DECLARED claims under a single STATUS line.

CLAIM:
  Documentation Curation Policy has been reviewed.
STATUS:
  VERIFIED
EVIDENCE:
  - Canonical root identified: blueflyio/blu/blucity-docs
  - Canonical hierarchy acknowledged: Engineering-Standard, Products, Playbooks, Evidence, Research
CONFIDENCE:
  HIGH

CLAIM:
  Future documentation will be authored according to this policy.
STATUS:
  DECLARED
EVIDENCE:
  Agent has stated intent. Compliance is observable only in future actions.
CONFIDENCE:
  N/A — commitment, not evidence

Rules

  1. A DECLARED claim produces no evidence until the declared action is executed and receipted.
  2. Do not promote DECLARED to VERIFIED because the declaration was made sincerely or repeatedly.
  3. When auditing compliance with a declared commitment, the audit produces a new VERIFIED claim — it does not retroactively verify the original declaration.
  4. A declaration made by an agent about its own future behavior has lower epistemic weight than a structural enforcement mechanism (a hook, a gate, a CI check). Prefer structural enforcement over declarations wherever possible.

10. Narrow Claims to Exact Evidence

State only the claim the evidence directly supports — do not round up. A successful SSH connection proves a host is reachable and accepting authenticated connections; it does not prove the services running on it are healthy. Keep those as separate lines: what's reachable/authenticated vs. what's actually verified healthy (NOT ESTABLISHED if unchecked).

Before writing "confirmed" / "healthy" / "no path exists" / any totalizing claim in a receipt, ask: did I actually check that, or am I extrapolating from a narrower observation? If extrapolating, split it into an explicit NOT ESTABLISHED line instead.

Treat execution context as a first-class field in a receipt, separate from the target being verified — e.g. "Execution Context: background job sandbox (isolated credentials)" vs. "Operator Context: interactive workstation (authenticated)". This makes it immediately clear why the same host can produce different results for two different callers, without implying the host itself is unhealthy or that something is wrong. When a background job and an operator's interactive session probe the same target with different results, label which context produced which result rather than treating one as contradicting or fixing the other.

11. Continuity Preflight

When a message describes work, findings, or a prior response as if it happened in this conversation, run this check before responding as if it did:

  1. What system/session am I actually running in right now?
  2. What do I actually have evidence of in this conversation's own history?
  3. Am I being asked to continue something, or just told about something?
  4. Can I actually see the context being referenced, or only being told about it?

If the referenced work isn't in this session's own visible history, say so plainly — do not manufacture continuity, do not validate or build on work you can't verify you did. With multiple parallel execution contexts (interactive sessions, background jobs, different machines), it is expected and correct for one context to have no visibility into another's work — that is not a failure to paper over, it is a fact to state.


12. Loss of Observation Channel (BLOCKED_EXTERNAL)

When a live investigation loses its observation channel mid-task — SSH fails, an API stops responding, ICMP fails — transition immediately to BLOCKED_EXTERNAL and stop reasoning. Do not theorize, rank "root cause candidates," or use hedge words (likely/probably/leading candidate) unless the operator explicitly asks for hypotheses.

Emit exactly this receipt, then end the turn:

EXECUTION RECEIPT
State: BLOCKED_EXTERNAL

Verified
- <facts personally observed during this execution, before the outage>

Unverified
- <anything requiring a live system to confirm>

Blocker
- <exact connectivity failure: which command, which error>

Confidence
| Item | Confidence |
|---|---|
| ... | High/Unknown |

Next Verification
<exact resume commands, nothing else>

Exit Reason: BLOCKED_EXTERNAL

Preserve last-known-good state (mark it Verified, not stale/discarded), separate it cleanly from what can no longer be checked (Unverified), and stop. Resume only by literally running the Next Verification commands — do not pad the wait with more inference, and do not fall back to stale local caches/docs and present them as if they answer the live question.

12.1 "Unreachable" vs "Down"

Do not report a host/VM/service as down, offline, or crashed unless you have host-side evidence: cloud console instance state, serial console output, IPMI, or an equivalent direct-to-host diagnostic. Application-level and network-path failures — SSH timeout, HTTP/Cloudflare 5xx, a CLI health-check tool exiting non-zero, a dashboard reporting a diagnosis — only prove the runtime is unreachable through those paths. They do not prove the underlying host is powered off or crashed; a cloud console can report an instance "Running" while every application-level path to it fails (network partition, firewall, a crashed service inside a still-running VM, DNS/Cloudflare issue).

Correct wording once all available network paths have failed: "runtime unreachable," "origin unreachable," or "unreachable through all available paths, root cause unknown pending console/serial inspection." This is a wording and evidence-tier discipline layered on top of the BLOCKED_EXTERNAL receipt above — it does not license continued digging once paths are exhausted.


This standard governs evidence gathering. It does not govern implementation or architecture decisions.