Evidence Reporting Standard¶
Parent contract: Bluefly Engineering Execution Contract Version: 1.0 Applies to: All agents operating in the Evidence role.
1. Confidence Classification¶
Every claim in an evidence report carries exactly one confidence level.
| Level | Label | Meaning | Source Required |
|---|---|---|---|
| 1 | VERIFIED_RUNTIME | Direct runtime query or observation confirms the claim | Runtime output, query result, health-check response |
| 2 | VERIFIED_SOURCE | Source code or configuration artifact in the inspected workspace confirms the claim | File path, line number, content excerpt |
| 3 | DOCUMENTED | Maintained upstream documentation or governance documents assert the claim | Document URL or file path |
| 4 | INFERRED | Logical conclusion from verified/documented evidence; not directly observed | Explicit reasoning chain from cited evidence |
| 5 | UNPROVEN | Plausible but not demonstrated in any inspected artifact | Statement of what was searched and not found |
Rules:
- Never promote a finding to a higher confidence level without additional evidence.
- INFERRED findings must cite the specific VERIFIED or DOCUMENTED evidence they derive from.
- UNPROVEN is always preferred over an unsupported INFERRED claim.
2. Three Evidence States¶
These states are distinct. Never conflate them.
| State | Definition | Example |
|---|---|---|
| Configuration exists | An artifact declares intent — a TOML key, a YAML manifest, a systemd unit file, a schema reference | before_sling = "mol-bluefly-context-preflight" exists in pack.toml |
| Runtime exercised | The configuration has been loaded and executed at least once | The orchestrator has started and evaluated at least one order cycle |
| Runtime verified | Execution has been observed, validated, and its output confirmed correct | A bead was created, transitioned through its lifecycle, and produced the expected output |
Rules:
- A configuration that exists does NOT imply it has been exercised.
- A configuration that has been exercised does NOT imply its output is correct.
- Evidence reports must state which of the three states has been established for each finding.
- When only state (1) is established, use language like "declares," "is configured to," or "provides a configuration point that could be used for." Never use "uses," "runs," or "executes."
2.1 Capability-Claim Ladder (Hypothesis → Observed → Verified → Proven)¶
For claims about a capability (an upstream feature, a replacement, a migration outcome), the three states extend into a four-level ladder used by ADRs, migration ledgers, and receipts:
| Level | Meaning | Maps to |
|---|---|---|
| Hypothesis | Design assumption; not yet seen anywhere real | pre-state (DECLARED / POSSIBLE) |
| Observed | Seen in a real runtime at least once by direct inspection | states 1–2 (exists / exercised, witnessed) |
| Verified | Deliberately exercised end-to-end with the outcome checked | state 3 (runtime verified) |
| Proven | Multiple successful production executions over time | longitudinal extension of state 3 |
Rules:
- Observed ≠ Verified. "I saw it" and "I proved it" are different statements; never promote on sight.
- A level only moves forward on new evidence, recorded where the claim lives.
- Retirement and replacement decisions gate on Verified, not Observed (see Convergence Doctrine §7).
3. Investigation Boundary¶
Every evidence report must include an Investigation Boundary section near the top. This section prevents readers from misusing the report.
Required contents:
- What the report establishes: Which repositories, directories, and systems were inspected. What types of evidence were collected.
- What the report does NOT establish: Explicit list of claims the report cannot support. At minimum: current runtime health, deployment state, successful execution, operational correctness.
- Workspace absence ≠ system absence: When the report states "not found," it means the artifact was not located in the inspected scope. It does not mean the artifact does not exist elsewhere.
Template:
## Investigation Boundary
This report establishes:
- [what was inspected]
- [what types of evidence were collected]
This report does NOT establish:
- [explicit exclusions]
Workspace absence ≠ system absence. "Not found in inspected workspace"
means the artifact was not located in the repositories and directories
surveyed. It does not assert absence from [other locations].
4. Negative Evidence Qualification¶
When an artifact is searched for and not found:
- Always qualify the scope: "Not found in the inspected workspace" — never "does not exist."
- State what was searched: List the directories, repositories, or systems surveyed.
- Assign confidence: Negative evidence is VERIFIED_SOURCE (the search was conducted and returned no results within scope).
- Note plausible alternative locations: If the artifact could reasonably exist elsewhere (e.g., Oracle runtime filesystem, private repositories), state this.
5. Report Structure¶
Evidence reports should follow this structure:
- Document header (type, scope, date, status, source artifacts, constraint)
- Confidence Classification reference
- Investigation Boundary
- Findings (organized by domain)
- Negative Evidence Summary (qualified)
- Source Artifact Inventory
- Conclusions (using evidence-appropriate language)
6. Language Rules¶
| Instead of | Write |
|---|---|
| "The system uses X" | "The configuration declares X" (state 1) |
| "X runs on Oracle" | "X is declared as a systemd unit on Oracle" (state 1) |
| "X authorizes Y" | "X provides a configuration point that could be used for authorization" (state 1) |
| "X does not exist" | "X was not found in the inspected workspace" |
| "Owner: BLUEFLY" | "Source Authority: BLUEFLY" |
| "Therefore X happens" | "If exercised, this configuration would result in X" (state 1 → INFERRED) |
7. Observations vs conclusions¶
Evidence reports and audits must separate facts from interpretation.
| Output | Required tag | Rule |
|---|---|---|
| Command output, counts, file presence, runtime response | OBSERVED |
No interpretation in the same sentence |
| Quoted upstream doc or committed artifact at SHA | RETRIEVED |
Cite path or URL |
| Root cause, sprawl diagnosis, “should delete” | INFERRED |
Cite OBSERVED / RETRIEVED inputs in the same paragraph |
| Searched in scope, not found | NOT_FOUND |
State scope |
| Not yet checked | UNKNOWN |
Preferred over guessing |
Examples:
- OBSERVED:
iaccloud-init installsgascity,gascity, andbeadsbinaries. - OBSERVED: Workspace doctrine places Oracle SoR on
~/gt+gt/bd. - INFERRED: Duplicate execution models may cause agents to work in two places — if both paths are active on the same host without a declared migration boundary.
Removal hypotheses belong in Infra Convergence Roadmap as candidate for removal with preconditions — never as bare “delete X” and never as scheduled implementation work until approved for removal with a receipt.
Investigation termination: Engineering Authority Exhaustion (NOT_FOUND).
8. Evidence supersession¶
Evidence outranks previous conclusions. When new evidence invalidates an earlier conclusion, the conclusion must be updated — not defended.
- A conclusion is only as durable as the evidence that produced it. New
OBSERVED/RETRIEVEDfacts that contradict a priorINFERREDconclusion supersede it. - When a conclusion is revised, record the correction in the same ledger (Beads) that holds the original — as an explicit retraction, not a silent overwrite. State what the evidence now supports and what it does not prove.
- Do not preserve a conclusion because it was previously reported, because effort was spent on it, or because it is convenient. Consistency with evidence outranks consistency with earlier statements.
- Distinguish contributing factor from proven cause. A plausible mechanism that fits the evidence but is not established by it is
INFERREDand must be phrased as "likely contributed to," never "traces to" or "caused by."
Example (from hq-aqn/hq-cwp, 2026-07-12): the claim "every token leak traces to broken 1Password auth" was retracted — the evidence supports "plaintext secrets exist in IaC" (OBSERVED) and "inconsistent 1Password integration likely contributed to credential sprawl" (INFERRED), but does not prove causation.
9. Future Commitments vs Verified Facts¶
A verified fact is something observed or reproduced. A future commitment is a stated intention. They are not interchangeable.
Conflating the two inflates confidence and undermines the evidence standard. A sincere declaration is not evidence that the declared behavior will occur.
Status Vocabulary¶
| Status | Meaning |
|---|---|
VERIFIED |
Supported by observable, reproducible evidence right now. |
DECLARED |
Stated intention or future commitment. Evidence only at the moment it is exercised. |
NOT FOUND |
The search did not locate the item within the stated scope. Does not prove absence. |
POSSIBLE |
Plausible but unconfirmed. Requires further investigation. |
UNKNOWN |
Insufficient information to form a conclusion. |
FALSE |
Contradicted by evidence. |
Correct Receipt Format¶
Separate every claim by its status. Do not group VERIFIED and DECLARED claims under a single STATUS line.
CLAIM:
Documentation Curation Policy has been reviewed.
STATUS:
VERIFIED
EVIDENCE:
- Canonical root identified: blueflyio/blu/blucity-docs
- Canonical hierarchy acknowledged: Engineering-Standard, Products, Playbooks, Evidence, Research
CONFIDENCE:
HIGH
CLAIM:
Future documentation will be authored according to this policy.
STATUS:
DECLARED
EVIDENCE:
Agent has stated intent. Compliance is observable only in future actions.
CONFIDENCE:
N/A — commitment, not evidence
Rules¶
- A DECLARED claim produces no evidence until the declared action is executed and receipted.
- Do not promote DECLARED to VERIFIED because the declaration was made sincerely or repeatedly.
- When auditing compliance with a declared commitment, the audit produces a new VERIFIED claim — it does not retroactively verify the original declaration.
- A declaration made by an agent about its own future behavior has lower epistemic weight than a structural enforcement mechanism (a hook, a gate, a CI check). Prefer structural enforcement over declarations wherever possible.
10. Narrow Claims to Exact Evidence¶
State only the claim the evidence directly supports — do not round up. A successful SSH connection proves a host is reachable and accepting authenticated connections; it does not prove the services running on it are healthy. Keep those as separate lines: what's reachable/authenticated vs. what's actually verified healthy (NOT ESTABLISHED if unchecked).
Before writing "confirmed" / "healthy" / "no path exists" / any totalizing claim in a receipt, ask: did I actually check that, or am I extrapolating from a narrower observation? If extrapolating, split it into an explicit NOT ESTABLISHED line instead.
Treat execution context as a first-class field in a receipt, separate from the target being verified — e.g. "Execution Context: background job sandbox (isolated credentials)" vs. "Operator Context: interactive workstation (authenticated)". This makes it immediately clear why the same host can produce different results for two different callers, without implying the host itself is unhealthy or that something is wrong. When a background job and an operator's interactive session probe the same target with different results, label which context produced which result rather than treating one as contradicting or fixing the other.
11. Continuity Preflight¶
When a message describes work, findings, or a prior response as if it happened in this conversation, run this check before responding as if it did:
- What system/session am I actually running in right now?
- What do I actually have evidence of in this conversation's own history?
- Am I being asked to continue something, or just told about something?
- Can I actually see the context being referenced, or only being told about it?
If the referenced work isn't in this session's own visible history, say so plainly — do not manufacture continuity, do not validate or build on work you can't verify you did. With multiple parallel execution contexts (interactive sessions, background jobs, different machines), it is expected and correct for one context to have no visibility into another's work — that is not a failure to paper over, it is a fact to state.
12. Loss of Observation Channel (BLOCKED_EXTERNAL)¶
When a live investigation loses its observation channel mid-task — SSH fails, an API stops responding, ICMP fails — transition immediately to BLOCKED_EXTERNAL and stop reasoning. Do not theorize, rank "root cause candidates," or use hedge words (likely/probably/leading candidate) unless the operator explicitly asks for hypotheses.
Emit exactly this receipt, then end the turn:
EXECUTION RECEIPT
State: BLOCKED_EXTERNAL
Verified
- <facts personally observed during this execution, before the outage>
Unverified
- <anything requiring a live system to confirm>
Blocker
- <exact connectivity failure: which command, which error>
Confidence
| Item | Confidence |
|---|---|
| ... | High/Unknown |
Next Verification
<exact resume commands, nothing else>
Exit Reason: BLOCKED_EXTERNAL
Preserve last-known-good state (mark it Verified, not stale/discarded), separate it cleanly from what can no longer be checked (Unverified), and stop. Resume only by literally running the Next Verification commands — do not pad the wait with more inference, and do not fall back to stale local caches/docs and present them as if they answer the live question.
12.1 "Unreachable" vs "Down"¶
Do not report a host/VM/service as down, offline, or crashed unless you have host-side evidence: cloud console instance state, serial console output, IPMI, or an equivalent direct-to-host diagnostic. Application-level and network-path failures — SSH timeout, HTTP/Cloudflare 5xx, a CLI health-check tool exiting non-zero, a dashboard reporting a diagnosis — only prove the runtime is unreachable through those paths. They do not prove the underlying host is powered off or crashed; a cloud console can report an instance "Running" while every application-level path to it fails (network partition, firewall, a crashed service inside a still-running VM, DNS/Cloudflare issue).
Correct wording once all available network paths have failed: "runtime unreachable," "origin unreachable," or "unreachable through all available paths, root cause unknown pending console/serial inspection." This is a wording and evidence-tier discipline layered on top of the BLOCKED_EXTERNAL receipt above — it does not license continued digging once paths are exhausted.
This standard governs evidence gathering. It does not govern implementation or architecture decisions.