STD-CONTEXT-002 — Agent Context Economy¶
Standard ID: STD-CONTEXT-002 Applies to: every agent execution context run against Bluefly systems — interactive, background, subagent, cloud, and scheduled.
Relationship to existing standards. STD-CONTEXT-001 governs the shape of a handoff between contexts. ADR-0019 governs execution-context identity. STD-DOC-001 §6 "Agent context budget" governs auto-loaded instruction files. This standard governs what none of them do: how much context an agent may acquire, how long it may hold it, and what that costs. It references those documents rather than restating them.
1. The failure this standard exists to stop¶
The conversation is not the knowledge store. The session is disposable execution context.
OBSERVED — a two-day fleet session consumed approximately 1.3M output tokens, 543.9M
cached-read tokens and 6.1M cache-write tokens, at roughly \$365. In that window 97% of sessions
ran 8+ hours, 76% exceeded 150k context, 53% of spend went to general-purpose subagents, and 14%
to a single MCP surface.
INFERRED — those numbers are not a documentation-volume problem. Agents were ingesting too much
documentation too early and carrying it too long, using the conversation as the durable store
that Git, Beads and ledger/ already are. The remedy is retrieval discipline, not fewer
documents.
2. Never load the whole company¶
An agent receives only its role, its project, its Bead, its immediate dependencies, and the minimum authoritative context the task requires. Everything else is retrieval-on-demand.
MUST NOT preload: the blucity-docs tree, whole Bead histories, large GitLab MR histories,
unrelated architecture, complete incident ledgers, full repository trees, or another project's
context.
3. Context budget¶
| Threshold | Meaning |
|---|---|
| 100k | Warning. Start reducing. |
| 150k | Hard intervention. Do not cross it without discharging state first. |
Before crossing 150k an agent MUST: summarise durable findings; write or update the Bead; update the owning document if a durable fact changed; commit and push; produce a compact handoff per STD-CONTEXT-001; then compact or start a fresh session.
A long-running session MUST NOT be kept alive merely because it already knows the history.
History that matters belongs in Git, Beads or ledger/ — see §8.
4. Task-bounded sessions¶
START → read Bead → retrieve required authority → execute one bounded objective
→ test → commit/push → update Bead → write compact handoff → END
A session SHOULD correspond to one Bead or a tightly bounded portion of one.
5. Subagents are expensive compute, not free workers¶
MUST NOT spawn a general-purpose high-end subagent for mechanical work: file listing, grep, branch inventory, duplicate detection, link checking, simple classification, formatting, or summarisation. Use the cheapest tool that returns the same evidence — in practice a shell command, not a reasoning model.
Reserve the strongest model for architecture, ambiguous debugging, synthesis, policy, high-risk reasoning, and complex implementation.
Default concurrency: one primary agent and zero to two bounded workers. More requires a stated reason. Optimise for independent useful work per token consumed, never for maximum concurrency.
6. Every worker gets a contract¶
A spawned worker MUST carry: ROLE, OBJECTIVE, INPUTS, AUTHORITY, OUTPUT,
STOP_CONDITION.
"general-purpose — investigate" is not a contract. "inventory these 20 systemd units and
return this 8-column table" is.
7. MCP output is not memory¶
MUST NOT repeatedly dump large MR objects, complete pipeline payloads, whole note threads or large API responses into the primary conversation. Retrieve the narrow field needed, summarise the useful state into the Bead or the owning document, then flush.
8. Session memory is non-authoritative¶
If losing the current session would lose something important, stop and persist it first — to
source, to the Bead, to ledger/, or to the owning document. Then continue. An agent that is the
only holder of a fact is an outage waiting to happen.
9. ContextControl is the retrieval layer¶
As ContextControl becomes available, the request is "what context applies to Bead X, as role R,
in project P?" and the answer names what is required, what is optional, what is out of
scope, plus authority, provenance and last_reviewed per item. Returning a corpus is a defect.
The contract and precedence rules are STD-DOC-001 §8.5.
NOT_ESTABLISHED: no running instance at the time of writing (bc-7e3).
10. Gemini Notebook is a research mirror¶
blucity-docs remains canonical. Curated evergreen material MAY be published into a notebook for
cross-document research, synthesis and discovery. Notebook output is not authority. A durable
conclusion travels: notebook research → evidence in ledger/ → review → owning Git document →
republication. See STD-DOC-001 §8.4.
11. Model routing¶
cheap worker gathers → structured evidence → strong model reasons
Not: several expensive workers each re-reading the same large context. Reduce the evidence first, reason second.
12. Handoffs are small¶
A handoff carries BEAD, OBJECTIVE, COMPLETED, PROVEN, CHANGED, MR, BLOCKERS,
NEXT, REQUIRED_CONTEXT_LINKS — and normally stays under 1,500 words. Never a transcript.
Field definitions are STD-CONTEXT-001's; do not redefine them here.
Metrics¶
Tracked weekly. These are the standard's compliance signal, not decoration.
| Metric | Target |
|---|---|
SESSIONS_OVER_150K |
exceptional, not routine |
SESSIONS_OVER_8H |
exceptional, not routine |
GENERAL_PURPOSE_SUBAGENT_SHARE |
sharply reduced |
OPUS_SUBAGENT_SHARE |
reasoning work only |
MCP_CONTEXT_SHARE |
declining |
AVG_SESSION_CONTEXT |
declining |
CACHE_READ_TOKENS / CACHE_WRITE_TOKENS |
tracked |
COST_PER_COMPLETED_BEAD |
declining |
TOKENS_PER_COMPLETED_BEAD |
declining |
NOT_ESTABLISHED: no collector currently emits these. Until one does, they are reported from
provider usage data at review time.