Skip to content

STD-CONTEXT-002 — Agent Context Economy

Standard ID: STD-CONTEXT-002 Applies to: every agent execution context run against Bluefly systems — interactive, background, subagent, cloud, and scheduled.

Relationship to existing standards. STD-CONTEXT-001 governs the shape of a handoff between contexts. ADR-0019 governs execution-context identity. STD-DOC-001 §6 "Agent context budget" governs auto-loaded instruction files. This standard governs what none of them do: how much context an agent may acquire, how long it may hold it, and what that costs. It references those documents rather than restating them.


1. The failure this standard exists to stop

The conversation is not the knowledge store. The session is disposable execution context.

OBSERVED — a two-day fleet session consumed approximately 1.3M output tokens, 543.9M cached-read tokens and 6.1M cache-write tokens, at roughly \$365. In that window 97% of sessions ran 8+ hours, 76% exceeded 150k context, 53% of spend went to general-purpose subagents, and 14% to a single MCP surface.

INFERRED — those numbers are not a documentation-volume problem. Agents were ingesting too much documentation too early and carrying it too long, using the conversation as the durable store that Git, Beads and ledger/ already are. The remedy is retrieval discipline, not fewer documents.


2. Never load the whole company

An agent receives only its role, its project, its Bead, its immediate dependencies, and the minimum authoritative context the task requires. Everything else is retrieval-on-demand.

MUST NOT preload: the blucity-docs tree, whole Bead histories, large GitLab MR histories, unrelated architecture, complete incident ledgers, full repository trees, or another project's context.

3. Context budget

Threshold Meaning
100k Warning. Start reducing.
150k Hard intervention. Do not cross it without discharging state first.

Before crossing 150k an agent MUST: summarise durable findings; write or update the Bead; update the owning document if a durable fact changed; commit and push; produce a compact handoff per STD-CONTEXT-001; then compact or start a fresh session.

A long-running session MUST NOT be kept alive merely because it already knows the history. History that matters belongs in Git, Beads or ledger/ — see §8.

4. Task-bounded sessions

START → read Bead → retrieve required authority → execute one bounded objective
      → test → commit/push → update Bead → write compact handoff → END

A session SHOULD correspond to one Bead or a tightly bounded portion of one.

5. Subagents are expensive compute, not free workers

MUST NOT spawn a general-purpose high-end subagent for mechanical work: file listing, grep, branch inventory, duplicate detection, link checking, simple classification, formatting, or summarisation. Use the cheapest tool that returns the same evidence — in practice a shell command, not a reasoning model.

Reserve the strongest model for architecture, ambiguous debugging, synthesis, policy, high-risk reasoning, and complex implementation.

Default concurrency: one primary agent and zero to two bounded workers. More requires a stated reason. Optimise for independent useful work per token consumed, never for maximum concurrency.

6. Every worker gets a contract

A spawned worker MUST carry: ROLE, OBJECTIVE, INPUTS, AUTHORITY, OUTPUT, STOP_CONDITION.

"general-purpose — investigate" is not a contract. "inventory these 20 systemd units and return this 8-column table" is.

7. MCP output is not memory

MUST NOT repeatedly dump large MR objects, complete pipeline payloads, whole note threads or large API responses into the primary conversation. Retrieve the narrow field needed, summarise the useful state into the Bead or the owning document, then flush.

8. Session memory is non-authoritative

If losing the current session would lose something important, stop and persist it first — to source, to the Bead, to ledger/, or to the owning document. Then continue. An agent that is the only holder of a fact is an outage waiting to happen.

9. ContextControl is the retrieval layer

As ContextControl becomes available, the request is "what context applies to Bead X, as role R, in project P?" and the answer names what is required, what is optional, what is out of scope, plus authority, provenance and last_reviewed per item. Returning a corpus is a defect.

The contract and precedence rules are STD-DOC-001 §8.5. NOT_ESTABLISHED: no running instance at the time of writing (bc-7e3).

10. Gemini Notebook is a research mirror

blucity-docs remains canonical. Curated evergreen material MAY be published into a notebook for cross-document research, synthesis and discovery. Notebook output is not authority. A durable conclusion travels: notebook research → evidence in ledger/ → review → owning Git document → republication. See STD-DOC-001 §8.4.

11. Model routing

cheap worker gathers  →  structured evidence  →  strong model reasons

Not: several expensive workers each re-reading the same large context. Reduce the evidence first, reason second.

12. Handoffs are small

A handoff carries BEAD, OBJECTIVE, COMPLETED, PROVEN, CHANGED, MR, BLOCKERS, NEXT, REQUIRED_CONTEXT_LINKS — and normally stays under 1,500 words. Never a transcript. Field definitions are STD-CONTEXT-001's; do not redefine them here.


Metrics

Tracked weekly. These are the standard's compliance signal, not decoration.

Metric Target
SESSIONS_OVER_150K exceptional, not routine
SESSIONS_OVER_8H exceptional, not routine
GENERAL_PURPOSE_SUBAGENT_SHARE sharply reduced
OPUS_SUBAGENT_SHARE reasoning work only
MCP_CONTEXT_SHARE declining
AVG_SESSION_CONTEXT declining
CACHE_READ_TOKENS / CACHE_WRITE_TOKENS tracked
COST_PER_COMPLETED_BEAD declining
TOKENS_PER_COMPLETED_BEAD declining

NOT_ESTABLISHED: no collector currently emits these. Until one does, they are reported from provider usage data at review time.