STD-COST-001: Bluefly Context Retrieval and Model Cost Optimization¶
Status: Approved Governing Standard
Authority: Thomas P. Scola Jr. — Factory Operating Contract & Economic Gate
Date: 2026-09-23
Bead: bl-cost01 / bc-cost01
Cross-References: STD-ECON-001, STD-AUTH-001, STD-CONTEXT-001, STD-GATE-001
Objective¶
Reduce paid-model usage, context-window waste, repeated repository exploration, and agent rediscovery without creating another orchestration platform or another source of truth.
The optimization hierarchy is:
DON'T SEND IT TO A MODEL
↓
RETRIEVE ONLY WHAT IS NEEDED
↓
USE LOCAL COMPUTE
↓
USE CHEAP MODEL
↓
USE STANDARD MODEL
↓
USE EXPENSIVE MODEL ONLY WHEN REQUIRED
The goal is not merely cheaper tokens. The goal is:
LESS_CONTEXT_INGESTION
LESS_REDISCOVERY
FEWER_TOOL_CALLS
FEWER_FAILED_RUNS
LESS_PROVIDER_DEPENDENCE
LOWER_PAID_MODEL_COST
BETTER_EVIDENCE
1. Authority Model¶
Nothing in this initiative becomes a new authority.
GitLab = source / CI / releases
Gas City = orchestration
Beads/Dolt = durable work
BluCity-Docs = governed doctrine
AgenticTools/agents = reusable Agent source
AgenticTools/skills = reusable Skill source
AgenticTools/plugins = reusable Plugin/tool integration source
ContextControl = governed human/context surface
LiteLLM = model gateway
QMD, Orbit, CodeGraph and similar systems are:
INDEXES
CACHES
RETRIEVAL TOOLS
They are NOT authorities. Everything they contain must be rebuildable from authoritative source.
2. Documentation Retrieval — Adopt QMD¶
QMD fits Bluefly extremely well. It combines local BM25 full-text retrieval, vector retrieval, query expansion, and local LLM reranking. It runs on-device rather than requiring another hosted retrieval service.
Use QMD for human-readable governed knowledge.
Primary collection: BluCity-Docs
Potential additional collections should be added only when they contain governed durable documentation.
Do NOT mirror into QMD:
- Beads as a replacement for bd
- GitLab source as a replacement for Git
- Live runtime state
- Generated logs
- Entire repositories of source code
Those have better native owners.
Desired retrieval order for documentation questions:¶
QMD
↓
Targeted governing document
↓
Targeted source/runtime evidence if required
Never: find → grep → cat 20 markdown files → feed 40,000 lines into a frontier model.
QMD Operating Modes¶
Use the cheapest retrieval method that works:
qmd search→ exact terminology / known standard (BM25)qmd vsearch→ semantic concept (vector)qmd query→ difficult cross-document question (reranked)
Do not automatically invoke the full reranking pipeline when BM25 already resolves the question. That reduces even local inference work.
3. Code Intelligence — Orbit Local vs CodeGraph Bakeoff¶
Rule: Do NOT establish ORBIT + CODEGRAPH EVERYWHERE. They substantially overlap. Evaluate both against actual Bluefly workloads.
GitLab Orbit Local¶
Orbit Local is particularly attractive because Bluefly is GitLab-centered: - Local indexing - DuckDB store - Definitions, files, directories, imports, cross-file references - CLI + MCP interfaces - Zero GitLab Credits required - Works without a GitLab connection (local-first) - Currently beta; experimental MCP surface makes it an accelerator, not production authority.
CodeGraph¶
CodeGraph is also local-first and supports PHP along with TypeScript, JavaScript, Python, Go, Rust, Java and other languages.
- Concentrated around codegraph_explore
- Returns relevant verbatim source, call paths, and blast-radius information in a single tool call instead of requiring repeated grep/read cycles
- High utility for PHP/Drupal service container and hook discovery.
Bakeoff Specification¶
Use representative Bluefly repositories:
1. DrupalWorks
2. contextcontrol-ai
3. bluefly.io
4. BluCity
5. blucity-packs
6. iac (infrastructure)
Test real engineering queries: - Where is this Drupal service instantiated? - What calls this method? - What would changing this interface affect? - How does this request reach this plugin? - Where is this Pack capability consumed? - Which code path handles this event?
Comparative Metrics:¶
ANSWER_CORRECT
TOOL_CALLS
INPUT_TOKENS
OUTPUT_TOKENS
LATENCY
INDEX_TIME
INDEX_SIZE
STALE_RESULT_RATE
PHP_DRUPAL_QUALITY
CROSS_FILE_QUALITY
AGENT_USABILITY
Convergence Rule:¶
PRIMARY_CODE_INDEXER = ORBIT | CODEGRAPH
SECONDARY_CODE_INDEXER = NONE
Only retain both if there is a proven non-overlapping capability worth owning.
4. GitLab Orbit Remote — Pilot, Not Authority¶
Orbit Remote is compelling for Bluefly because it indexes the wider SDLC: - Projects - Source - Work items - Merge requests - Pipelines - Users - Security findings - Cross-entity relationships
GitLab labels Orbit Remote beta (available for testing, not production use).
Therefore:
ORBIT_REMOTE_PRODUCTION_DEPENDENCY = NO
ORBIT_REMOTE_AUTHORITY = NO
ORBIT_REMOTE_EXPERIMENT = YES
Use experimentally for SDLC-wide queries: - Which MRs are connected to this work? - Which projects depend on this component? - What pipeline failures cluster around this subsystem? - Which work items relate to this source?
Do not make Gas City execution depend upon Orbit availability. GitLab remains authoritative. Orbit is an accelerated query surface over that authority.
5. Adopt the Useful Orbit Documentation Pattern¶
Do not copy GitLab's monolithic documentation files. Adopt its structural pattern: - A small agent-facing map containing: - HOW THIS SYSTEM WORKS - WHAT CI ENFORCES - WHERE TO FIND THINGS - CURRENT INVARIANTS - WHAT NOT TO READ - WHAT OWNS WHAT
Design Document Rule:¶
IF behavior changes AND governing documentation describes that behavior:
UPDATE BOTH IN THE SAME DELIVERY
Do not force every code MR to modify BluCity-Docs.
DOC_CHANGE_REQUIRED = YES (only when governed architecture/policy/behavior changed)
DOC_CHANGE_REQUIRED = NO (routine implementation/bugfix within existing contract)
Repository instruction files must remain small projections of canonical doctrine, not 1,500-line prompt dumps.
6. Oh My OpenAgent — Mine the Ideas, Do Not Adopt the Orchestrator¶
Do NOT make Oh My OpenAgent another Bluefly orchestration layer.
Oh My OpenAgent already supplies its own agents, planning, delegation, skills, hooks, MCPs, background workers, model fallback, team mode, and continuation logic.
That directly overlaps: - Gas City - Bluefly configured Agents - BluCity-Packs - AgenticTools Skills - AgenticTools Plugins - Beads
Therefore:
OMO_AS_ORCHESTRATOR = NO
OMO_AS_WORK_AUTHORITY = NO
OMO_AGENT_DEFINITIONS_AS_AUTHORITY = NO
OMO_SKILLS_AS_AUTHORITY = NO
Useful Patterns to Harvest:¶
- On-demand skill-scoped MCP activation (bring scoped MCPs rather than permanent context load)
- Context compaction recovery
- Model fallback chains
- Structured plan review gates
- Tool-output truncation and streaming safety
- Continuation after context limits
- LSP integration and AST-aware code transformations
Harvest Filter:
DOES GAS CITY ALREADY OWN THIS?
DOES OUR AGENTICTOOLS STACK ALREADY OWN THIS?
IS THERE AN EXISTING UPSTREAM COMPONENT?
CAN WE CONSUME IT WITHOUT ADOPTING OMO ORCHESTRATION?
If yes: consume into agentictools/skills or agentictools/plugins. Do not fork OMO into Bluefly infrastructure.
7. AgenticTools Remains Canonical¶
Reusable behavior discovered through CodeGraph, Orbit, OMO or any upstream project must land in the correct canonical repository:
agent definition → agentictools/agents
reusable skill → agentictools/skills
runtime/tool integration→ agentictools/plugins
Gas City composition → BluCity-Packs
BluCity-Packs references and configures capabilities. It does not independently recreate reusable Agent/Skill/Plugin source.
Generated Claude, Codex, Antigravity, OpenCode, and OpenClaw representations remain disposable projections.
8. Keep LiteLLM — Do Not Add OpenGateway Yet¶
Current Architecture:
Agent → Gas City execution → LiteLLM → Provider / Self-hosted endpoint
Keep that single model-egress layer.
LiteLLM already supports: - Provider routing & fallbacks - Retries & timeout management - Rate limits & client quotas - Usage accounting & spend tracking - Response caching - Budget enforcement - Routing to zero-cost self-hosted models when paid budgets exhaust
OpenGateway provides OpenAI-compatible routing but adds provider pass-through costs plus a platform fee. Adding OpenGateway behind or alongside LiteLLM is redundant and economically unjustified.
OPENGATEWAY_CURRENT_DECISION = DO_NOT_ADOPT
Re-evaluate only if OpenGateway demonstrates a measurable capability or total-cost advantage unavailable from the existing stack.
9. Model Routing for Actual Cost Reduction¶
Meaningful dollar savings come from tier aliasing rather than letting every Agent name arbitrary frontier models.
Standard Capability Aliases:¶
blu-local(Self-hosted via LM Studio / Ollama / vLLM):- Bounded tasks: classification, search expansion, summarization, simple extraction, lint interpretation, log classification, Bead triage, documentation retrieval synthesis, simple code navigation.
blu-fast(Low-cost hosted: Flash / Haiku):- Routing, triage, simple transformations, commit-message synthesis, receipt normalization, small documentation operations.
blu-standard(Capable engineering: Sonnet / GPT-4o / Gemini Pro):- Implementation, code review, debugging, Drupal work, test generation, refactoring.
blu-deep(Frontier reasoning: Opus / O1 / Pro Thinking):- Novel architecture, difficult incidents, security analysis, high-risk migrations, complex cross-system debugging, policy design.
Binding Policy: BLU determines required capability tier. Agents must not select frontier reasoning merely because "Opus is better."
10. Budget-Aware Fallback to Local Models¶
LiteLLM supports treating explicitly configured zero-cost models as available after a paid model budget is exhausted.
PAID FAST MODEL
↓ (Budget / Rate limit / Upstream failure)
SELF-HOSTED LOCAL MODEL
Applicable to workloads whose policy allows local fallback.
Exception: Security-sensitive, cryptographic, or high-reasoning work must yield WAIT_FOR_APPROVED_CAPABILITY rather than silently degrading quality.
11. Context Cost Reduction Pipeline¶
Before an Agent sends a large request to a paid model:
QUESTION
↓
CLASSIFY
↓
┌──────────────────────────────────────────────┐
│ docs? → QMD │
│ code? → chosen code graph (Orbit / CG) │
│ work state? → Beads │
│ source/MR? → GitLab │
│ runtime? → direct observation │
└──────────────────────────────────────────────┘
↓
20–200 relevant lines / structured facts
↓
MODEL
Do not send the whole knowledge estate merely because a large context window exists.
12. Do Not Invent Savings Numbers¶
No speculative savings figures (e.g. "80-90% savings") are permitted in receipts or doctrine without direct measurement.
Bluefly must measure its own workloads across real tasks.
13. Cost Evidence & Telemetry Schema¶
Every model execution must carry sufficient attribution to calculate:
BEAD=
AGENT=
RIG=
FORMULA=
MODEL_ALIAS=
ACTUAL_MODEL=
PROVIDER=
LOCAL_OR_PAID=
INPUT_TOKENS=
OUTPUT_TOKENS=
CACHE_READ_TOKENS=
CACHE_WRITE_TOKENS=
COST=
LATENCY=
SUCCESS=
RETRY_COUNT=
Cost is an engineering variable, not an unmeasured monthly utility bill.
14. Measure Before and After¶
Capture baseline telemetry prior to altering production routing:
| Metric | Before | After | Delta |
|---|---|---|---|
| Paid input tokens per completed Bead | |||
| Paid output tokens per completed Bead | |||
| Paid model cost per completed Bead | |||
| Tool calls per completed Bead | |||
| Files read per completed Bead | |||
| Context bytes injected | |||
| Failed / restarted sessions | |||
| Median completion latency | |||
| Local-model share (%) | |||
| Expensive-model share (%) | |||
| Witness failure rate (%) |
Primary Success Metric:¶
TARGET_METRIC = COST_PER_VERIFIED_COMPLETED_BEAD
A cheap model that causes three failed runs is not cheap.
15. Implementation Order¶
- QMD first: Index
BluCity-Docsand establish a small reusable AgenticTools skill for governed-document retrieval. - Orbit Local vs CodeGraph bakeoff: Test representative Bluefly repositories and choose one default code-intelligence layer.
- Slim agent instructions: Apply the useful Orbit
AGENTS.mdpattern: architecture, ownership, CI gates, navigation, not giant duplicated doctrine. - LiteLLM cost baseline: Establish Agent/Rig/Bead/model-level spend and token telemetry before changing routing.
- Local-first aliases: Route proven low-risk workloads to self-hosted models with paid fallback.
- Tier paid models: Establish
blu-fast,blu-standard,blu-deepcapability aliases rather than hardcoded provider names in Agents. - OMO capability harvest: Evaluate individual concepts (scoped MCP activation, context compaction); adopt only where Bluefly lacks the capability.
- Orbit Remote experiment: Enable only as a read-only/context experiment. Delivery must not depend on it.
- OpenGateway remains rejected: Revisit only with measured comparative economic evidence.
Target Architecture¶
BLUEFLY FACTORY
┌───────────────┐
│ Gas City │
│ orchestration │
└───────┬───────┘
│ configured Agent
┌───────────────┼───────────────┐
│ │ │
▼ ▼ ▼
QMD CODE GRAPH GitLab / Beads
governed chosen winner durable / live
docs (Orbit OR CG) truth
│ │ │
└───────────────┼───────────────┘
│ minimal context (20-200 lines)
▼
LiteLLM
│
┌───────────────┼───────────────┐
│ │ │
▼ ▼ ▼
LOCAL MODELS CHEAP MODELS DEEP MODELS
(LM Studio/ (blu-fast: (blu-deep:
vLLM/Ollama) Flash/Haiku) Opus/Thinking)
SEPARATION OF CAPABILITY:
AgenticTools (agents/ skills/ plugins/)
↓
BluCity-Packs (composition & rig wiring)
↓
Gas City (execution runtime)
Acceptance Vector¶
QMD_GOVERNED_DOC_RETRIEVAL = PASS
CODE_INDEX_BAKEOFF_COMPLETE = YES
PRIMARY_CODE_INDEXER = ORBIT | CODEGRAPH
DUPLICATE_CODE_INDEXER_REQUIRED = NO | PROVEN_REASON
AGENTICTOOLS_AUTHORITY_PRESERVED = YES
BLUCITY_PACKS_AS_SKILL_AUTHORITY = NO
OMO_ORCHESTRATION_ADOPTED = NO
USEFUL_OMO_PATTERNS_EVALUATED = YES
LITELLM_REMAINS_SINGLE_GATEWAY = YES
OPENGATEWAY_ADOPTED = NO
LOCAL_MODEL_ROUTING_PROVEN = YES
MODEL_ALIASES_PROVIDER_INDEPENDENT = YES
MODEL_COST_PER_BEAD_MEASURABLE = YES
PAID_TOKENS_PER_BEAD_MEASURABLE = YES
COST_PER_VERIFIED_COMPLETED_BEAD_BEFORE = BASELINE_ESTABLISHED
COST_PER_VERIFIED_COMPLETED_BEAD_AFTER = MEASURED_LOWER
NEXT_RUN_CHEAPER_AND_MORE_REUSABLE = YES