Model Routing and Availability Architecture¶
Ultimate Outcome: SELF_HOSTED_FIRST=YES, PAID_DEFAULT=NO, ONE_ROUTING_AUTHORITY=YES, ONE_MODEL_FOR_ALL_WORKLOADS=NO. The factory operates natively on self-hosted inference, escalating to paid models only when explicitly authorized and required by the workload.
THE ROUTING AUTHORITY MATRIX¶
- LiteLLM / Aperture = concrete provider/model routing authority
- Gas City = selects semantic workload/model class
- OSSA / agentictools/agents = agent capability intent
- agent-docker = runtime projection/deployment
- OpenClaw = consumer only
- 1Password = secret authority
No consumer should hardcode concrete cloud model IDs unless the use is explicitly an approved escalation path.
SEMANTIC MODEL CLASSES¶
Each semantic class must resolve to the best available self-hosted model for that specific workload.
- ROUTER_MODEL: Smallest reliable self-hosted model
- DEFAULT_AGENT_MODEL: General-purpose self-hosted model
- CODE_MODEL: Strongest self-hosted coding model
- REASONING_MODEL: Strongest self-hosted reasoning model
- VISION_MODEL: Self-hosted multimodal model (if available)
- EMBEDDING_MODEL: Self-hosted embedding model (if available)
- PAID_ESCALATION_MODEL: Governed external model
Rule: Do not assume llama3.1 is suitable for all classes.
PROVIDER DISCOVERY & TOPOLOGY¶
Do not assume Oracle Ollama is the singular inference host. Inventory all reachable providers before mapping classes: - Oracle Ollama - LM Studio / blu host - Existing NAS inference - vLLM (if deployed) - Any current OpenAI-compatible self-hosted endpoint
LiteLLM must hide this topology from consumers.
MIGRATION AND BENCHMARK GATES¶
1. No Blind Bulk Rewrites¶
For every concrete model pin being replaced (e.g., 50+ OSSA manifests), establish: - AGENT= - CURRENT_MODEL= - CURRENT_PURPOSE= - STALE_OR_INTENTIONAL= - REQUIRED_CAPABILITIES (Min context, tool calling, vision, code, reasoning)= - TARGET_SEMANTIC_CLASS=
2. Benchmark Gate¶
Candidate local models must pass a bounded representative test suite for their class: - ROUTING_CLASSIFICATION - TOOL_CALLING - STRUCTURED_JSON - CODE_EDIT - CODE_REVIEW - LONG_CONTEXT - REASONING - DRUPAL_TASK - GITLAB_TASK - GAS_CITY_TASK
Use the smallest model that reliably passes the specific workload class.
3. Failover & Escalation¶
- Paid Provider Disabled: Factory execution must continue normally for all standard workloads.
PAID_API_REQUIRED=NO. - Paid Escalation: Must be explicit. (Task exceeds local capability -> agent records reason -> authorized escalation policy ->
PAID_ESCALATION_MODEL-> evidence/cost recorded). Silently falling back to paid inference is prohibited.
DEPLOYMENT¶
Do not mutate Oracle directly. All routing changes flow through: feature/* -> release/v0.1.x -> CI -> main -> immutable artifact/config -> governed deployment -> runtime verification -> Witness