Skip to content

Inference Topology

One owner for where model weights live, who serves them, and which gateway is the AI front door. config/litellm.yaml in agent-platform/infra/agent-docker is a projection of this document, not the reverse. Do not create a parallel LLM, Ollama, model-storage, or Aperture standard.

Storage placement on the NAS share is Wave 6 of NAS-STORAGE-CONVERGENCE-PLAN.md. Product endpoint notes live in control-room-POC.md.

Binding topology

Layer Authority Must not
Model storage SoR NAS AgentPlatform LLM: /volume1/AgentPlatform/LLM (Mac mount when present: /Volumes/AgentPlatform/LLM) Mac disk as the authoritative store; Oracle disk as a model store
Registry /volume1/AgentPlatform/LLM/llms.txt — register before download Pull first and document later
On-prem inference NAS Ollama at blueflynas.tailcf98b3.ts.net:11434; LM Studio at http://blu.tailcf98b3.ts.net:1234 Make the Mac the always-on inference host; treat /volume1/docker/services/ollama-models as the registry SoR
Mac Thin client / development workstation. Local model copies may exist only as cache or temporary test state Require the Mac to be awake for shared inference; make Mac-local Ollama the production coder backend
Oracle Gas City/runtime and proxy/control-plane workloads only Ollama/vLLM model weights, GPU-farm role, or filling/slowing Oracle with local models
Northbound AI front door Tailscale Aperture at https://ai.tailcf98b3.ts.net/ Build another gateway or expose per-device model servers as separate authorities
Other devices Phone, iPad, and other computers consume inference over the private network through the governed front door / NAS-backed route Depend on the Mac being online

Operator outcome

The durable outcome is not merely that weights sit on NAS storage. The NAS must store and serve the shared local models so that:

  • the Mac is not slowed by acting as the shared model host;
  • models continue serving when the Mac is shut down;
  • phone, iPad, and other computers can use the same models over Tailscale;
  • Oracle remains free of local model weights and does not become a GPU/model host;
  • one governed route is used instead of adding another gateway.

Future on-prem provider canaries must target NAS-hosted inference, not a Mac-local listener.

Ownership chain

This document (topology + model-storage SoR)
        |
        v
agent-platform/infra/agent-docker
  (LiteLLM configuration + compose projections)
        |
        v
NAS Ollama + NAS/Oracle LiteLLM projections
        |
        v
Aperture (ai.tailcf98b3.ts.net) — northbound front door
  • Weights: /volume1/AgentPlatform/LLM. Not Git. Not Oracle. Not Mac $HOME.
  • LiteLLM file owner: git@gitlab-bluefly:blueflyio/agent-platform/infra/agent-docker.git, config/litellm.yaml.
  • Obsolete second copy: blueflyio/blu/blutown agent_docker/config/litellm.yaml is not consulted at runtime and is not a competing authority.
  • Vast.ai path: docker-compose.nas.yml has historically named vastai-gpu.tailcf98b3.ts.net:11434; that is not the model-storage SoR and must not become a replacement architecture by accident.

Implementation belongs in agent-docker / IaC. This document does not itself mutate Aperture, Oracle, Gas City, compose, LiteLLM, Ollama, or device clients.

Current observed state — 2026-09-09

The following is evidence, not desired architecture.

Observation Classification
NAS Ollama reachable at blueflynas.tailcf98b3.ts.net:11434 LIVE
NAS live tags observed: qwen3-coder:30b, qwen2.5:7b, qwen2.5:0.5b, nomic-embed-text LIVE / verified serving coder model
LM Studio Tailscale Upstream configured in Gas City (http://blu.tailcf98b3.ts.net:1234/v1) LIVE / verified upstream routing
NAS LiteLLM coding alias routed to ollama/qwen3-coder:30b at blueflynas:11434 with Mac failover CONVERGED / NAS primary
OpenCode CLI runtime configured to use ollama-nas/qwen3-coder:30b as default model CONVERGED / client pointed to NAS
NAS LiteLLM port 4050 bound to 127.0.0.1:4050 instead of 0.0.0.0:4050 in compose DEFECT / Tailscale remote reachability blocked from Mac
/volume1/AgentPlatform/LLM/ollama/data serving store populated for coder model CONVERGED
NAS Docker model volume /volume1/docker/services/ollama-models contains blobs/manifests but is not the AgentPlatform registry SoR NON-AUTHORITATIVE STORE
NAS already holds LM Studio/MLX model material under the AgentPlatform LLM library, including Qwen3-Coder material DISK ASSET / not proof of Ollama serving
Mac-local Ollama retains qwen3-coder:30b from ~/.ollama as local cache/failover only LOCAL CACHE ONLY — shared inference unblocked on NAS
Oracle :11434 is not serving Ollama and has no required local model weights CORRECT
Aperture is reachable as the front door; reachability alone does not prove a working NAS-backed model request PARTIAL

Do not interpret a successful Mac-local model test as convergence. It proves only that the laptop cache/listener works.

Model identity and registration rules

Different tags and formats are different artifacts. Do not silently treat these as equivalent:

qwen3-coder:30b
qwen3:30b-a3b
Qwen3-Coder-30B-A3B-Instruct-MLX-4bit
Qwen3-Coder-Next-GGUF

For a shared Ollama model:

  1. register the exact intended model identity in the NAS LLM registry;
  2. populate the NAS-backed Ollama store through supported Ollama behavior;
  3. prove NAS-local inference;
  4. prove remote access from the Mac over Tailscale;
  5. repoint clients away from Mac-local listeners;
  6. prove Aperture / LiteLLM routing where those layers are intended;
  7. prove phone/iPad/other-computer access;
  8. only then classify Mac-local weights for cleanup.

Do not manually copy arbitrary Ollama blobs between stores as a substitute for a supported model-registration / population path.

Decision matrix

Local groups exist as LiteLLM aliases. They must target NAS-hosted inference once the required weights are registered and serving from the governed NAS model store. They must not depend on the Mac. Cloud groups remain vendor-backed through the governed gateway/proxy path. Oracle may proxy; it must not store local model weights.

Model group Owning service Runs where Target state
chat / reasoning / coding / fast / cheap NAS Ollama + LiteLLM alias + Aperture front door NAS for local models NAS-backed; no Mac dependency
embedding NAS Ollama NAS NAS-backed
claude / claude-fast Anthropic through governed proxy/gateway path Cloud No local-weight requirement

Completion contract

Inference convergence is complete only when all of the following are true:

NAS_MODEL_STORE=/volume1/AgentPlatform/LLM
NAS_OLLAMA_RUNNING=YES
NAS_REQUIRED_CODER_MODEL_REGISTERED=YES
NAS_REQUIRED_CODER_MODEL_SERVING=YES
NAS_LOCAL_INFERENCE=PASS
MAC_TO_NAS_INFERENCE=PASS
QWEN_OR_OTHER_PRIMARY_CLIENT_TARGETS_NAS=YES
APERTURE_OR_GOVERNED_FRONT_DOOR_TO_NAS=PASS
IPHONE_ACCESS=PASS
IPAD_ACCESS=PASS
OTHER_COMPUTER_ACCESS=PASS
MAC_REQUIRED_FOR_SHARED_INFERENCE=NO
ORACLE_LOCAL_MODEL_WEIGHTS=0

A model copied onto NAS but still served by the Mac does not satisfy this contract. NAS must be both the durable model store and the always-on local inference host.

Historical intent

Date Change
2026-07-07 LiteLLM created; NAS/local and cloud-provider routes began converging.
2026-07-14 Coding/model groups were deliberately repointed toward Mac-hosted inference.
2026-07-20 Mac routing was investigated as an intentional choice rather than silent drift.
2026-09-02 Operator closed the target matrix: NAS weights + NAS serving + Aperture front door; Mac client; Oracle no model weights.
2026-09-09 Live state still showed the Mac serving the primary coder while NAS served only small models; convergence therefore remained incomplete.

July Mac routing was a committed choice, not silent drift. It is no longer the target architecture.

Provenance

Durable findings consolidated from the existing BluCity-Docs architecture, NAS/agent-docker estate inspection, and live operator verification through 2026-09-09. Temporary chat, canvas, Scratch, and agent-session artifacts are not parallel authorities; durable findings belong here or in the other existing BluCity-Docs owner documents.