12 KiB
VAUCO OS — ARCHITECTURE.md
Eier: PLAN-Claude-Architect
Sist oppdatert: 2026-06-29 CEST
Godkjent av: Chris Christiansen
Hensikt
Single source of truth for systemarkitektur, komponent-diagram og design-kontrakter. Ingen andre tråder skriver denne filen. Andre tråder kan foreslå patcher via PR.
Komponent-diagram (ASCII)
[Chris / Operatør]
|
-----+------+----------+----------+
| | | |
OPS EXEC PLAN AUDIT
Perplexity Gemini Claude Nemotron
Computer 2.5 Pro Sonnet/ (NMT)
Opus
| | | |
+-----+-----+----------+----------+
|
[GitHub]
OSVauco (main)
vauco-bootstrap (main)
|
[GCP / Cloud Run]
project: propane-will-491900-m5
region: us-central1
|
[OPAX-MCP — opax.vauco.no]
service: opax-mcp-core
protocol: OPAX Protocol
|
[Klient-OS]
<klient>-oss.vauco.no (stage)
<klient>-os.vauco.no (prod)
VM-infrastruktur (GCP Compute Engine)
| VM | Formål | Specs | Modeller |
|---|---|---|---|
osvauco-dev-vm |
Utvikling, OSV-pipeline (lokal Ollama), MCP-testing | 16 GB RAM, CPU-only, us-central1-b |
gemma3:4b (3.3 GB), qwen2.5:7b (4.7 GB), nomic-embed-text, llama3.2 |
emma-gpu-vm |
Tung ML, Emma co-pilot (Gemma 4 27B) | GPU, høy RAM | Gemma 4 27B (int4) — kun ved eksplisitt eskalering |
OSV-pipeline (agents/osv/pipeline.py)
Lokal tre-lags Ollama-pipeline på osvauco-dev-vm for å minimere API-kostnader:
Lag 1 — gemma3:4b (front) → forstår intent, strukturerer oppgave
Lag 2 — qwen2.5:7b (analyse) → dyptgående analyse, kode, planlegging
Lag 3 — gemma3:4b (output) → formaterer og leverer svar til bruker
Eskaleringspolitikk (Emma):
emmatrigger: kun hvis Chris eksplisitt skriver "emma" i meldingenask_emma: pipeline spør Chris hvis konfidans < 0.4- Aldri automatisk eskalering
- Emma kjører på
emma-gpu-vm:11434— kostnad påløper ved oppstart
Viktig — RAM-krav:
osvauco-dev-vmMÅ ha ≥16 GB RAM for å kjøre gemma3:4b + qwen2.5:7b samtidig- 8 GB er ikke tilstrekkelig — modellene timeout-er under lasting
- Øk VM til 16 GB i GCP Console ved behov
MCP Identity
| Parameter | Verdi |
|---|---|
| MCP_NAME | OPAX-MCP |
| MCP_PROTOCOL | OPAX Protocol |
| MCP_SERVICE | opax-mcp-core |
| Cloud Run | us-central1 |
| HUB_URL | https://opax.vauco.no (Phase 2) |
Domenekonvensjon — LOCKED
| Subdomain | Type | Formål |
|---|---|---|
opax.vauco.no |
Hub/MCP | OPAX-MCP operator hub — kun Vauco internt |
<klient>-os.vauco.no |
Prod OS | Kundens live produksjonssystem |
<klient>-oss.vauco.no |
Stage OS | Kundens staging/demo-system (pre-go-live) |
-os= produksjon.-oss= staging. Aldri omvendt.
Auth-strategi
- Nå: Google OAuth på alle miljøer
- Senere: BankID på
-os-domener for kliniske kunder (Medioteq først)
Kjente / planlagte domener
| Domene | Status | Cloud Run-tjeneste |
|---|---|---|
opax.vauco.no |
DNS pending | opax-mcp-core ✅ |
medioteq-oss.vauco.no |
Planlagt | medioteq-oss-core |
medioteq-os.vauco.no |
Planlagt | medioteq-os-core |
Tråd-kontrakter
| Tråd | Modell | Eid fil | Konnektor |
|---|---|---|---|
| OPS-Computer-Hub | Perplexity Computer | AGENT_RULEBOOK.md, LEARNINGS.md, CHANGELOG.md | GitHub web, search |
| EXEC-Gemini-Workstation | Gemini 2.5 Pro | SYSTEM_STATE_AGENT_BOOT.md, TODO.md | shell, gcloud, git |
| PLAN-Claude-Architect | Claude Sonnet/Opus | ARCHITECTURE.md (denne), ROADMAP.md | GitHub UI/API |
| AUDIT-NMT-Critic | Nemotron | RISKREGISTER.md, PLANBOARD.md | read-only alle filer |
Regel: En fil — én eier. Andre tråder foreslår via PR, eier merger.
Dataflyt
OPS --[MD-patch PR]--> PLAN merger
EXEC --[SYSTEM_STATE]--> alle leser
PLAN --[arch-spec PR]--> EXEC implementerer
AUDIT--[PLANBOARD]--> Chris prioriterer
Filer flyter via: GitHub commits på OSVauco/main
Aldri: direkte tråd-til-tråd
Klientmodell — Prosjektisolasjon
Prinsipp: Hver klient med sensitiv/medisinsk data får eget GCP-prosjekt.
| Lag | GCP-prosjekt | Rolle |
|---|---|---|
| OPAX/OSVauco | propane-will-491900-m5 |
Management plane — deploy, monitor, orchestrate |
| Medioteq klinisk | <medioteq-project-id> |
Data plane — kliniske data forblir her |
- OPAX mottar aldri rå pasientdata
- Data-residency:
europe-north1for norske helsedata - Alle agentkall logges til BigQuery i klientprosjektet (input/output hash, ikke råinnhold)
Avhengigheter
OSVauco/mainer canonical branch for dette prosjektet- Cloud Run avhenger av
vauco-gemini-tui-bridgeog OPAX-MCP-tjenesten - Bridge avhenger av
bridge/-kode + GCP service accountvauco-dev - Token economy avhenger av
scripts/token_economy.py(planlagt) - Pre-commit hook v2 avhenger av
.githooks/pre-commit(planlagt) - OSV-pipeline avhenger av Ollama på
osvauco-dev-vmmed ≥16 GB RAM
Åpne spørsmål (max 3)
-
Langsiktig repo-struktur?
- A: Behold OSVauco + tjeneste-repoer (nåværende)
- B: Slå sammen til monorepo
- C: Splitt ytterligere per tjeneste
-
Cloud Run min-instances=0 eller 1?
- A: 0 — kutter idle-kostnad (anbefalt H0)
- B: 1 — ingen cold start latency
-
IAM: vauco-dev service account scope?
- A: Minimal — kun Vertex AI + Firestore
- B: Per-skill separate accounts
EOF v1.2 — 2026-06-29 — OPS (osvauco-dev-vm + OSV-pipeline + Emma-eskaleringspolitikk lagt til)
Dual-layer deploymodell: B over A
Vauco Architecture Manifest
Core Thesis
Vauco should use a dual-layer deployment model in which Model B governs Model A. The internal OPAX-MCP control plane should own policy, approvals, orchestration, and capability boundaries, while the simpler Cloud Build trigger pattern remains the standardized execution path for services that benefit from a repeatable and low-friction deployment flow.
This approach fits the documented profile model, where operator capabilities are broad and privileged while external, family, customer, or blueprint-driven profiles receive narrower and safer capabilities. It also matches the existing direction that OPAX-MCP is the external gateway and governance surface, while Gitea remains the source of truth for code.
Layer Model
Layer 1: Gitea as code truth
Gitea is the intended primary source of truth for code, and the repository flow has already been shifted toward Gitea-first operation rather than GitHub-first day-to-day work. This means repository state, branches, commits, and promotion events should originate from Gitea even when downstream deployment mechanisms differ by service type.
Layer 2: Model A as standard execution plane
Model A is the standardized deploy mechanism: a declarative Cloud Build pipeline builds the image, pushes it to Artifact Registry, renders a temporary service manifest from opax-mcp.yaml, and applies it with gcloud run services replace. This makes Model A suitable as the simple operational path for services that benefit from repeatable, low-drama deployment behavior.
Model A should therefore be treated as the default execution substrate for client-facing or repeatable services, especially where simplicity, supportability, and low operator friction matter more than deep custom orchestration. It is an execution mechanism, not the long-term brain of the platform.
Layer 3: Model B as internal control organ
Model B is the internal OPAX-MCP-owned control organ. Its documented architecture places webhook verification, approval handling, SMS-based HITL, and deploy orchestration inside OPAX-MCP itself, culminating in a controlled gcloud builds submit using the existing declarative pipeline file.
In this architecture, Model B does not need to replace Model A everywhere. Instead, Model B should govern when Model A is created, updated, permitted, or bypassed, allowing Vauco to keep a proprietary control plane without forcing every service into the same high-complexity execution path.
Design Principle
The key principle is complexity in the core, simplicity at the edge. OPAX-MCP should centralize policy, operator powers, capability boundaries, and architectural governance, while simpler deployment channels remain available for services that do not need bespoke orchestration.
This is operationally healthier than making every deployment path equally advanced. A system where every service must go through the most complex path increases cognitive load and failure analysis burden, while a system with no internal control organ increases fragmentation across triggers, consoles, and undocumented workflows.
Why This Fits Vauco
This dual model supports the existing profile strategy. The documented access model already distinguishes operator access from family, blueprint, and customer access, which implies a future where capabilities are governed centrally rather than exposed raw to every actor.
It also supports the documented learning-loop strategy: quality should come from verified events, clear diffs, explicit blockers, and curated operational truth instead of uncontrolled growth in tools or context. A layered model makes it easier to preserve that discipline because the control plane can observe, gate, and record the simpler execution plane rather than becoming identical to it.
Trade-offs
| Topic | Model A alone | Model B alone | Hybrid: B governs A |
|---|---|---|---|
| Speed to restore deploy flow | High | Low to medium | High for standard paths, medium for advanced paths |
| Proprietary control | Lower | Higher | High where it matters, moderate where standardization is enough |
| Cognitive load | Lower initially | Higher initially | Lower at edge, concentrated in operator core |
| Platform value over time | Moderate | High | High, with incremental rollout |
| Suitability for profiles/capabilities | Limited | Strong | Strong |
Governance Rules
To avoid dual-truth drift, the following architectural rules should hold:
-
Gitea is the source of truth for code and repository events.
-
opax-mcp.yaml is the deploy configuration truth for opax-mcp and should not be bypassed by ad-hoc imperative deploys.
-
OPAX-MCP is the policy and control truth for operator-level orchestration and future capability governance.
-
Model A is a standard execution mechanism used by policy, not an independent architectural authority.
Recommended Path
The recommended path is to restore and stabilize Model A first as the immediate execution layer, because the declarative pipeline file already exists and is the fastest way to stop manual deployment drift. After that, Model B should be built deliberately as the internal governance and orchestration layer that can manage or supersede Model A where appropriate.
This sequence preserves momentum without giving up the proprietary direction. It creates a platform in which Vauco can scale simple services through standardized execution while reserving deeper operator control for the places where control, policy, and identity matter most.