docs: session continuity file + claude-max-bridge /v1/responses fix

- Add RESUME.md for cross-session continuity (open items, key state, constraints)
- claude-max-bridge: implement /v1/responses (OpenAI Responses API) with
  previous_response_id chaining via server-side history injection into system prompt.
  Root cause of "session already in use": CLI leaves JSONL in un-resumable
  "dequeued" state after each --print run; fix avoids session reuse entirely.
  Also fixed: assistant content must be array-of-blocks not plain string (silent
  JS crash otherwise).
- LiteLLM: add pass_through_endpoints for /v1/responses → claude-max-bridge
- Storage, volumes, architecture docs reconciled (Vaultwarden → local zfs,
  Pocket-ID backup, WAL-G fix, apps/ decommission, Nextcloud CIFS→NFS)
- Add ideas/, proxmox-memory-audit.md, llm-benchmark.md (new docs this session)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
2026-05-22 22:42:17 +02:00
parent e0199042c1
commit 143343decf
12 changed files with 496 additions and 34 deletions
+33
View File
@@ -4,7 +4,40 @@ All notable infrastructure / service / doc changes. Newest first.
## 2026-05-22 ## 2026-05-22
- **Pocket-ID backup wired into Backrest**: CT 110 SQLite-only (no Postgres). Pre-backup hook script at `/opt/backrest/scripts/pocketid-prestage.sh` on CT 103: SSHs to CT 110, runs `pocket-id export` inside container, `docker cp`s ZIP to host, scp's ZIP + signing keys to `/mnt/pve/unas/services/pocketid-backup/`. `services-backup-plan` snapshots staging → jottacloud. Verified: hook fires before snapshot, snapshot includes Pocket-ID data.
- **Vaultwarden data → local zfs**: Moved `/mnt/pve/unas/services/vaultwarden` to `/opt/stacks/vaultwarden/data` (local NVMe). DB was already on CT 113 postgres; only `rsa_key.pem`, `config.json`, `attachments/` remain in the data dir. Stale `db.sqlite3` deleted. NFS copy at `services/vaultwarden/` is now orphaned and can be cleaned up.
- **WAL-G backup script fixed (CT 104)**: All nightly postgres backups were failing silently. Three bugs fixed: (1) `set -e` causing script abort when `shared-postgres` container not found (it was migrated to CT 113); (2) missing `-u postgres` on all `docker exec` calls — WAL-G requires the postgres role; (3) missing explicit data path `/var/lib/postgresql/data` on `backup-push`. `shared-postgres` commented out (CT 113 has its own WAL-G cron). `immich_postgres` and `lobe-postgres` now back up successfully to Garage S3.
- **apps/ directory decommissioned (CT 104)**: `/opt/stacks/apps/` contained dead duplicate compose files for gotify, karakeep, memos, paperless-ngx, traccar, vaultwarden, shared-db — all without `.env` files, never running. Renamed to `.DECOMMISSIONED-2026-05-22` following established pattern. The canonical stacks at `/opt/stacks/<service>/` are unaffected.
- **ComfyUI `--async-offload` removed**: Flag caused full CPU fallback on GGUF img2img with XPU (3.5 min/step instead of ~7 s/step) and pure-noise output. Removed from `ai/comfyui.yml`. Current args: `--listen 0.0.0.0 --enable-cors-header --use-pytorch-cross-attention --disable-smart-memory --force-fp16`. See `infra/proxmox-memory-audit.md`.
- **CT 105 Nextcloud CIFS → NFS**: Nextcloud was crash-looping ("Appdata is not present") because `unas_smb` CIFS mount (`//192.168.1.31/storage`) was not mounted on the host. Root cause: fstab entry was missing `_netdev`; mount dropped after a host event and was never re-established. Fixed by migrating CT 105 `mp0` from `/mnt/pve/unas_smb` to `/mnt/pve/unas` (same NFS share all other CTs use). CIFS entry removed from fstab. All 9 Nextcloud containers healthy post-restart.
- **arr stack started**: `vpn_gluetun`, `shelfarr` (`:13004`), `rdtclient` (via VPN, `:13001`), `prowlarr` (`:13002`), `audiobookshelf` (`:13003`). Fixed empty `labels:` in VPN compose.
- **claude-max-bridge**: new service on CT 104 (`ai-internal`, no host port). FastAPI shim wrapping `claude` CLI as an OpenAI-compatible `/v1/chat/completions` endpoint. Enables LiteLLM to route to Claude Max subscription at zero API cost. Key details: `--strict-mcp-config` (no MCP servers), `--input-format stream-json` (multi-turn), `--include-partial-messages` (real streaming deltas). Bind-mounts `/usr/local/bin/claude` + `/root/.claude` (rw for OAuth token refresh). Source at `git.nuclide.systems/fkrebs/claude-max-bridge`.
- **LiteLLM Claude models**: `claude-sonnet-4-6`, `claude-opus-4-7`, `claude-haiku-4-5` + short aliases (`sonnet`, `opus`, `haiku`) + effort variants (`claude-sonnet-4-6-high`, `claude-opus-4-7-high`) added via `openai/` provider pointing at `claude-max-bridge:8000/v1`. Priced at Anthropic list rates for cost-visibility — actual cost $0 (Max). Bridge features: temperature→effort mapping, `extra_body.effort` override, `extra_body.fallback_model`, `x_claude_cost_usd` and `x_claude_rate_limit` in responses.
- **claude-max-bridge v2**: additional CLI parameter mappings: `extra_body.system_prompt_mode` (`replace`/`append``--system-prompt`/`--append-system-prompt`), `extra_body.max_turns``--max-turns`, `extra_body.max_budget_usd``--max-budget-usd`, `response_format.json_schema``--json-schema`, `extra_body.exclude_dynamic_system_prompt_sections``--exclude-dynamic-system-prompt-sections`. Build context moved from `/tmp/` to `/opt/stacks/ai/claude-max-bridge/` (survives reboots). Source committed to Gitea at `a781abb`.
- **Shepard MongoDB pinned to 8.0.4**: `mongo:8.0` (latest) introduced a fatal kernel-version check in 8.0.5+ that rejects Proxmox's `7.0.2-5-pve` kernel string (parsed as major version 7 ≥ 6.19). Pinned to `mongo:8.0.4` in `/opt/shepard/infrastructure/docker-compose.yml`. Track MongoDB bug SERVER-121912 for a patched upstream release before unpinning.
- **ComfyUI XPU optimisation**: removed `--lowvram` + `--reserve-vram 1.0`; added `--async-offload` (Intel-patched XPU streams, ~10% speedup) + `--force-fp16`; memory limit 20 G → 24 G (FLUX peaks ~22 GB in shared XPU RAM). `--disable-smart-memory` kept since `get_free_memory()` returns the full 58 GB shared pool on Arc — smart eviction is blind to the cgroup limit. See `infra/proxmox-memory-audit.md`.
- **WaveSpeed FBC workflow**: `ApplyFBCacheOnModel(residual_diff_threshold=0.12)` + `FluxGuidance(guidance=3.5)` + `BasicGuider` + `SamplerCustomAdvanced` added to FLUX.1-schnell pipeline. Workflow saved at `/mnt/pve/unas/services/comfyui/input/flux-schnell-wavespeed.json`. Estimated 3050% additional speedup by caching first-block activations. WaveSpeed FBCache confirmed device-agnostic (no CUDA guards); `torch.compile` node must NOT be used on XPU.
- **LiteLLM benchmark**: `bench.py` at `/opt/stacks/ai/benchmark/` measures TTFT + TPS for all configured models using streaming requests; renders xychart diagrams via internal Kroki (`http://kroki:8000`). First run results: Mistral fastest TTFT (43 ms), Claude bridge competitive at 101143 ms. See `services/llm-benchmark.md`.
- **MCP gateway**: added `git`, `gitlab` (DLR, `--pass-environment` fix for mcp-proxy env passthrough bug), `paper-search` (Docker catalog image replaces inline `python:3.12-slim`); `UNPAYWALL_EMAIL` wired into paper-search. Removed duplicate `papersearch` container. Updated `services/mcp-gateway.md`. - **MCP gateway**: added `git`, `gitlab` (DLR, `--pass-environment` fix for mcp-proxy env passthrough bug), `paper-search` (Docker catalog image replaces inline `python:3.12-slim`); `UNPAYWALL_EMAIL` wired into paper-search. Removed duplicate `papersearch` container. Updated `services/mcp-gateway.md`.
- **Nextcloud MCP re-enabled**: `cbcoutinho/nextcloud-mcp-server` running in `single_user_basic` mode with `NEXTCLOUD_VERIFY_SSL=false`. App-password generated non-interactively via `occ user:add-app-password`. Gateway patched: (1) `NEXTCLOUD_VERIFY_SSL` / `MCP_DEPLOYMENT_MODE` now baked into hardcoded SERVERS env; (2) startup overlay extended to persist `enabled` flag; (3) proxy route bypasses per-user credential provisioning when `MCP_DEPLOYMENT_MODE=single_user_basic`. 26 servers healthy.
- **CT 109 renamed `observe``ops`**: updated `ct-inventory.md`, `homelab-architecture.md`, `proxmox-state.md` (IP changed to `.8` after conflict resolution).
- **CT 101 Shepard disk expanded**: 100 GiB → 500 GiB (ZFS online resize, no restart).
- **CT 112 Infisical marked live**: `ct-inventory.md` and `portmap.md` updated; LAN-only, no Zoraxy route.
- **litellm DB password rotated**: placeholder `litellm_password_here` replaced in both `ai/.env` and `litellm-config/config.yaml` (config.yaml takes precedence over env var). Container restarted, healthy.
- **UNAS migration dumps cleaned**: 7 × `*-migration-20260521.sql` (~492 MB) deleted from `/mnt/pve/unas/dump/`.
- **Arcane OIDC secret rotated**: new secret updated in Pocket-ID SQLite DB (bcrypt, cost 10) and `/opt/stacks/arcane/docker-compose.yml`. Container restarted.
- **n8n OIDC secret rotated**: Pocket-ID DB bcrypt update; secret applied to `/opt/stacks/n8n/docker-compose.yml` (authoritative stack). `/opt/stacks/apps/n8n/` was a dead duplicate (no `.env`, never ran) — compose renamed `.DECOMMISSIONED-2026-05-22`.
- **n8n encryption key rotated**: `N8N_ENCRYPTION_KEY` + `N8N_USER_MANAGEMENT_JWT_SECRET` replaced with 32-byte hex values in `/opt/stacks/n8n/.env` and `/opt/stacks/n8n/data/config`. Container restarted, healthy. Any previously stored workflow credentials are now invalid and must be re-entered.
- **Backrest repo passwords rotated**: `tapirnase` (weak placeholder) replaced on both `media-repo` and `services-repo` (JottaCloud). Used `restic key add` + `restic key remove`; `config.json` updated; daemon restarted.
- **Obsidian MCP (`mcpvault`)**: new server added. `bitbonsai/mcpvault` bridged via `mcp-proxy` stdio→HTTP in a custom `mcp-obsidian-bridge` image (Node 3.12-Alpine + mcp-proxy). Vault: Nextcloud `Notizen/` folder at `/mnt/pve/unas/services/nextcloud/fkrebs@nucli.de/files/Notizen`, bind-mounted read-write. 15 tools: read/write/patch/search/tags/frontmatter/stats. TTL cache 30s. Registered in Claude Code settings.json and LobeHub for both users.
- **LobeHub MCP sync**: all 27 gateway servers synced into `user_installed_plugins` for both users; stale `papersearch` entry removed; all manifests updated to long-lived `mcp_*` gateway token. Skipped: `crawl4ai` (no `/mcp` endpoint in `unclecode/crawl4ai:latest`), `n8n` (uses own auth).
- **CT 109 plan updated**: IP conflict noted (`.6` taken by CT 113); RAM 8 GiB, disk 50 GiB; Loki added; sshwifty added (multi-tab web SSH); Alloy replaces Promtail on all hosts; sidecar table expanded to CT 103/111/113. - **CT 109 plan updated**: IP conflict noted (`.6` taken by CT 113); RAM 8 GiB, disk 50 GiB; Loki added; sshwifty added (multi-tab web SSH); Alloy replaces Promtail on all hosts; sidecar table expanded to CT 103/111/113.
- **Portmap**: CT 113 (`db`, `.6`) section added; QNAP TS-251D (Klipper/Mainsail) section added. - **Portmap**: CT 113 (`db`, `.6`) section added; QNAP TS-251D (Klipper/Mainsail) section added.
- **New docs**: `services/backrest.md`, `services/databases.md`, `services/arcane.md`. - **New docs**: `services/backrest.md`, `services/databases.md`, `services/arcane.md`.
+11 -8
View File
@@ -4,6 +4,7 @@ Single source of truth for the homelab. Git-tracked since 2026-05-20.
## Read this first ## Read this first
- **[`/CLAUDE.md`](../CLAUDE.md)** (host root) — operator doctrine for any AI agent SSHing in. Skim before touching anything. - **[`/CLAUDE.md`](../CLAUDE.md)** (host root) — operator doctrine for any AI agent SSHing in. Skim before touching anything.
- **[`RESUME.md`](RESUME.md)** — open items + recent changes; update at end of every session.
- **[`CHANGELOG.md`](CHANGELOG.md)** — what changed and when. - **[`CHANGELOG.md`](CHANGELOG.md)** — what changed and when.
- **[`ct-inventory.md`](ct-inventory.md)** — canonical roster of LXCs / VM, IPs, role, sizing. - **[`ct-inventory.md`](ct-inventory.md)** — canonical roster of LXCs / VM, IPs, role, sizing.
@@ -26,15 +27,17 @@ Single source of truth for the homelab. Git-tracked since 2026-05-20.
│ ├── zoraxy.md reverse proxy │ ├── zoraxy.md reverse proxy
│ ├── adguard-dns.md DNS + rewrites │ ├── adguard-dns.md DNS + rewrites
│ ├── cloud-gpu.md GPU passthrough + future external GPU │ ├── cloud-gpu.md GPU passthrough + future external GPU
│ └── pocket-id.md [planned] OIDC IdP + client registry │ └── llm-benchmark.md LiteLLM model TTFT + TPS benchmark
├── stacks/ ← CT 104 docker stacks notes ├── infra/ (continued)
│ ├── PORTMAP.md canonical service → port → public hostname registry │ ├── portmap.md canonical service → port → public hostname registry
│ ├── CLAUDE.md agent breadcrumbs (CT 104 specific) │ ├── storage.md storage class per service (source of truth)
│ ├── storage.md [being rewritten] storage class per service │ ├── volumes.md volume bindings per stack
│ ├── volumes.md [being rewritten] volume bindings per stack │ ├── proxmox-memory-audit.md CT allocations, host budget, ComfyUI memory analysis
── docker-networks.md ── docker-networks.md
└── todo.md
├── stacks/ ← CT 104 agent notes + legacy stack docs
│ └── CLAUDE.md agent breadcrumbs (CT 104 specific)
├── history/ ← frozen / archived ├── history/ ← frozen / archived
│ ├── traefik-migration.md ABANDONED 2026-05-16 (Zoraxy chosen) │ ├── traefik-migration.md ABANDONED 2026-05-16 (Zoraxy chosen)
+58
View File
@@ -0,0 +1,58 @@
# Session Resume
Last updated: 2026-05-22. Update this at the end of every session.
## Open items (priority order)
### Backup gaps
- [x] ~~**Pocket-ID backup**~~ — done 2026-05-22. Pre-hook on CT 103 exports via `pocket-id export` + copies keys to UNAS staging; `services-backup-plan` → jottacloud. Verified snapshot includes Pocket-ID data.
- [ ] **Backrest media-repo**`video-projects-plan` and `media-backup-plan` were at 0 snapshots but both restic processes were actively running as of 2026-05-22 19:10. Verify snapshots landed after they complete.
- [x] ~~**Gitea repos**~~ — already in `services-backup-plan` → jottacloud. DB on CT 113 (WAL-G).
- [x] ~~**Coder workspace homes**~~ — already in `services-backup-plan` → jottacloud.
### Infrastructure
- [ ] **CT 109 "ops"** — planned LXC for Prometheus, Loki, Grafana, Alloy, Homarr, sshwifty, Gotify. Not yet built. See `infra/proxmox-state.md` §16.
- [ ] **Second NVMe** — needed before Postgres consolidation (CT 113) and rpool mirror. Unblocks moving remaining per-stack PG containers to CT 113.
- [x] ~~**Vaultwarden NFS cleanup**~~`/mnt/pve/unas/services/vaultwarden/` deleted 2026-05-22.
- [x] ~~**apps/backrest rename**~~ — done 2026-05-22 (renamed from PVE host via ZFS subvolume path; idmapped LXC blocked it from inside CT 104).
### Audits / housekeeping
- [ ] **Zoraxy audit** — CT 108 routes: WS headers, ACME coverage, decommissioned routes, auth consistency (Tinyauth gaps). See `services/zoraxy.md`.
- [ ] **Arcane OIDC secret** — rotation was done 2026-05-22 but Arcane has known issues with OIDC. Verify it still authenticates cleanly.
- [ ] **AdGuard split-horizon DNS** — some internal services still resolve to public IPs instead of LAN IPs for internal clients. Needs rewrite rules audit in `services/adguard-dns.md`.
## Recently completed (session continued, 2026-05-22)
- **claude-max-bridge `/v1/responses` endpoint**: implemented OpenAI Responses API with `previous_response_id` chaining. Root cause of "session already in use": Claude CLI leaves session JSONL files in an un-resumable "dequeued" state after each `--print` run. Fix: maintain conversation history server-side as formatted text; inject into system prompt for continuations; always use `--no-session-persistence`. Also fixed: assistant turns in `--input-format stream-json` require `content` as array-of-blocks not string (JS error otherwise). Assistant content array format also fixed in `_parse_messages` / `_parse_responses_input`.
- **LiteLLM pass-through for `/v1/responses`**: added `pass_through_endpoints` entry in `/opt/stacks/ai/litellm-config/config.yaml``http://claude-max-bridge:8000/v1/responses`. LiteLLM appends its token hash to the response `id`, so `previous_response_id` chaining only works when calling the bridge directly (not through LiteLLM).
## Recently completed (this session, 2026-05-22)
- **Pocket-ID backup**: pre-hook on CT 103 runs `pocket-id export` + scp keys → UNAS staging; `services-backup-plan` snapshots → jottacloud. Verified.
- **Stale NFS vaultwarden dir deleted**: `/mnt/pve/unas/services/vaultwarden/` removed.
- **Nextcloud CIFS → NFS**: crash-loop fixed; CT 105 mp0 migrated from `/mnt/pve/unas_smb` to `/mnt/pve/unas`; CIFS fstab entry removed.
- **Vaultwarden → local zfs**: `/mnt/pve/unas/services/vaultwarden``/opt/stacks/vaultwarden/data`; stale `db.sqlite3` deleted; DB already on CT 113 postgres.
- **WAL-G fixed**: `immich_postgres` and `lobe-postgres` backups now succeed. Fixed: missing `-u postgres`, missing data path, `set -e` abort, dead `shared-postgres` reference.
- **ComfyUI `--async-offload` removed**: caused CPU fallback + pure-noise output on GGUF img2img with XPU.
- **apps/ decommissioned**: dead duplicate compose files in `/opt/stacks/apps/` renamed `.DECOMMISSIONED-2026-05-22`.
- **Docs consolidated**: storage.md, volumes.md, homelab-architecture.md, proxmox-state.md, proxmox-memory-audit.md, ct-inventory.md, README.md all reconciled.
## Key system state
| Host | IP | Role |
|---|---|---|
| nuc (PVE) | 192.168.1.20 | Proxmox host — SSH gateway to all CTs |
| CT 104 docker | 192.168.1.40 | Main Docker host — AI/ML, media, ~65 containers |
| CT 113 db | 192.168.1.6 | Shared Postgres 17 + WAL-G → Garage S3 (cron 02:00) |
| CT 110 id | 192.168.1.5 | Pocket-ID OIDC IdP |
| CT 103 backrest | 192.168.1.3 | Backrest — jottacloud via rclone; UNAS at `/mnt/pve/unas` |
| CT 111 dev | 192.168.1.42 | Coder + Gitea |
| CT 108 zoraxy | 192.168.1.4 | Reverse proxy + ACME — **always confirm before changes** |
## Constraints to remember
- Zoraxy (CT 108): always get explicit user confirmation before any config change.
- Infisical (CT 112): LAN-only, no Zoraxy route — must not be internet-exposed.
- CT 104 has a separate agent reconciling compose files — rename decommissioned files (don't just stop), update `/opt/stacks/CLAUDE.md`.
- ComfyUI container limit: 24 G. Do not raise without measuring host impact.
- Pocket-ID is SQLite-only — no Postgres migration possible.
+5 -2
View File
@@ -9,17 +9,20 @@ Verified 2026-05-20 via `pct list`, `pct config <id>`, `qm config 100` on `nuc`.
| 102 | dns | LXC unpriv | 192.168.1.2/24 | 2 | 1024 | 4 GiB | **AdGuard Home** — LAN DNS resolver + filter | — | no | running | | 102 | dns | LXC unpriv | 192.168.1.2/24 | 2 | 1024 | 4 GiB | **AdGuard Home** — LAN DNS resolver + filter | — | no | running |
| 103 | backrest | LXC unpriv | 192.168.1.3/24 | 1 | 512 | 8 GiB | **Backrest** (restic) backup scheduler | mp0 NFS | no | running | | 103 | backrest | LXC unpriv | 192.168.1.3/24 | 1 | 512 | 8 GiB | **Backrest** (restic) backup scheduler | mp0 NFS | no | running |
| 104 | docker | LXC unpriv (idmapped) | 192.168.1.40/24 | 16 | 49152 | 200 GiB | Main Docker host — AI/ML + media + identity-adjacent (~65 containers) | mp0 NFS | Intel iGPU (card+render) | running | | 104 | docker | LXC unpriv (idmapped) | 192.168.1.40/24 | 16 | 49152 | 200 GiB | Main Docker host — AI/ML + media + identity-adjacent (~65 containers) | mp0 NFS | Intel iGPU (card+render) | running |
| 105 | nextcloud | LXC priv | 192.168.1.41/24 | 4 | 8196 | 100 GiB | Nextcloud AIO | mp0 **CIFS** (`//192.168.1.31/storage`) | Intel iGPU (render only) | running | | 105 | nextcloud | LXC priv | 192.168.1.41/24 | 4 | 8196 | 100 GiB | Nextcloud AIO | mp0 NFS | Intel iGPU (render only) | running |
| 108 | zoraxy | LXC unpriv | 192.168.1.4/24 | 2 | 2048 | 6 GiB | **Zoraxy** reverse proxy + ACME (`*.nuclide.systems`) | — | no | running | | 108 | zoraxy | LXC unpriv | 192.168.1.4/24 | 2 | 2048 | 6 GiB | **Zoraxy** reverse proxy + ACME (`*.nuclide.systems`) | — | no | running |
| 109 | ops | LXC | ⚠️ TBD (`.6`=CT 113, `.7`=CT 112) — likely `.8` | 4 | 8192 | 50 GiB | **Planned** — Prometheus, Loki, Grafana, Arcane Manager, Dozzle, Homarr, Diun, sshwifty, Tinyauth, socket-proxy; Alloy agents on all hosts | — | no | **planned** (see `infra/proxmox-state.md` §16) | | 109 | ops | LXC | ⚠️ TBD (`.6`=CT 113, `.7`=CT 112) — likely `.8` | 4 | 8192 | 50 GiB | **Planned** — Prometheus, Loki, Grafana, Arcane Manager, Dozzle, Homarr, Diun, sshwifty, Tinyauth, socket-proxy; Alloy agents on all hosts | — | no | **planned** (see `infra/proxmox-state.md` §16) |
| 110 | id | LXC unpriv | 192.168.1.5/24 | 1 | 1024 | 4 GiB | **Pocket-ID** OIDC IdP (`id.nuclide.systems`); migrated off CT 104 on 2026-05-20 | — | no | running | | 110 | id | LXC unpriv | 192.168.1.5/24 | 1 | 1024 | 4 GiB | **Pocket-ID** OIDC IdP (`id.nuclide.systems`); migrated off CT 104 on 2026-05-20 | — | no | running |
| 111 | dev | LXC unpriv | 192.168.1.42/24 | 12 | 32768 | 60 GiB | Dev host — **Coder + Gitea** + workspaces | mp0 NFS | Intel iGPU (render only; symlinked at `/dev/dri/renderD128`) | running | | 111 | dev | LXC unpriv | 192.168.1.42/24 | 12 | 32768 | 60 GiB | Dev host — **Coder + Gitea** + workspaces | mp0 NFS | Intel iGPU (render only; symlinked at `/dev/dri/renderD128`) | running |
| 112 | secrets | LXC unpriv | 192.168.1.7/24 | 2 | 4096 | 20 GiB | **Infisical** secrets manager (own Postgres 16 + Redis 7; replaces scattered .env files) — LAN-only, no Zoraxy route | — | no | running | | 112 | secrets | LXC unpriv | 192.168.1.7/24 | 2 | 4096 | 20 GiB | **Infisical** secrets manager (own Postgres 16 + Redis 7; replaces scattered .env files) — LAN-only, no Zoraxy route | — | no | running |
| 113 | db | LXC unpriv | 192.168.1.6/24 | 2 | 4096 | 40 GiB | Shared Postgres 17 + pgAdmin + WAL-G → Garage S3; Arcane edge agent | — | no | running |
UNAS NFS = `192.168.1.31:/var/nfs/shared/storage` over NFSv3. CT 105 (Nextcloud) is the lone outlier — it mounts the same UNAS share over **CIFS/SMB 3.1.1**, not NFS. UNAS NFS = `192.168.1.31:/var/nfs/shared/storage` over NFSv3. CT 105 (Nextcloud) is the lone outlier — it mounts the same UNAS share over **CIFS/SMB 3.1.1**, not NFS.
CT 112 ("secrets") — Infisical deployed 2026-05-22; see `services/secrets-manager.md`. LAN: `http://192.168.1.7:8200`. No Zoraxy route — secrets must not be internet-exposed. CT 112 ("secrets") — Infisical deployed 2026-05-22; see `services/secrets-manager.md`. LAN: `http://192.168.1.7:8200`. No Zoraxy route — secrets must not be internet-exposed.
CT 105 ("nextcloud") — migrated from CIFS (`//192.168.1.31/storage`) to NFS (`192.168.1.31:/var/nfs/shared/storage`) on 2026-05-22. CIFS mount was not remounting after host reboots, causing Nextcloud crash-loops. NFS is consistent with all other CTs.
## SSH access ## SSH access
The Proxmox host (`nuc`, `192.168.1.20`) has root SSH access to all LXC containers via key auth. To grant your own key access to every running container in one step, run on the host: The Proxmox host (`nuc`, `192.168.1.20`) has root SSH access to all LXC containers via key auth. To grant your own key access to every running container in one step, run on the host:
@@ -44,4 +47,4 @@ echo "ssh-ed25519 AAAA... you@yourmachine" >> ~/.ssh/authorized_keys
ssh-rsa AAAAB3NzaC1yc2EAAAADAQABAAACAQCuVqBW3VXg... root@nuc ssh-rsa AAAAB3NzaC1yc2EAAAADAQABAAACAQCuVqBW3VXg... root@nuc
``` ```
CT 113 ("db") is mentioned in memory notes (`postgres-consolidation`) as a future consolidation target. Not yet planned in any source-of-truth doc. CT 113 ("db") provisioned 2026-05-21; see `services/databases.md`. Future role: consolidate per-stack Postgres instances once second NVMe lands.
+190
View File
@@ -0,0 +1,190 @@
# Design: Claude Max Subscription → LiteLLM Gateway Bridge
**Status**: Proposal — not implemented
**Context**: LiteLLM at `https://ai.nuclide.systems` currently has no Anthropic models. Claude Max subscription (claude.ai) provides high-rate access to Sonnet 4.5/4.6 and Opus 4.7 but is decoupled from Anthropic API billing. This doc explores bridging the two.
## Problem
Claude Max and the Anthropic API are separate products with separate billing:
| | Claude Max | Anthropic API |
|---|---|---|
| Auth | OAuth / browser session | API key |
| Billing | Flat monthly subscription | Per-token |
| Rate limits | 5h rolling windows, model-specific | Tier-based RPM/TPM |
| Access | claude.ai web + Claude Code CLI | Any HTTP client |
The goal is to surface Max-subscription capacity through LiteLLM so that LobeHub, n8n, Coder workspaces, and other internal tools can call `claude-sonnet-4-6` at zero marginal cost and fall back to paid providers only when Max limits are hit.
## Approaches
### A — Session Cookie Reverse-Engineering (not recommended)
Several community projects (e.g. `claude-unofficial-api`) scrape the claude.ai WebSocket/HTTP protocol and expose an OpenAI-compatible endpoint. LiteLLM would point at this as a custom `openai/` provider.
**Pros**: Exposes the full web model lineup; streaming works.
**Cons**: Violates Anthropic ToS; breaks on any claude.ai front-end change; auth flow requires persisting browser cookies; no multimodal or tool-use parity guarantees.
**Verdict**: Avoid. Fragile and non-compliant.
---
### B — Claude Code CLI Bridge (recommended)
Claude Code CLI (`claude`) is already installed on CT 104 and authenticated with the Max subscription via `~/.claude/`. It ships a `--print` / `--output-format stream-json` mode designed for non-interactive use, and Anthropic explicitly supports programmatic use of the CLI.
A small **`claude-max-bridge`** service wraps this CLI as an OpenAI-compatible HTTP endpoint. LiteLLM registers it as a custom `openai/` base URL. No ToS issues — this is the supported surface.
```
LobeHub / n8n / Coder / Claude Code
LiteLLM Gateway (ai.nuclide.systems)
│ model: claude-sonnet-4-6 → openai/claude-sonnet-4-6
│ api_base: http://claude-max-bridge:8000
claude-max-bridge (new container, CT 104 ai-internal)
│ subprocess: claude --model ... --print --output-format stream-json
~/.claude/ (Max subscription session)
Anthropic (claude.ai)
```
**Pros**:
- Uses the officially supported programmatic interface
- Auth is already set up; no cookie management
- `claude` CLI handles retries, token limits, context window management
- Subprocess overhead is ~200400 ms cold; warm invocations faster
**Cons**:
- One subprocess per request — cannot multiplex a single session (unlike streaming HTTP)
- CLI is tied to the single authenticated user; no multi-user isolation
- Max rate limits apply per-account, same pool as interactive use
- Claude Code SDK (TypeScript) is cleaner but adds Node dependency
---
### C — Official API + Budget Cap (stop-gap)
Add Anthropic API key to LiteLLM with a hard budget cap (e.g. $20/month). Use it for Claude-specific features (tool use, long context) and let the existing free SAIA/Gemini fallback chain absorb general-purpose traffic.
**Pros**: Zero implementation work; full API feature parity.
**Cons**: Still costs money; no benefit from Max subscription.
**Verdict**: Valid fallback if B proves too complex, or as a complement for tool-heavy workloads that need the official API surface.
---
## Recommended Architecture (Option B)
### claude-max-bridge service
```
/opt/stacks/ai/claude-max-bridge/
Dockerfile
server.py # FastAPI, ~150 lines
docker-compose.yml (or entry in ai/docker-compose.yml)
```
**Dockerfile** — reuse the existing `claude` CLI install:
```dockerfile
FROM python:3.12-slim
RUN pip install fastapi uvicorn
# Mount ~/.claude from host; claude binary from host PATH or copied in
COPY server.py /app/server.py
CMD ["uvicorn", "app.server:app", "--host", "0.0.0.0", "--port", "8000"]
```
**server.py** — OpenAI `/v1/chat/completions` shim:
```python
# POST /v1/chat/completions
# Translates messages[] → claude --print --model ... --output-format stream-json
# Streams NDJSON lines back as SSE (text/event-stream)
# Maps finish_reason, usage tokens from CLI output headers
```
Key translation notes:
- `messages` → write to a temp file, pass via `--input-file` (avoids shell quoting issues)
- `stream: true` → parse `stream-json` lines, emit `data: {...}` SSE chunks
- `stream: false` → buffer all chunks, return single response object
- `max_tokens`, `temperature`, `system` → map to `--max-tokens`, `--temperature`, `--system`
- Tool use: **not supported** in first iteration; return 501 for requests with `tools`
- Model name passthrough: `claude-sonnet-4-6``--model claude-sonnet-4-6`
### LiteLLM config additions
```yaml
model_list:
- model_name: claude-sonnet-4-6
litellm_params:
model: openai/claude-sonnet-4-6
api_base: http://claude-max-bridge:8000
api_key: "dummy" # bridge ignores it; LiteLLM requires a value
stream_timeout: 120
timeout: 120
- model_name: claude-opus-4-7
litellm_params:
model: openai/claude-opus-4-7
api_base: http://claude-max-bridge:8000
api_key: "dummy"
stream_timeout: 180
timeout: 180
- model_name: claude-haiku-4-5
litellm_params:
model: openai/claude-haiku-4-5-20251001
api_base: http://claude-max-bridge:8000
api_key: "dummy"
stream_timeout: 60
timeout: 60
```
Add to `router_settings.fallbacks` — Max hits rate limit → fall to SAIA:
```yaml
fallbacks:
- claude-sonnet-4-6: [qwen3.5-397b-a17b]
- claude-opus-4-7: [qwen3.5-397b-a17b, mistral-large-latest]
- claude-haiku-4-5: [qwen3.5-122b-a10b, cerebras-llama-3.1-8b]
```
### Rate limit handling
The bridge should detect the CLI's rate-limit exit code / stderr message and return HTTP 429. LiteLLM's router will then trigger the fallback chain. No custom logic needed in LiteLLM itself.
Max limits as of 2026 (approximate, vary by plan tier):
- Sonnet 4.6: ~50 messages / 5-hour window at Pro; higher at Max
- Opus 4.7: ~1020 messages / 5-hour window
- Haiku 4.5: effectively unlimited under Max
## Volume estimate
Typical homelab workloads (LobeHub chat, n8n automations, Coder coding assist) generate maybe 200500 completions/day. Max's 5-hour windows reset 45× daily, so the practical limit is not usually hit unless Opus is used heavily.
## Risks
| Risk | Mitigation |
|---|---|
| CLI API changes between Claude Code releases | Pin the `claude` binary version; watch for breaking changes in release notes |
| Max rate limit shared with interactive use | Monitor via `claude usage`; bridge adds `X-Max-Usage` header from CLI output |
| Auth session expiry | Bridge returns 401 on auth failure; detect and alert via Gotify |
| Subprocess latency (cold start) | Keep a warm subprocess pool (12 persistent processes) using `--interactive` + message passing, or accept the 200400 ms overhead |
| No tool use support | Fallback chain routes tool-use requests to API providers (Gemini, Mistral) |
## Implementation plan
1. Write `server.py` shim (~150 lines) + Dockerfile
2. Build image on CT 104, add to `ai/docker-compose.yml`
3. Bind-mount `~/.claude` read-only into the container
4. Add model entries to `litellm-config/config.yaml`
5. Test streaming via `curl` against the bridge directly
6. Test end-to-end via LiteLLM → LobeHub
7. Watch `docker logs claude-max-bridge` for rate-limit / auth errors for 48h
## Open questions
- Does `claude --print` support all message roles (`system`, `user`, `assistant` turns)? Needs verification with multi-turn conversation format.
- Should the bridge be on CT 104 (access to `~/.claude`) or CT 109 ops (remote docker socket to CT 104)? CT 104 is simpler for v1.
- Is there an official Claude Code SDK for Python? (TypeScript SDK exists at `@anthropic-ai/claude-code`; Python wrapper may be cleaner than shelling out.)
+80
View File
@@ -0,0 +1,80 @@
# Proxmox Memory Audit
Last audited: 2026-05-22
## Host physical resources
| Resource | Total | Used (idle) | Available |
|----------|-------|-------------|-----------|
| RAM | 62 GiB | ~33 GiB | ~28 GiB |
| Swap | 31 GiB | 0 GiB | 31 GiB |
## CT memory allocations
| CT | Name | Allocated (MiB) | Swap (MiB) | Typical use | Notes |
|----|------|----------------|------------|-------------|-------|
| 101 | shepard | 32,768 | 8,192 | ~4 GB | Heavy stack: Mongo, Neo4j, TimescaleDB, Keycloak |
| 102 | dns | 1,024 | 512 | ~100 MB | AdGuard Home |
| 103 | backrest | 2,048 | 512 | ~200 MB | Restic scheduler |
| 104 | docker | **49,152** | 32,000 | **622 GB** | Main Docker host; FLUX spikes to ~22 GB |
| 105 | nextcloud | 8,196 | 8,196 | ~2 GB | Nextcloud AIO |
| 108 | zoraxy | 2,048 | 512 | ~300 MB | Reverse proxy |
| 110 | id | 1,024 | 512 | ~200 MB | Pocket-ID |
| 111 | dev | 32,768 | 8,192 | ~3 GB | Coder + Gitea workspaces |
| 112 | secrets | 4,096 | 512 | ~600 MB | Infisical |
| **Sum** | | **133,124 MiB (130 GiB)** | | | **2× overprovisioned vs physical RAM** |
## Key findings
### Overprovisioning is safe — until it isn't
Proxmox uses balloon drivers so CTs only consume what they actually use. At idle the host sits at ~33 GB used with 28 GB available. This is healthy. However, two specific CTs represent risk:
- **CT 104** (49 GiB limit) + **CT 111** (32 GiB limit) together could claim 81 GiB — well over the 62 GiB physical. If both hit peak simultaneously the host would start swapping heavily.
- CT 104's ComfyUI (FLUX generation) spikes from 6 GB baseline to ~22 GB. That 16 GB spike on top of the 33 GB idle baseline = 49 GB total, within physical RAM but tight.
### Why 40G Docker container limit destabilised the system
Setting ComfyUI's container limit to 40 G was the trigger. At generation time:
- FLUX model + activations: ~22 GB in the container
- Other ~65 Docker containers on CT 104: ~7 GB
- CT 104 OS + kernel: ~1 GB
- Other CTs idle: ~26 GB
- **Total: ~56 GB** → host started swapping, degrading all services
### ComfyUI XPU memory accounting gap
ComfyUI's `get_free_memory()` for Intel XPU queries `torch.xpu.get_device_properties().total_memory` = **58 GB** (the full shared memory pool — Arc shares system RAM). It has no awareness of the Docker cgroup limit. Smart memory management (`free_memory()`) calculates `memory_required - get_free_memory()` which is always hugely negative → never evicts models. This makes `--disable-smart-memory` irrelevant for XPU; smart memory is already broken.
Consequence: ComfyUI will always try to load the full model into XPU memory regardless of container limit. The cgroup OOM killer is the only backstop.
### Container limit recommendation for ComfyUI (CT 104)
| Scenario | Limit | Safe? |
|----------|-------|-------|
| `--lowvram` (original) | 20 G | ✅ Safe but slow (231 s/image) |
| No `--lowvram`, FLUX only | **24 G** | ✅ Fits FLUX peak (~22 GB) + 2 GB headroom |
| No `--lowvram` + img2img after FLUX | 26 G | ✅ FLUX stays resident, SD1.5 loads on top |
| 28 G | ⚠️ | Marginal — OOM triggered in testing |
| 40 G | ❌ | Destabilises host when generating |
Peak host usage at 24 G container limit during FLUX generation: `24 + 7 (other containers) + 26 (other CTs idle) ≈ 57 GB` — stays under 62 GB physical.
## Recommendations
1. **ComfyUI container limit**: set to **24 G** when running without `--lowvram`. Current revert to 20 G + `--lowvram` is stable but slower.
2. **CT 104 LXC allocation (49 GiB)**: appropriately sized given Docker workload, but is by far the largest single consumer. Do not raise further without measuring host impact.
3. **CT 111 (dev, 32 GiB)**: Coder workspaces could spike if users run heavy jobs. Consider adding a per-workspace memory limit in the Coder template.
4. **Watch list**: CT 101 (Shepard, 32 GiB) + CT 104 simultaneously at peak = 54 GB → host would need to swap. Unlikely in practice but possible during CI runs on CT 111 + FLUX generation on CT 104.
5. **Long-term**: when CT 109 (ops) is built, run Prometheus `node_exporter` on the PVE host and alert when host available RAM drops below 8 GiB.
## ComfyUI memory optimisation log
| Date | Change | Effect |
|------|--------|--------|
| 2026-05-22 | Removed `--lowvram` + raised limit to 28 G | OOM at 28 G (XPU DRM buffers counted against cgroup) |
| 2026-05-22 | Raised limit to 40 G | Host destabilised — reverted |
| 2026-05-22 | Reverted to `--lowvram` + 20 G | Stable, slow (231 s/image) |
| 2026-05-22 | Downloaded `t5-v1_1-xxl-encoder-Q4_K_S.gguf` (2.6 GB vs 3.2 GB Q5_K_M) | Saves 600 MB at load time |
| 2026-05-22 | Kept Q5_K_M T5 (Q4_K_S degrades prompt following per city96); removed `--lowvram`; set limit to 24 G; added `--async-offload --force-fp16` | ~730 s generation, safe within host memory budget |
| 2026-05-22 | Removed `--async-offload` | Flag incompatible with GGUF img2img on XPU — caused full CPU fallback (3.5 min/step) and pure-noise output. Removed; XPU generation now correct at ~7 s/step for txt2img. Current CLI: `--listen 0.0.0.0 --enable-cors-header --use-pytorch-cross-attention --disable-smart-memory --force-fp16` |
+3 -2
View File
@@ -5,7 +5,7 @@ Host: `nuc` · PVE `9.1.11` · Kernel `6.17.13-4-pve` · Debian 13 (trixie)
CPU: Intel Core Ultra 7 155H (16C / 22T, hybrid P+E+LP-E) · 1 socket · 1 NUMA CPU: Intel Core Ultra 7 155H (16C / 22T, hybrid P+E+LP-E) · 1 socket · 1 NUMA
RAM: 62 GiB physical · 31 GiB zram swap (50 % of RAM, zstd, prio 100) RAM: 62 GiB physical · 31 GiB zram swap (50 % of RAM, zstd, prio 100)
Storage: single Crucial P3 2 TB NVMe (QLC, DRAM-less) → `rpool` (ZFS, ashift=12, no redundancy) Storage: single Crucial P3 2 TB NVMe (QLC, DRAM-less) → `rpool` (ZFS, ashift=12, no redundancy)
Workload: 1 VM (HAOS) + 6 LXCs (Docker, Adguard, Backrest, Nextcloud, Zoraxy, Dockge) Workload: 1 VM (HAOS) + 9 LXCs (Docker, AdGuard, Backrest, Nextcloud, Zoraxy, Pocket-ID, Dev, Secrets, DB) + 1 planned (Ops/CT109)
--- ---
@@ -332,7 +332,7 @@ tail /var/log/unattended-upgrades/unattended-upgrades.log
## 10. Container-specific issues ## 10. Container-specific issues
### CT 104 (docker / AI / image-gen) — 128 GiB cap, 16 cores, GPU passthrough ### CT 104 (docker / AI / image-gen) — 48 GiB cap, 16 cores, GPU passthrough
Measured: 18.8 GiB current, peak 29.3 GiB, **9.3 GiB in swap**, load 8.0, ~65 Docker containers (Immich + ML, ComfyUI, LobeChat, n8n, Daytona, LiteLLM, Vaultwarden, Paperless+AI, many MCP servers, Speaches-OpenVINO). Measured: 18.8 GiB current, peak 29.3 GiB, **9.3 GiB in swap**, load 8.0, ~65 Docker containers (Immich + ML, ComfyUI, LobeChat, n8n, Daytona, LiteLLM, Vaultwarden, Paperless+AI, many MCP servers, Speaches-OpenVINO).
@@ -733,6 +733,7 @@ Single LXC holding everything monitoring/ops-shaped. Sizing target: **4 cores /
| Tinyauth | forward-auth gate for non-OIDC apps (Backrest, raw Dozzle, raw Prometheus, raw Grafana/Loki). OIDC client to Pocket-ID | | Tinyauth | forward-auth gate for non-OIDC apps (Backrest, raw Dozzle, raw Prometheus, raw Grafana/Loki). OIDC client to Pocket-ID |
| docker-socket-proxy | local + remote (CT 101/104/110) — hardened read-only docker.sock for Homarr/Arcane discovery | | docker-socket-proxy | local + remote (CT 101/104/110) — hardened read-only docker.sock for Homarr/Arcane discovery |
| **sshwifty** | web SSH client, multi-tab — multiple concurrent shells to different hosts (PVE, CT 104, CT 103, CT 111, etc.). LAN-only, port 8182. SSH key auth per host, no password prompt. Zoraxy + Tinyauth gate deferred. | | **sshwifty** | web SSH client, multi-tab — multiple concurrent shells to different hosts (PVE, CT 104, CT 103, CT 111, etc.). LAN-only, port 8182. SSH key auth per host, no password prompt. Zoraxy + Tinyauth gate deferred. |
| **docs-server** | mkdocs Material site (`/docs` git repo → static HTML); migrating here from CT 111 where it currently runs at port 13080. LAN-only, no Zoraxy route. |
### Sidecars deployed on each host ### Sidecars deployed on each host
+10 -11
View File
@@ -10,8 +10,7 @@
| Local zfs (rootfs) | NVMe in nuc | `/opt/stacks/<stack>/` on each LXC | Tier-1 hot state: databases, SQLite, secret stores, anything latency-sensitive | | Local zfs (rootfs) | NVMe in nuc | `/opt/stacks/<stack>/` on each LXC | Tier-1 hot state: databases, SQLite, secret stores, anything latency-sensitive |
| Local zfs (subvol) | NVMe in nuc | `rpool/data/subvol-<NNN>-disk-0` per LXC | Container rootfs, image cache | | Local zfs (subvol) | NVMe in nuc | `rpool/data/subvol-<NNN>-disk-0` per LXC | Container rootfs, image cache |
| Docker named volume | NVMe in nuc | `/var/lib/docker/volumes/...` on each Docker CT | Per-container persistent state managed by Docker | | Docker named volume | NVMe in nuc | `/var/lib/docker/volumes/...` on each Docker CT | Per-container persistent state managed by Docker |
| UNAS over **NFSv3** | `192.168.1.31` | `/mnt/pve/unas` on CTs 101, 103, 104, 111 | Bulk media, workspace home dirs, document storage | | UNAS over **NFSv3** | `192.168.1.31` | `/mnt/pve/unas` on CTs 101, 103, 104, 105, 111 | Bulk media, workspace home dirs, document storage |
| UNAS over **CIFS/SMB 3.1.1** | `192.168.1.31` | `/mnt/pve/unas` on CT 105 (Nextcloud) | Nextcloud-only path — earlier docs mis-labelled this as NFS |
## What lives where (verified 2026-05-21) ## What lives where (verified 2026-05-21)
@@ -24,7 +23,6 @@ Bulk + media + non-latency-sensitive app data.
| Immich (uploads, thumbs, derived) | 104 | `services/immich/{encoded-video,profile,thumbs}` + `media/images` + `backup/immich` | Immich own pg_dump → `media/images/db-dumps`; no off-host copy | | | Immich (uploads, thumbs, derived) | 104 | `services/immich/{encoded-video,profile,thumbs}` + `media/images` + `backup/immich` | Immich own pg_dump → `media/images/db-dumps`; no off-host copy | |
| Paperless-ngx (documents) | 104 | `media/documents/public/paperless-ngx/{consume,export,library}` | none | | | Paperless-ngx (documents) | 104 | `media/documents/public/paperless-ngx/{consume,export,library}` | none | |
| Paperless-AI | 104 | `services/paperless-ai``/app/data` | none | | | Paperless-AI | 104 | `services/paperless-ai``/app/data` | none | |
| Vaultwarden (attachments + SQLite DB) | 104 | `services/vaultwarden``/data` | none | tier-1 DB on NFS — risky, plan to move local |
| Traccar (logs, config) | 104 | `services/traccar/{logs,traccar.xml}` | none | data dir reverted to local `/opt/stacks/traccar/data` | | Traccar (logs, config) | 104 | `services/traccar/{logs,traccar.xml}` | none | data dir reverted to local `/opt/stacks/traccar/data` |
| Memos | 104 | `services/memos``/var/opt/memos` | none | | | Memos | 104 | `services/memos``/var/opt/memos` | none | |
| Arr-stack (Audiobookshelf, Prowlarr, RDTClient, Shelfarr) + media | 104 | `services/arr-stack/*` + `media/{audiobooks,ebooks,podcasts,Torrents}` | none | already migrated (older doc claimed "still local") | | Arr-stack (Audiobookshelf, Prowlarr, RDTClient, Shelfarr) + media | 104 | `services/arr-stack/*` + `media/{audiobooks,ebooks,podcasts,Torrents}` | none | already migrated (older doc claimed "still local") |
@@ -52,7 +50,8 @@ Tier-1 state and anything that should NOT be NFS-backed.
| SearXNG config | 104 | `/opt/stacks/ai/searxng` | none | | | SearXNG config | 104 | `/opt/stacks/ai/searxng` | none | |
| Flaresolverr | 104 | `/var/lib/flaresolver` | none | | | Flaresolverr | 104 | `/var/lib/flaresolver` | none | |
| AdGuard / Zoraxy / DNS / Shepard / Backrest binaries | 102/108/103/101 | local zfs only | none | Backrest itself has no self-backup | | AdGuard / Zoraxy / DNS / Shepard / Backrest binaries | 102/108/103/101 | local zfs only | none | Backrest itself has no self-backup |
| Nextcloud config + sidecars | 105 | local zfs (CT rootfs) | none | data path is CIFS — see below | | Vaultwarden (attachments + key material) | 104 | `/opt/stacks/vaultwarden/data``/data` | none | **moved off NFS 2026-05-22**; DB is on CT 113 postgres; stale `db.sqlite3` deleted |
| Nextcloud config + sidecars | 105 | local zfs (CT rootfs) | none | app data on NFS — see below |
### Garage S3 buckets (on local zfs, CT 104) ### Garage S3 buckets (on local zfs, CT 104)
@@ -76,9 +75,9 @@ Local on `/var/lib/docker/volumes/`. Mostly databases and caches.
| `coder-db` + `gitea-db` (docker volume `coder-db`) | / 111 | none | | `coder-db` + `gitea-db` (docker volume `coder-db`) | / 111 | none |
| (orphaned) `daytona-minimal_db_data` + runner anon vol | / 104 | none — clean up post-decommission | | (orphaned) `daytona-minimal_db_data` + runner anon vol | / 104 | none — clean up post-decommission |
### Nextcloud (CT 105) — the CIFS outlier ### Nextcloud (CT 105) — data on NFS
Nextcloud AIO mounts `//192.168.1.31/storage` over **CIFS/SMB 3.1.1**, not NFS like every other CT. Application data path is under that mount. The Postgres + Redis sidecars stay on local docker volumes. Nextcloud AIO mounts `/mnt/pve/unas/services/nextcloud` via NFSv3 (same share as all other CTs). Migrated from CIFS on 2026-05-22 after the CIFS mount caused a crash-loop. The Postgres + Redis sidecars stay on local docker volumes.
## Backup reality ## Backup reality
@@ -96,8 +95,8 @@ This is a known gap. Plans:
## Drift cleanup TODO ## Drift cleanup TODO
- [ ] Remove stale `services/n8n/database.sqlite` (408 MB, last write 2026-05-15) from UNAS. - [x] ~~Remove stale `services/n8n/database.sqlite` from UNAS~~ — gone (verified 2026-05-22)
- [ ] Remove stale `services/shared-db/garage/` copy from UNAS if present. - [x] ~~Remove stale `services/shared-db/garage/` copy from UNAS~~ — gone (verified 2026-05-22)
- [ ] Remove orphaned `daytona-minimal_db_data` Docker volume on CT 104. - [x] ~~Remove orphaned `daytona-minimal_db_data` Docker volume on CT 104~~ — gone (verified 2026-05-22)
- [ ] Vaultwarden `data/` on UNAS — plan to migrate to local zfs (latency + tier-1). - [x] ~~Vaultwarden `data/` on UNAS~~ — migrated to local zfs 2026-05-22. Stale NFS copy at `services/vaultwarden/` can be cleaned up.
- [ ] Add a Backrest plan for Pocket-ID DB and Gitea data → jottacloud. - [x] ~~Add a Backrest plan for Pocket-ID DB~~ — done 2026-05-22. Pre-backup hook runs `pocket-id export` + copies keys to UNAS staging; `services-backup-plan` snapshots staging → jottacloud nightly.
+11 -6
View File
@@ -46,9 +46,11 @@
### Pocket ID — `pocketid` ### Pocket ID — `pocketid`
| Type | Source | Container | UNAS | > **CT 110, not CT 104.** Pocket-ID migrated off CT 104 to its own LXC on 2026-05-20.
| Type | Source | Container | Notes |
|---|---|---|---| |---|---|---|---|
| **bind** | **`/mnt/pve/unas/services/pocketid`** | **`/app/data`** | **Already on UNAS** | | bind | `/opt/stacks/pocketid/data` | `/app/data` | Local zfs on CT 110 — no UNAS mount |
### Vaultwarden — `vaultwarden` ### Vaultwarden — `vaultwarden`
@@ -309,10 +311,12 @@ No DB-specific mounts needed (connects via env vars).
### Garage S3 — `garage` ### Garage S3 — `garage`
> Moved off NFS to local zfs on 2026-05-19 after a WAL-G outage. Stale copy at `services/shared-db/garage/` on UNAS may still exist — clean up.
| Type | Source | Container | Notes | | Type | Source | Container | Notes |
|---|---|---|---| |---|---|---|---|
| bind | `/mnt/pve/unas/services/shared-db/garage/data` | `/var/lib/garage/data` | S3 object data | | bind | `/opt/stacks/shared-db/garage/data` | `/var/lib/garage/data` | S3 object data — local NVMe |
| bind | `/mnt/pve/unas/services/shared-db/garage/meta` | `/var/lib/garage/meta` | S3 metadata (LMDB) | | bind | `/opt/stacks/shared-db/garage/meta` | `/var/lib/garage/meta` | S3 metadata (LMDB) — local NVMe |
--- ---
@@ -336,9 +340,10 @@ No DB-specific mounts needed (connects via env vars).
| Status | Count | Services | | Status | Count | Services |
|---|---|---| |---|---|---|
| ✅ Already on UNAS | 11 | gluetun, immich*(4), karakeep*(2), ntfy*(2), paperless-ai, pocketid, stremio, traccar*(3), vaultwarden, paperless-docs*(3), garage*(2) | | ✅ Already on UNAS | 9 | gluetun, immich*(4), karakeep*(2), ntfy*(2), paperless-ai, stremio, traccar*(3), vaultwarden, paperless-docs*(3) |
| ✅ Migrated (Phase 1) | 3 | memos, arcane, n8n | | ✅ Migrated (Phase 1) | 3 | memos, arcane, n8n |
| 🔷 Shared infrastructure | 2 | shared-postgres (local volume), garage (on UNAS) | | 🔷 Shared infrastructure (local) | 2 | shared-postgres (local volume), garage (local NVMe — moved off UNAS 2026-05-19) |
| 📋 Own CT, local only | 1 | pocketid (CT 110, no UNAS) |
| ⏳ Phase 2 planned (PG consolidation) | 5 | immich → shared-postgres, paperless → shared-postgres, daytona → shared-postgres, litellm → shared-postgres, lobehub → shared-postgres | | ⏳ Phase 2 planned (PG consolidation) | 5 | immich → shared-postgres, paperless → shared-postgres, daytona → shared-postgres, litellm → shared-postgres, lobehub → shared-postgres |
| ⏳ Phase 3 planned | 1 | arr-stack*(8 mounts) | | ⏳ Phase 3 planned | 1 | arr-stack*(8 mounts) |
| 📋 Keep local | ~5 | homepage, dozzle, litellm-config, searxng-config, saia-image-proxy | | 📋 Keep local | ~5 | homepage, dozzle, litellm-config, searxng-config, saia-image-proxy |
+4 -4
View File
@@ -19,7 +19,7 @@ across ~23 compose stacks. Principle: **self-host everything**, OIDC SSO,
(confirmed) — so ad/tracker blocking is network-wide despite DHCP not (confirmed) — so ad/tracker blocking is network-wide despite DHCP not
handing out AdGuard's IP directly. handing out AdGuard's IP directly.
- `103` lxc **backrest** — Restic/Backrest backups (1c/0.5G) - `103` lxc **backrest** — Restic/Backrest backups (1c/0.5G)
- `104` lxc **docker****the primary `/opt/stacks` host** (16c/32G/215G) → `.40` - `104` lxc **docker****the primary `/opt/stacks` host** (16c/48G/200G) → `.40`
- `105` lxc **nextcloud** — Nextcloud (4c/8G/107G) → `.41` - `105` lxc **nextcloud** — Nextcloud (4c/8G/107G) → `.41`
- `108` lxc **zoraxy** — reverse proxy (2c/2G) → `.4` - `108` lxc **zoraxy** — reverse proxy (2c/2G) → `.4`
- `110` lxc **id** — Pocket-ID OIDC IdP (1c/1G/4G) → `.5`; migrated off 104 on 2026-05-20 - `110` lxc **id** — Pocket-ID OIDC IdP (1c/1G/4G) → `.5`; migrated off 104 on 2026-05-20
@@ -392,14 +392,14 @@ flowchart TB
dns["CT 102 · dns<br/>(.2)<br/>AdGuard Home"] dns["CT 102 · dns<br/>(.2)<br/>AdGuard Home"]
backrest["CT 103 · backrest<br/>(.3)<br/>Backrest / restic"] backrest["CT 103 · backrest<br/>(.3)<br/>Backrest / restic"]
docker104["<b>CT 104 · docker</b><br/>(.40) · 16c/48G/200G<br/>~65 containers · Intel Arc passthrough"] docker104["<b>CT 104 · docker</b><br/>(.40) · 16c/48G/200G<br/>~65 containers · Intel Arc passthrough"]
nc["CT 105 · nextcloud<br/>(.41)<br/>Nextcloud AIO (CIFS to UNAS)"] nc["CT 105 · nextcloud<br/>(.41)<br/>Nextcloud AIO (NFS)"]
zoraxy["CT 108 · zoraxy<br/>(.4)<br/>reverse proxy + ACME"] zoraxy["CT 108 · zoraxy<br/>(.4)<br/>reverse proxy + ACME"]
obs(["CT 109 · ops<br/>(.8) · <i>planned</i><br/>Prom / Graf / Arcane / Homarr / Tinyauth"]) obs(["CT 109 · ops<br/>(.8) · <i>planned</i><br/>Prom / Graf / Arcane / Homarr / Tinyauth"])
id["CT 110 · id<br/>(.5)<br/>Pocket-ID (moved here 2026-05-20)"] id["CT 110 · id<br/>(.5)<br/>Pocket-ID (moved here 2026-05-20)"]
dev["CT 111 · dev<br/>(.42) · 12c/32G/60G<br/>Coder + Gitea + workspaces · Intel Arc"] dev["CT 111 · dev<br/>(.42) · 12c/32G/60G<br/>Coder + Gitea + workspaces · Intel Arc"]
db(["CT 113 · db<br/>(.?) · <i>planned</i><br/>shared Postgres + pgbouncer"]) db["CT 113 · db<br/>(.6) · provisioned 2026-05-21<br/>shared Postgres + pgAdmin"]
end end
UNAS[("UNAS<br/>192.168.1.31<br/>NFSv3 + CIFS")] UNAS[("UNAS<br/>192.168.1.31<br/>NFSv3")]
UDM[["UDM-SE · 192.168.1.1<br/>UniFi gateway · DNS → AdGuard"]] UDM[["UDM-SE · 192.168.1.1<br/>UniFi gateway · DNS → AdGuard"]]
Inet([Internet · ACME challenges · jottacloud · LiteLLM upstreams]) Inet([Internet · ACME challenges · jottacloud · LiteLLM upstreams])
+89
View File
@@ -0,0 +1,89 @@
# LLM Benchmark
Streaming performance benchmark for all LiteLLM-routed models. Measures **TTFT** (time to first token) and **TPS** (tokens per second) end-to-end from the LiteLLM proxy.
## Script
`/opt/stacks/ai/benchmark/bench.py` on CT 104.
## How to run
Must run inside a container on `ai-internal` to reach both LiteLLM and the internal Kroki instance:
```bash
docker run --rm --network ai-internal \
-v /opt/stacks/ai/benchmark/bench.py:/bench.py \
-v /opt/stacks/ai/benchmark:/charts \
-e LITELLM_API_KEY=sk-tapirnase \
-e KROKI_URL=http://kroki:8000 \
python:3.12-slim \
bash -c 'pip install httpx -q && python /bench.py --charts /charts'
```
Charts (`ttft.svg`, `tps.svg`) land in `/opt/stacks/ai/benchmark/`.
### Options
| Flag | Default | Description |
|------|---------|-------------|
| `--models a,b` | all 11 | Comma-separated subset to test |
| `--prompt` | `medium` | `ttft_short`, `medium`, or `code_medium` |
| `--runs N` | `2` | Runs per model (mean reported) |
| `--out PATH` | `/tmp/bench_results.json` | Raw JSON output |
| `--charts DIR` | `/tmp/bench_charts` | SVG chart output directory |
### Environment
| Var | Default | Notes |
|-----|---------|-------|
| `LITELLM_API_KEY` | — | Required (`sk-tapirnase`) |
| `LITELLM_BASE_URL` | `https://ai.nuclide.systems/v1` | Override for in-network: `http://litellm:4000` |
| `KROKI_URL` | `http://kroki:8000` | Internal Kroki for chart rendering |
## Models tested
| Model ID | Provider | Notes |
|----------|----------|-------|
| `claude-haiku-4-5` | claude-max-bridge | Via `claude` CLI subprocess |
| `claude-sonnet-4-6` | claude-max-bridge | |
| `claude-opus-4-7` | claude-max-bridge | |
| `gemini-2.5-flash-lite` | Google | |
| `gemini-2.5-flash` | Google | |
| `mistral-small-latest` | Mistral | |
| `mistral-large-latest` | Mistral | |
| `cerebras-llama-3.1-8b` | Cerebras | Hardware-accelerated inference |
| `cerebras-qwen-3-235b` | Cerebras | |
| `qwen3.5-122b-a10b` | — | Returns empty responses — likely extended-thinking mode with no visible content; excluded from charts |
| `deepseek-r1-distill-llama-70b` | DeepSeek | Reasoning model; total latency high due to thinking tokens |
## Methodology
- Each run sends a single streaming `POST /v1/chat/completions` request with `max_tokens: 512`
- **TTFT**: `perf_counter()` delta from request start to first SSE chunk containing `content`
- **TPS**: completion chars ÷ total elapsed seconds (char-based; divide by ~3.5 for true token TPS). Overridden by `usage.completion_tokens` if the provider returns it.
- Run 1 of each model includes cold-start overhead (subprocess fork for Claude bridge, connection setup for cloud APIs). The mean across all runs is reported in the summary.
- Models that return no `content` deltas (e.g. thinking-only responses) are marked `ERROR: no content tokens in response` and excluded from charts.
## Results — 2026-05-22
Prompt: `"Explain what TCP/IP is in 3 sentences."` · 2 runs each · warm LiteLLM cache
| Model | TTFT (ms) | TPS (chars/s) | Notes |
|-------|----------:|--------------|-------|
| **mistral-small-latest** | **43** | 10,111 | Fastest TTFT |
| **mistral-large-latest** | **58** | 11,042 | Fastest TPS |
| claude-opus-4-7 | 106 | 4,184 | |
| cerebras-qwen-3-235b | 118 | 3,777 | |
| claude-haiku-4-5 | 125 | 3,454 | |
| gemini-2.5-flash | 139 | 3,519 | |
| claude-sonnet-4-6 | 130 | 3,856 | |
| gemini-2.5-flash-lite | 202 | 2,396 | |
| deepseek-r1-distill-llama-70b | 297 | 10,561 | Total 5.4 s (thinking tokens) |
| cerebras-llama-3.1-8b | 155 | 2,947 | |
| qwen3.5-122b-a10b | ERROR | — | Empty response |
**Key observations:**
- Mistral leads on both TTFT and raw throughput — likely a combination of low-latency EU endpoints and LiteLLM response caching
- Claude models via the subprocess bridge perform well: 101143 ms TTFT is competitive with direct cloud APIs
- DeepSeek R1 has a fast TTFT (297 ms) but 5+ s total due to streaming reasoning tokens before the final answer
- TPS figures are char/s estimates; actual token/s ≈ TPS ÷ 3.5
+2 -1
View File
@@ -14,7 +14,7 @@ DinD MCP gateway on CT 104. Single OAuth-gated entry point fronting 24 MCP serve
## Server inventory (as of 2026-05-22) ## Server inventory (as of 2026-05-22)
26 servers configured in `mcp-gateway/config.json`. **Status: 26 healthy, 1 disabled (nextcloud re-enabled 2026-05-22).** 26 servers configured in `mcp-gateway/config.json`. **Status: 27 healthy.**
| Group | Server | Kind | Notes | | Group | Server | Kind | Notes |
|---|---|---|---| |---|---|---|---|
@@ -35,6 +35,7 @@ DinD MCP gateway on CT 104. Single OAuth-gated entry point fronting 24 MCP serve
| | wikipedia-mcp | catalog | | | | wikipedia-mcp | catalog | |
| | youtube-transcript | catalog | `mcp/youtube-transcript` | | | youtube-transcript | catalog | `mcp/youtube-transcript` |
| **personal** | bluesky | static | `ariel-mcp:8000` | | **personal** | bluesky | static | `ariel-mcp:8000` |
| | obsidian | spawn | `mcp-obsidian-bridge` + mcpvault; vault at Nextcloud `fkrebs@nucli.de/files/Notizen` (UNAS mount) |
| | gotify | spawn | | | | gotify | spawn | |
| | home-assistant | static | `http://192.168.1.60:9583/private_...` (HA add-on) | | | home-assistant | static | `http://192.168.1.60:9583/private_...` (HA add-on) |
| | immich | spawn | | | | immich | spawn | |