docs: session continuity file + claude-max-bridge /v1/responses fix
- Add RESUME.md for cross-session continuity (open items, key state, constraints) - claude-max-bridge: implement /v1/responses (OpenAI Responses API) with previous_response_id chaining via server-side history injection into system prompt. Root cause of "session already in use": CLI leaves JSONL in un-resumable "dequeued" state after each --print run; fix avoids session reuse entirely. Also fixed: assistant content must be array-of-blocks not plain string (silent JS crash otherwise). - LiteLLM: add pass_through_endpoints for /v1/responses → claude-max-bridge - Storage, volumes, architecture docs reconciled (Vaultwarden → local zfs, Pocket-ID backup, WAL-G fix, apps/ decommission, Nextcloud CIFS→NFS) - Add ideas/, proxmox-memory-audit.md, llm-benchmark.md (new docs this session) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -4,7 +4,40 @@ All notable infrastructure / service / doc changes. Newest first.
|
||||
|
||||
## 2026-05-22
|
||||
|
||||
- **Pocket-ID backup wired into Backrest**: CT 110 SQLite-only (no Postgres). Pre-backup hook script at `/opt/backrest/scripts/pocketid-prestage.sh` on CT 103: SSHs to CT 110, runs `pocket-id export` inside container, `docker cp`s ZIP to host, scp's ZIP + signing keys to `/mnt/pve/unas/services/pocketid-backup/`. `services-backup-plan` snapshots staging → jottacloud. Verified: hook fires before snapshot, snapshot includes Pocket-ID data.
|
||||
|
||||
- **Vaultwarden data → local zfs**: Moved `/mnt/pve/unas/services/vaultwarden` to `/opt/stacks/vaultwarden/data` (local NVMe). DB was already on CT 113 postgres; only `rsa_key.pem`, `config.json`, `attachments/` remain in the data dir. Stale `db.sqlite3` deleted. NFS copy at `services/vaultwarden/` is now orphaned and can be cleaned up.
|
||||
|
||||
- **WAL-G backup script fixed (CT 104)**: All nightly postgres backups were failing silently. Three bugs fixed: (1) `set -e` causing script abort when `shared-postgres` container not found (it was migrated to CT 113); (2) missing `-u postgres` on all `docker exec` calls — WAL-G requires the postgres role; (3) missing explicit data path `/var/lib/postgresql/data` on `backup-push`. `shared-postgres` commented out (CT 113 has its own WAL-G cron). `immich_postgres` and `lobe-postgres` now back up successfully to Garage S3.
|
||||
|
||||
- **apps/ directory decommissioned (CT 104)**: `/opt/stacks/apps/` contained dead duplicate compose files for gotify, karakeep, memos, paperless-ngx, traccar, vaultwarden, shared-db — all without `.env` files, never running. Renamed to `.DECOMMISSIONED-2026-05-22` following established pattern. The canonical stacks at `/opt/stacks/<service>/` are unaffected.
|
||||
|
||||
- **ComfyUI `--async-offload` removed**: Flag caused full CPU fallback on GGUF img2img with XPU (3.5 min/step instead of ~7 s/step) and pure-noise output. Removed from `ai/comfyui.yml`. Current args: `--listen 0.0.0.0 --enable-cors-header --use-pytorch-cross-attention --disable-smart-memory --force-fp16`. See `infra/proxmox-memory-audit.md`.
|
||||
|
||||
- **CT 105 Nextcloud CIFS → NFS**: Nextcloud was crash-looping ("Appdata is not present") because `unas_smb` CIFS mount (`//192.168.1.31/storage`) was not mounted on the host. Root cause: fstab entry was missing `_netdev`; mount dropped after a host event and was never re-established. Fixed by migrating CT 105 `mp0` from `/mnt/pve/unas_smb` to `/mnt/pve/unas` (same NFS share all other CTs use). CIFS entry removed from fstab. All 9 Nextcloud containers healthy post-restart.
|
||||
|
||||
- **arr stack started**: `vpn_gluetun`, `shelfarr` (`:13004`), `rdtclient` (via VPN, `:13001`), `prowlarr` (`:13002`), `audiobookshelf` (`:13003`). Fixed empty `labels:` in VPN compose.
|
||||
- **claude-max-bridge**: new service on CT 104 (`ai-internal`, no host port). FastAPI shim wrapping `claude` CLI as an OpenAI-compatible `/v1/chat/completions` endpoint. Enables LiteLLM to route to Claude Max subscription at zero API cost. Key details: `--strict-mcp-config` (no MCP servers), `--input-format stream-json` (multi-turn), `--include-partial-messages` (real streaming deltas). Bind-mounts `/usr/local/bin/claude` + `/root/.claude` (rw for OAuth token refresh). Source at `git.nuclide.systems/fkrebs/claude-max-bridge`.
|
||||
- **LiteLLM Claude models**: `claude-sonnet-4-6`, `claude-opus-4-7`, `claude-haiku-4-5` + short aliases (`sonnet`, `opus`, `haiku`) + effort variants (`claude-sonnet-4-6-high`, `claude-opus-4-7-high`) added via `openai/` provider pointing at `claude-max-bridge:8000/v1`. Priced at Anthropic list rates for cost-visibility — actual cost $0 (Max). Bridge features: temperature→effort mapping, `extra_body.effort` override, `extra_body.fallback_model`, `x_claude_cost_usd` and `x_claude_rate_limit` in responses.
|
||||
- **claude-max-bridge v2**: additional CLI parameter mappings: `extra_body.system_prompt_mode` (`replace`/`append` → `--system-prompt`/`--append-system-prompt`), `extra_body.max_turns` → `--max-turns`, `extra_body.max_budget_usd` → `--max-budget-usd`, `response_format.json_schema` → `--json-schema`, `extra_body.exclude_dynamic_system_prompt_sections` → `--exclude-dynamic-system-prompt-sections`. Build context moved from `/tmp/` to `/opt/stacks/ai/claude-max-bridge/` (survives reboots). Source committed to Gitea at `a781abb`.
|
||||
- **Shepard MongoDB pinned to 8.0.4**: `mongo:8.0` (latest) introduced a fatal kernel-version check in 8.0.5+ that rejects Proxmox's `7.0.2-5-pve` kernel string (parsed as major version 7 ≥ 6.19). Pinned to `mongo:8.0.4` in `/opt/shepard/infrastructure/docker-compose.yml`. Track MongoDB bug SERVER-121912 for a patched upstream release before unpinning.
|
||||
- **ComfyUI XPU optimisation**: removed `--lowvram` + `--reserve-vram 1.0`; added `--async-offload` (Intel-patched XPU streams, ~10% speedup) + `--force-fp16`; memory limit 20 G → 24 G (FLUX peaks ~22 GB in shared XPU RAM). `--disable-smart-memory` kept since `get_free_memory()` returns the full 58 GB shared pool on Arc — smart eviction is blind to the cgroup limit. See `infra/proxmox-memory-audit.md`.
|
||||
- **WaveSpeed FBC workflow**: `ApplyFBCacheOnModel(residual_diff_threshold=0.12)` + `FluxGuidance(guidance=3.5)` + `BasicGuider` + `SamplerCustomAdvanced` added to FLUX.1-schnell pipeline. Workflow saved at `/mnt/pve/unas/services/comfyui/input/flux-schnell-wavespeed.json`. Estimated 30–50% additional speedup by caching first-block activations. WaveSpeed FBCache confirmed device-agnostic (no CUDA guards); `torch.compile` node must NOT be used on XPU.
|
||||
- **LiteLLM benchmark**: `bench.py` at `/opt/stacks/ai/benchmark/` measures TTFT + TPS for all configured models using streaming requests; renders xychart diagrams via internal Kroki (`http://kroki:8000`). First run results: Mistral fastest TTFT (43 ms), Claude bridge competitive at 101–143 ms. See `services/llm-benchmark.md`.
|
||||
|
||||
- **MCP gateway**: added `git`, `gitlab` (DLR, `--pass-environment` fix for mcp-proxy env passthrough bug), `paper-search` (Docker catalog image replaces inline `python:3.12-slim`); `UNPAYWALL_EMAIL` wired into paper-search. Removed duplicate `papersearch` container. Updated `services/mcp-gateway.md`.
|
||||
- **Nextcloud MCP re-enabled**: `cbcoutinho/nextcloud-mcp-server` running in `single_user_basic` mode with `NEXTCLOUD_VERIFY_SSL=false`. App-password generated non-interactively via `occ user:add-app-password`. Gateway patched: (1) `NEXTCLOUD_VERIFY_SSL` / `MCP_DEPLOYMENT_MODE` now baked into hardcoded SERVERS env; (2) startup overlay extended to persist `enabled` flag; (3) proxy route bypasses per-user credential provisioning when `MCP_DEPLOYMENT_MODE=single_user_basic`. 26 servers healthy.
|
||||
- **CT 109 renamed `observe` → `ops`**: updated `ct-inventory.md`, `homelab-architecture.md`, `proxmox-state.md` (IP changed to `.8` after conflict resolution).
|
||||
- **CT 101 Shepard disk expanded**: 100 GiB → 500 GiB (ZFS online resize, no restart).
|
||||
- **CT 112 Infisical marked live**: `ct-inventory.md` and `portmap.md` updated; LAN-only, no Zoraxy route.
|
||||
- **litellm DB password rotated**: placeholder `litellm_password_here` replaced in both `ai/.env` and `litellm-config/config.yaml` (config.yaml takes precedence over env var). Container restarted, healthy.
|
||||
- **UNAS migration dumps cleaned**: 7 × `*-migration-20260521.sql` (~492 MB) deleted from `/mnt/pve/unas/dump/`.
|
||||
- **Arcane OIDC secret rotated**: new secret updated in Pocket-ID SQLite DB (bcrypt, cost 10) and `/opt/stacks/arcane/docker-compose.yml`. Container restarted.
|
||||
- **n8n OIDC secret rotated**: Pocket-ID DB bcrypt update; secret applied to `/opt/stacks/n8n/docker-compose.yml` (authoritative stack). `/opt/stacks/apps/n8n/` was a dead duplicate (no `.env`, never ran) — compose renamed `.DECOMMISSIONED-2026-05-22`.
|
||||
- **n8n encryption key rotated**: `N8N_ENCRYPTION_KEY` + `N8N_USER_MANAGEMENT_JWT_SECRET` replaced with 32-byte hex values in `/opt/stacks/n8n/.env` and `/opt/stacks/n8n/data/config`. Container restarted, healthy. Any previously stored workflow credentials are now invalid and must be re-entered.
|
||||
- **Backrest repo passwords rotated**: `tapirnase` (weak placeholder) replaced on both `media-repo` and `services-repo` (JottaCloud). Used `restic key add` + `restic key remove`; `config.json` updated; daemon restarted.
|
||||
- **Obsidian MCP (`mcpvault`)**: new server added. `bitbonsai/mcpvault` bridged via `mcp-proxy` stdio→HTTP in a custom `mcp-obsidian-bridge` image (Node 3.12-Alpine + mcp-proxy). Vault: Nextcloud `Notizen/` folder at `/mnt/pve/unas/services/nextcloud/fkrebs@nucli.de/files/Notizen`, bind-mounted read-write. 15 tools: read/write/patch/search/tags/frontmatter/stats. TTL cache 30s. Registered in Claude Code settings.json and LobeHub for both users.
|
||||
- **LobeHub MCP sync**: all 27 gateway servers synced into `user_installed_plugins` for both users; stale `papersearch` entry removed; all manifests updated to long-lived `mcp_*` gateway token. Skipped: `crawl4ai` (no `/mcp` endpoint in `unclecode/crawl4ai:latest`), `n8n` (uses own auth).
|
||||
- **CT 109 plan updated**: IP conflict noted (`.6` taken by CT 113); RAM 8 GiB, disk 50 GiB; Loki added; sshwifty added (multi-tab web SSH); Alloy replaces Promtail on all hosts; sidecar table expanded to CT 103/111/113.
|
||||
- **Portmap**: CT 113 (`db`, `.6`) section added; QNAP TS-251D (Klipper/Mainsail) section added.
|
||||
- **New docs**: `services/backrest.md`, `services/databases.md`, `services/arcane.md`.
|
||||
|
||||
Reference in New Issue
Block a user