# Session Resume Last updated: 2026-05-22. Update this at the end of every session. ## Open items (priority order) ### Backup gaps - [x] ~~**Pocket-ID backup**~~ — done 2026-05-22. Pre-hook on CT 103 exports via `pocket-id export` + copies keys to UNAS staging; `services-backup-plan` → jottacloud. Verified snapshot includes Pocket-ID data. - [ ] **Backrest media-repo** — `video-projects-plan` and `media-backup-plan` were at 0 snapshots but both restic processes were actively running as of 2026-05-22 19:10. Verify snapshots landed after they complete. - [x] ~~**Gitea repos**~~ — already in `services-backup-plan` → jottacloud. DB on CT 113 (WAL-G). - [x] ~~**Coder workspace homes**~~ — already in `services-backup-plan` → jottacloud. ### Infrastructure - [ ] **CT 109 "ops"** — planned LXC for Prometheus, Loki, Grafana, Alloy, Homarr, sshwifty, Gotify. Not yet built. See `infra/proxmox-state.md` §16. - [ ] **Second NVMe** — needed before Postgres consolidation (CT 113) and rpool mirror. Unblocks moving remaining per-stack PG containers to CT 113. - [x] ~~**Vaultwarden NFS cleanup**~~ — `/mnt/pve/unas/services/vaultwarden/` deleted 2026-05-22. - [x] ~~**apps/backrest rename**~~ — done 2026-05-22 (renamed from PVE host via ZFS subvolume path; idmapped LXC blocked it from inside CT 104). ### MCP / AI stack - [ ] **Consolidate syncstack → MCP gateway** — `syncstack.py` cron is fragile; move LiteLLM model polling + LobeChat model-list sync into `mcp-gateway/server.py` alongside the existing plugin sync loop. Delete syncstack container + cron after. - [ ] **crawl4ai SSE StreamConsumed** — intermittent 502s on first request after gateway restart (race with in-flight requests from old code). Investigate if still present after a clean restart cycle; likely harmless but worth confirming. ### Security / secrets - [ ] **Password rotation: `tapirnase`** — used as WiFi PSK, LiteLLM root key (`sk-tapirnase`), D-Link admin, and Backrest repo password. Single sniff = broad blast radius. Rotate all per-service. See `services/homelab-architecture.md` §operational rules. - [ ] **LiteLLM DB password** — placeholder `litellm_password_here` in `.env` + CT 113 user. Generate real password and update both. See `services/backrest.md` blind spot #6. - [ ] **Infisical secrets migration** — CT 112 provisioned but migration not started. Phase 1-4: move 85+ `.env` keys into central store. See `services/secrets-manager.md`. - [ ] **UniFi MFA JWT in homepage logs** — JWT found in `homepage/config/logs/homepage.log`. Rotate JWT + scrub log. See `services/homelab-architecture.md`. - [ ] **Vaultwarden OIDC SSO** — Pocket-ID client created 2026-05-21 but auth flow not wired up. Complete integration. - [ ] **Nextcloud Borg passphrase → Vaultwarden** — passphrase stored only in container env; CT 105 loss = unrecoverable backup. Move to Vaultwarden. See `services/backrest.md` blind spot #4b. ### Backup gaps (extended) - [ ] **UNAS personal data off-host** — ~820 GB Immich library + UNAS bulk not replicated off-host. Deploy rclone sync jobs: `media/images/library` → jottacloud:Photos + full UNAS → jottacloud:UNAS. See `services/backrest.md` phase 1a. - [ ] **Config-to-git** — daily cron to push zoraxy, adguard, pve, ha configs to `fkrebs/{zoraxy,adguard,pve,ha}-conf` Gitea repos. Repos exist, scripts not deployed. See `services/backrest.md` phase 3. - [ ] **Migration dump cleanup** — `/mnt/pve/unas/dump/*-migration-20260521.sql` (~492 MB, 2026-05-21). Decide keep or delete. See `services/backrest.md` blind spot #7. ### Database cleanup - [ ] **Drop stale `daytona` DB** — Daytona decommissioned 2026-05-20; DB still on shared-postgres (CT 113). Drop it. See `services/databases.md`. - [ ] **Remove old CT 111 Docker volumes** — `coder-db` + `gitea-db` Docker volumes on CT 111 may be stale after migration to CT 113. Verify and remove. See `services/databases.md`. ### Infrastructure (extended) - [ ] **ZFS ARC tuning** — `zfs_arc_max` artificially capped at 6.2 GiB; raise to 16 GiB once CT 101 memory is right-sized. See `infra/proxmox-state.md` §3. - [ ] **Docker disk reclaim** — ~13 GB reclaimable on CT 104 (8.3 GB unused images, 4.6 GB build cache). Run `docker system prune -a` after confirming no needed images. See `infra/proxmox-state.md` §10c. - [ ] **Dormant stacks audit** — ~10 stacks defined but not running on CT 104 (homepage, dozzle, arr-stack, qdrant, etc.). Document: parked vs. abandoned, or delete. See `infra/proxmox-state.md` §10c. ### Services / integrations - [ ] **Document ingestion n8n workflow** — Docling MCP deployed; n8n orchestration not built. Nextcloud/Paperless → Docling → LiteLLM embeddings → LobeChat knowledge base. See `services/doc-ingestion.md`. - [ ] **WAL-G monitoring** — silent stall went undetected for 13 h (failed_count=0). Add textfile collector + alert for hung archiver. See `services/homelab-architecture.md` roadmap. - [ ] **ComfyUI async queue** — current MCP tool blocks 90–300 s per call. Implement job-queue pattern: `generate_image()` returns job ID; `get_job_status()` polls. See `services/comfyui.md`. - [ ] **Arcane → CT 109 migration** — when CT 109 lands, move Arcane from CT 104; edge agents on CT 104/113 reconnect. Procedure documented in `services/arcane.md`. - [ ] **Additional MCP servers** — community MCPs exist for: Gitea, Paperless, Karakeep, Vaultwarden, Proxmox-VE, Audiobookshelf. Evaluate and wire into gateway. ### Audits / housekeeping - [ ] **Zoraxy audit** — CT 108 routes: WS headers, ACME coverage, decommissioned routes, auth consistency (Tinyauth gaps). See `services/zoraxy.md`. - [ ] **Arcane OIDC secret** — rotation was done 2026-05-22 but Arcane has known issues with OIDC. Verify it still authenticates cleanly. - [ ] **AdGuard split-horizon DNS** — some internal services still resolve to public IPs instead of LAN IPs for internal clients. Needs rewrite rules audit in `services/adguard-dns.md`. - [ ] **D-Link hardening** — admin UI reachable from any LAN device; SNMP community strings unverified; HTTP-only admin. Enable trusted-host allowlist, disable/harden SNMP, add TLS. See `services/homelab-architecture.md`. ## Recently completed (session continued ×2, 2026-05-22) - **crawl4ai MCP fixed**: was returning 404 on `/mcp` (streamable-HTTP). Current 0.8.6 image already has MCP over SSE at `/mcp/sse`. Added SSE transport support to `mcp-gateway/server.py`: per-request SSE round-trip with auto-initialize handshake; `transport: sse` in hardcoded SERVERS + config.json. Gateway health now shows crawl4ai ok; lobe-sync shows 27/27 servers. - **syncstack cron fixed**: missing `cd /opt/stacks/ai &&` prefix caused "no configuration file" every 15 min since ~May 18. Fixed. Manual run confirmed 60 models including 5 Claude models; LobeChat recreated. ## Recently completed (session continued, 2026-05-22) - **claude-max-bridge `/v1/responses` endpoint**: implemented OpenAI Responses API with `previous_response_id` chaining. Root cause of "session already in use": Claude CLI leaves session JSONL files in an un-resumable "dequeued" state after each `--print` run. Fix: maintain conversation history server-side as formatted text; inject into system prompt for continuations; always use `--no-session-persistence`. Also fixed: assistant turns in `--input-format stream-json` require `content` as array-of-blocks not string (JS error otherwise). Assistant content array format also fixed in `_parse_messages` / `_parse_responses_input`. - **LiteLLM pass-through for `/v1/responses`**: added `pass_through_endpoints` entry in `/opt/stacks/ai/litellm-config/config.yaml` → `http://claude-max-bridge:8000/v1/responses`. LiteLLM appends its token hash to the response `id`, so `previous_response_id` chaining only works when calling the bridge directly (not through LiteLLM). ## Recently completed (this session, 2026-05-22) - **Pocket-ID backup**: pre-hook on CT 103 runs `pocket-id export` + scp keys → UNAS staging; `services-backup-plan` snapshots → jottacloud. Verified. - **Stale NFS vaultwarden dir deleted**: `/mnt/pve/unas/services/vaultwarden/` removed. - **Nextcloud CIFS → NFS**: crash-loop fixed; CT 105 mp0 migrated from `/mnt/pve/unas_smb` to `/mnt/pve/unas`; CIFS fstab entry removed. - **Vaultwarden → local zfs**: `/mnt/pve/unas/services/vaultwarden` → `/opt/stacks/vaultwarden/data`; stale `db.sqlite3` deleted; DB already on CT 113 postgres. - **WAL-G fixed**: `immich_postgres` and `lobe-postgres` backups now succeed. Fixed: missing `-u postgres`, missing data path, `set -e` abort, dead `shared-postgres` reference. - **ComfyUI `--async-offload` removed**: caused CPU fallback + pure-noise output on GGUF img2img with XPU. - **apps/ decommissioned**: dead duplicate compose files in `/opt/stacks/apps/` renamed `.DECOMMISSIONED-2026-05-22`. - **Docs consolidated**: storage.md, volumes.md, homelab-architecture.md, proxmox-state.md, proxmox-memory-audit.md, ct-inventory.md, README.md all reconciled. ## Key system state | Host | IP | Role | |---|---|---| | nuc (PVE) | 192.168.1.20 | Proxmox host — SSH gateway to all CTs | | CT 104 docker | 192.168.1.40 | Main Docker host — AI/ML, media, ~65 containers | | CT 113 db | 192.168.1.6 | Shared Postgres 17 + WAL-G → Garage S3 (cron 02:00) | | CT 110 id | 192.168.1.5 | Pocket-ID OIDC IdP | | CT 103 backrest | 192.168.1.3 | Backrest — jottacloud via rclone; UNAS at `/mnt/pve/unas` | | CT 111 dev | 192.168.1.42 | Coder + Gitea | | CT 108 zoraxy | 192.168.1.4 | Reverse proxy + ACME — **always confirm before changes** | ## Constraints to remember - Zoraxy (CT 108): always get explicit user confirmation before any config change. - Infisical (CT 112): LAN-only, no Zoraxy route — must not be internet-exposed. - CT 104 has a separate agent reconciling compose files — rename decommissioned files (don't just stop), update `/opt/stacks/CLAUDE.md`. - ComfyUI container limit: 24 G. Do not raise without measuring host impact. - Pocket-ID is SQLite-only — no Postgres migration possible.