7.8 KiB
7.8 KiB
Session Resume
Last updated: 2026-05-22. Update this at the end of every session.
Open items (urgency order)
🔴 Critical — security risk or unrecoverable data loss
- Password rotation:
tapirnase— shared across WiFi PSK, LiteLLM root key (sk-tapirnase), D-Link admin, Backrest repo password. Single sniff = broad blast radius. Rotate per-service. Seeservices/homelab-architecture.md§operational rules. - UNAS personal data off-host — ~820 GB Immich library + UNAS bulk have zero off-host copy. Deploy rclone sync:
media/images/library→ jottacloud:Photos + full UNAS → jottacloud:UNAS. Seeservices/backrest.mdphase 1a. - Nextcloud Borg passphrase → Vaultwarden — passphrase only in container env; CT 105 loss = unrecoverable backup. Move to Vaultwarden. See
services/backrest.mdblind spot #4b.
🟠 High — known broken / verification needed
- Arcane OIDC secret — rotation done 2026-05-22 but Arcane has known OIDC issues. Verify it still authenticates.
- Backrest media-repo —
video-projects-planandmedia-backup-planwere at 0 snapshots as of 2026-05-22 19:10. Verify snapshots landed. - AdGuard split-horizon DNS — internal clients still resolve to public IPs. Rewrite rules audit. See
services/adguard-dns.md. - WAL-G monitoring — silent stall went undetected for 13 h. Add textfile collector + alert for hung archiver. See
services/homelab-architecture.mdroadmap. - Zoraxy audit — CT 108 routes: WS headers, ACME coverage, decommissioned routes, auth consistency. See
services/zoraxy.md.
🟡 Medium — incomplete migrations / cleanup debt
- Docker disk reclaim — ~17 GB reclaimable on CT 104 (10.3 GB unused images, 6.9 GB build cache).
docker system prune -aafter confirming no needed images. Seeinfra/proxmox-state.md§10c. - Dormant stacks audit — ~10 stacks defined but not running (homepage, dozzle, arr-stack, qdrant, etc.). Document: parked vs. abandoned. See
infra/proxmox-state.md§10c. - Config-to-git — push zoraxy, adguard, pve, ha configs daily to Gitea repos (repos exist, cron not deployed). See
services/backrest.mdphase 3. - D-Link hardening — admin UI open to full LAN, SNMP unverified, HTTP-only. Enable trusted-host allowlist, harden SNMP, add TLS. See
services/homelab-architecture.md. - Vaultwarden OIDC SSO — Pocket-ID client created 2026-05-21, auth flow not wired up.
🔵 Planned — requires infrastructure or significant effort
- CT 109 "ops" — LXC for Prometheus, Loki, Grafana, Alloy, Homarr, sshwifty, Gotify. Blocks monitoring stack + Arcane migration. See
infra/proxmox-state.md§16. - Second NVMe — needed before rpool mirror + Postgres consolidation (CT 113). Blocks several items.
- Infisical secrets migration — CT 112 provisioned; migration not started. Phase 1-4: move 85+
.envkeys. Seeservices/secrets-manager.md. - Consolidate syncstack → MCP gateway — move LiteLLM model polling + LobeChat sync into gateway's loop; delete syncstack cron.
- ZFS ARC tuning — raise
zfs_arc_maxfrom 6.2 GiB → 16 GiB once CT 101 memory right-sized. Seeinfra/proxmox-state.md§3. - Document ingestion n8n workflow — Docling MCP deployed; n8n orchestration not built. Nextcloud/Paperless → Docling → embeddings → LobeChat KB. See
services/doc-ingestion.md. - ComfyUI async queue — MCP tool blocks 90–300 s. Job-queue pattern:
generate_image()returns ID;get_job_status()polls. Seeservices/comfyui.md. - Additional MCP servers — community MCPs exist for Gitea, Paperless, Karakeep, Vaultwarden, Proxmox-VE, Audiobookshelf. Evaluate + wire in.
- Arcane → CT 109 migration — blocked by CT 109. Procedure documented in
services/arcane.md. crawl4ai SSE StreamConsumed— resolved 2026-05-22. Errors were from old code in log history; clean restart confirms no errors. Health: ok (7 tools), lobe-sync: 27/27 servers.
Recently completed (session continued ×2, 2026-05-22)
- crawl4ai MCP fixed: was returning 404 on
/mcp(streamable-HTTP). Current 0.8.6 image already has MCP over SSE at/mcp/sse. Added SSE transport support tomcp-gateway/server.py: per-request SSE round-trip with auto-initialize handshake;transport: ssein hardcoded SERVERS + config.json. Gateway health now shows crawl4ai ok; lobe-sync shows 27/27 servers. - syncstack cron fixed: missing
cd /opt/stacks/ai &&prefix caused "no configuration file" every 15 min since ~May 18. Fixed. Manual run confirmed 60 models including 5 Claude models; LobeChat recreated.
Recently completed (session continued, 2026-05-22)
- claude-max-bridge
/v1/responsesendpoint: implemented OpenAI Responses API withprevious_response_idchaining. Root cause of "session already in use": Claude CLI leaves session JSONL files in an un-resumable "dequeued" state after each--printrun. Fix: maintain conversation history server-side as formatted text; inject into system prompt for continuations; always use--no-session-persistence. Also fixed: assistant turns in--input-format stream-jsonrequirecontentas array-of-blocks not string (JS error otherwise). Assistant content array format also fixed in_parse_messages/_parse_responses_input. - LiteLLM pass-through for
/v1/responses: addedpass_through_endpointsentry in/opt/stacks/ai/litellm-config/config.yaml→http://claude-max-bridge:8000/v1/responses. LiteLLM appends its token hash to the responseid, soprevious_response_idchaining only works when calling the bridge directly (not through LiteLLM).
Recently completed (this session, 2026-05-22)
- Pocket-ID backup: pre-hook on CT 103 runs
pocket-id export+ scp keys → UNAS staging;services-backup-plansnapshots → jottacloud. Verified. - Stale NFS vaultwarden dir deleted:
/mnt/pve/unas/services/vaultwarden/removed. - Nextcloud CIFS → NFS: crash-loop fixed; CT 105 mp0 migrated from
/mnt/pve/unas_smbto/mnt/pve/unas; CIFS fstab entry removed. - Vaultwarden → local zfs:
/mnt/pve/unas/services/vaultwarden→/opt/stacks/vaultwarden/data; staledb.sqlite3deleted; DB already on CT 113 postgres. - WAL-G fixed:
immich_postgresandlobe-postgresbackups now succeed. Fixed: missing-u postgres, missing data path,set -eabort, deadshared-postgresreference. - ComfyUI
--async-offloadremoved: caused CPU fallback + pure-noise output on GGUF img2img with XPU. - apps/ decommissioned: dead duplicate compose files in
/opt/stacks/apps/renamed.DECOMMISSIONED-2026-05-22. - Docs consolidated: storage.md, volumes.md, homelab-architecture.md, proxmox-state.md, proxmox-memory-audit.md, ct-inventory.md, README.md all reconciled.
Key system state
| Host | IP | Role |
|---|---|---|
| nuc (PVE) | 192.168.1.20 | Proxmox host — SSH gateway to all CTs |
| CT 104 docker | 192.168.1.40 | Main Docker host — AI/ML, media, ~65 containers |
| CT 113 db | 192.168.1.6 | Shared Postgres 17 + WAL-G → Garage S3 (cron 02:00) |
| CT 110 id | 192.168.1.5 | Pocket-ID OIDC IdP |
| CT 103 backrest | 192.168.1.3 | Backrest — jottacloud via rclone; UNAS at /mnt/pve/unas |
| CT 111 dev | 192.168.1.42 | Coder + Gitea |
| CT 108 zoraxy | 192.168.1.4 | Reverse proxy + ACME — always confirm before changes |
Constraints to remember
- Zoraxy (CT 108): always get explicit user confirmation before any config change.
- Infisical (CT 112): LAN-only, no Zoraxy route — must not be internet-exposed.
- CT 104 has a separate agent reconciling compose files — rename decommissioned files (don't just stop), update
/opt/stacks/CLAUDE.md. - ComfyUI container limit: 24 G. Do not raise without measuring host impact.
- Pocket-ID is SQLite-only — no Postgres migration possible.