Files
docs/RESUME.md
T

8.4 KiB
Raw Blame History

Session Resume

Last updated: 2026-05-22. Update this at the end of every session.

Open items (urgency order)

🔴 Critical — security risk or unrecoverable data loss

  • Password rotation: tapirnase — shared across WiFi PSK, LiteLLM root key (sk-tapirnase), D-Link admin, Backrest repo password. Single sniff = broad blast radius. Rotate per-service. See services/homelab-architecture.md §operational rules.
  • UniFi MFA JWT in homepage logs — JWT found in homepage/config/logs/homepage.log. Rotate + scrub. See services/homelab-architecture.md.
  • UNAS personal data off-host — ~820 GB Immich library + UNAS bulk have zero off-host copy. Deploy rclone sync: media/images/library → jottacloud:Photos + full UNAS → jottacloud:UNAS. See services/backrest.md phase 1a.
  • Nextcloud Borg passphrase → Vaultwarden — passphrase only in container env; CT 105 loss = unrecoverable backup. Move to Vaultwarden. See services/backrest.md blind spot #4b.
  • LiteLLM DB password — placeholder litellm_password_here in .env + CT 113 user. Generate real password; update both. See services/backrest.md blind spot #6.

🟠 High — known broken / verification needed

  • Arcane OIDC secret — rotation done 2026-05-22 but Arcane has known OIDC issues. Verify it still authenticates.
  • Backrest media-repovideo-projects-plan and media-backup-plan were at 0 snapshots as of 2026-05-22 19:10. Verify snapshots landed.
  • AdGuard split-horizon DNS — internal clients still resolve to public IPs. Rewrite rules audit. See services/adguard-dns.md.
  • WAL-G monitoring — silent stall went undetected for 13 h. Add textfile collector + alert for hung archiver. See services/homelab-architecture.md roadmap.
  • Zoraxy audit — CT 108 routes: WS headers, ACME coverage, decommissioned routes, auth consistency. See services/zoraxy.md.

🟡 Medium — incomplete migrations / cleanup debt

  • Drop stale daytona DB — Daytona decommissioned 2026-05-20; DB still on shared-postgres CT 113. Drop it.
  • Remove old CT 111 Docker volumescoder-db + gitea-db may be stale after migration to CT 113. Verify and remove.
  • Migration dump cleanup/mnt/pve/unas/dump/*-migration-20260521.sql (~492 MB). Keep or delete.
  • Docker disk reclaim — ~13 GB reclaimable on CT 104 (8.3 GB unused images, 4.6 GB build cache). docker system prune -a after confirming. See infra/proxmox-state.md §10c.
  • Dormant stacks audit — ~10 stacks defined but not running (homepage, dozzle, arr-stack, qdrant, etc.). Document: parked vs. abandoned. See infra/proxmox-state.md §10c.
  • Config-to-git — push zoraxy, adguard, pve, ha configs daily to Gitea repos (repos exist, cron not deployed). See services/backrest.md phase 3.
  • D-Link hardening — admin UI open to full LAN, SNMP unverified, HTTP-only. Enable trusted-host allowlist, harden SNMP, add TLS. See services/homelab-architecture.md.
  • Vaultwarden OIDC SSO — Pocket-ID client created 2026-05-21, auth flow not wired up.

🔵 Planned — requires infrastructure or significant effort

  • CT 109 "ops" — LXC for Prometheus, Loki, Grafana, Alloy, Homarr, sshwifty, Gotify. Blocks monitoring stack + Arcane migration. See infra/proxmox-state.md §16.
  • Second NVMe — needed before rpool mirror + Postgres consolidation (CT 113). Blocks several items.
  • Infisical secrets migration — CT 112 provisioned; migration not started. Phase 1-4: move 85+ .env keys. See services/secrets-manager.md.
  • Consolidate syncstack → MCP gateway — move LiteLLM model polling + LobeChat sync into gateway's loop; delete syncstack cron.
  • ZFS ARC tuning — raise zfs_arc_max from 6.2 GiB → 16 GiB once CT 101 memory right-sized. See infra/proxmox-state.md §3.
  • Document ingestion n8n workflow — Docling MCP deployed; n8n orchestration not built. Nextcloud/Paperless → Docling → embeddings → LobeChat KB. See services/doc-ingestion.md.
  • ComfyUI async queue — MCP tool blocks 90300 s. Job-queue pattern: generate_image() returns ID; get_job_status() polls. See services/comfyui.md.
  • Additional MCP servers — community MCPs exist for Gitea, Paperless, Karakeep, Vaultwarden, Proxmox-VE, Audiobookshelf. Evaluate + wire in.
  • Arcane → CT 109 migration — blocked by CT 109. Procedure documented in services/arcane.md.
  • crawl4ai SSE StreamConsumed — intermittent 502s on gateway restart. Confirm resolved after clean restart cycle.

Recently completed (session continued ×2, 2026-05-22)

  • crawl4ai MCP fixed: was returning 404 on /mcp (streamable-HTTP). Current 0.8.6 image already has MCP over SSE at /mcp/sse. Added SSE transport support to mcp-gateway/server.py: per-request SSE round-trip with auto-initialize handshake; transport: sse in hardcoded SERVERS + config.json. Gateway health now shows crawl4ai ok; lobe-sync shows 27/27 servers.
  • syncstack cron fixed: missing cd /opt/stacks/ai && prefix caused "no configuration file" every 15 min since ~May 18. Fixed. Manual run confirmed 60 models including 5 Claude models; LobeChat recreated.

Recently completed (session continued, 2026-05-22)

  • claude-max-bridge /v1/responses endpoint: implemented OpenAI Responses API with previous_response_id chaining. Root cause of "session already in use": Claude CLI leaves session JSONL files in an un-resumable "dequeued" state after each --print run. Fix: maintain conversation history server-side as formatted text; inject into system prompt for continuations; always use --no-session-persistence. Also fixed: assistant turns in --input-format stream-json require content as array-of-blocks not string (JS error otherwise). Assistant content array format also fixed in _parse_messages / _parse_responses_input.
  • LiteLLM pass-through for /v1/responses: added pass_through_endpoints entry in /opt/stacks/ai/litellm-config/config.yamlhttp://claude-max-bridge:8000/v1/responses. LiteLLM appends its token hash to the response id, so previous_response_id chaining only works when calling the bridge directly (not through LiteLLM).

Recently completed (this session, 2026-05-22)

  • Pocket-ID backup: pre-hook on CT 103 runs pocket-id export + scp keys → UNAS staging; services-backup-plan snapshots → jottacloud. Verified.
  • Stale NFS vaultwarden dir deleted: /mnt/pve/unas/services/vaultwarden/ removed.
  • Nextcloud CIFS → NFS: crash-loop fixed; CT 105 mp0 migrated from /mnt/pve/unas_smb to /mnt/pve/unas; CIFS fstab entry removed.
  • Vaultwarden → local zfs: /mnt/pve/unas/services/vaultwarden/opt/stacks/vaultwarden/data; stale db.sqlite3 deleted; DB already on CT 113 postgres.
  • WAL-G fixed: immich_postgres and lobe-postgres backups now succeed. Fixed: missing -u postgres, missing data path, set -e abort, dead shared-postgres reference.
  • ComfyUI --async-offload removed: caused CPU fallback + pure-noise output on GGUF img2img with XPU.
  • apps/ decommissioned: dead duplicate compose files in /opt/stacks/apps/ renamed .DECOMMISSIONED-2026-05-22.
  • Docs consolidated: storage.md, volumes.md, homelab-architecture.md, proxmox-state.md, proxmox-memory-audit.md, ct-inventory.md, README.md all reconciled.

Key system state

Host IP Role
nuc (PVE) 192.168.1.20 Proxmox host — SSH gateway to all CTs
CT 104 docker 192.168.1.40 Main Docker host — AI/ML, media, ~65 containers
CT 113 db 192.168.1.6 Shared Postgres 17 + WAL-G → Garage S3 (cron 02:00)
CT 110 id 192.168.1.5 Pocket-ID OIDC IdP
CT 103 backrest 192.168.1.3 Backrest — jottacloud via rclone; UNAS at /mnt/pve/unas
CT 111 dev 192.168.1.42 Coder + Gitea
CT 108 zoraxy 192.168.1.4 Reverse proxy + ACME — always confirm before changes

Constraints to remember

  • Zoraxy (CT 108): always get explicit user confirmation before any config change.
  • Infisical (CT 112): LAN-only, no Zoraxy route — must not be internet-exposed.
  • CT 104 has a separate agent reconciling compose files — rename decommissioned files (don't just stop), update /opt/stacks/CLAUDE.md.
  • ComfyUI container limit: 24 G. Do not raise without measuring host impact.
  • Pocket-ID is SQLite-only — no Postgres migration possible.