143343decf
- Add RESUME.md for cross-session continuity (open items, key state, constraints) - claude-max-bridge: implement /v1/responses (OpenAI Responses API) with previous_response_id chaining via server-side history injection into system prompt. Root cause of "session already in use": CLI leaves JSONL in un-resumable "dequeued" state after each --print run; fix avoids session reuse entirely. Also fixed: assistant content must be array-of-blocks not plain string (silent JS crash otherwise). - LiteLLM: add pass_through_endpoints for /v1/responses → claude-max-bridge - Storage, volumes, architecture docs reconciled (Vaultwarden → local zfs, Pocket-ID backup, WAL-G fix, apps/ decommission, Nextcloud CIFS→NFS) - Add ideas/, proxmox-memory-audit.md, llm-benchmark.md (new docs this session) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
5.0 KiB
5.0 KiB
Session Resume
Last updated: 2026-05-22. Update this at the end of every session.
Open items (priority order)
Backup gaps
Pocket-ID backup— done 2026-05-22. Pre-hook on CT 103 exports viapocket-id export+ copies keys to UNAS staging;services-backup-plan→ jottacloud. Verified snapshot includes Pocket-ID data.- Backrest media-repo —
video-projects-planandmedia-backup-planwere at 0 snapshots but both restic processes were actively running as of 2026-05-22 19:10. Verify snapshots landed after they complete. Gitea repos— already inservices-backup-plan→ jottacloud. DB on CT 113 (WAL-G).Coder workspace homes— already inservices-backup-plan→ jottacloud.
Infrastructure
- CT 109 "ops" — planned LXC for Prometheus, Loki, Grafana, Alloy, Homarr, sshwifty, Gotify. Not yet built. See
infra/proxmox-state.md§16. - Second NVMe — needed before Postgres consolidation (CT 113) and rpool mirror. Unblocks moving remaining per-stack PG containers to CT 113.
Vaultwarden NFS cleanup—/mnt/pve/unas/services/vaultwarden/deleted 2026-05-22.apps/backrest rename— done 2026-05-22 (renamed from PVE host via ZFS subvolume path; idmapped LXC blocked it from inside CT 104).
Audits / housekeeping
- Zoraxy audit — CT 108 routes: WS headers, ACME coverage, decommissioned routes, auth consistency (Tinyauth gaps). See
services/zoraxy.md. - Arcane OIDC secret — rotation was done 2026-05-22 but Arcane has known issues with OIDC. Verify it still authenticates cleanly.
- AdGuard split-horizon DNS — some internal services still resolve to public IPs instead of LAN IPs for internal clients. Needs rewrite rules audit in
services/adguard-dns.md.
Recently completed (session continued, 2026-05-22)
- claude-max-bridge
/v1/responsesendpoint: implemented OpenAI Responses API withprevious_response_idchaining. Root cause of "session already in use": Claude CLI leaves session JSONL files in an un-resumable "dequeued" state after each--printrun. Fix: maintain conversation history server-side as formatted text; inject into system prompt for continuations; always use--no-session-persistence. Also fixed: assistant turns in--input-format stream-jsonrequirecontentas array-of-blocks not string (JS error otherwise). Assistant content array format also fixed in_parse_messages/_parse_responses_input. - LiteLLM pass-through for
/v1/responses: addedpass_through_endpointsentry in/opt/stacks/ai/litellm-config/config.yaml→http://claude-max-bridge:8000/v1/responses. LiteLLM appends its token hash to the responseid, soprevious_response_idchaining only works when calling the bridge directly (not through LiteLLM).
Recently completed (this session, 2026-05-22)
- Pocket-ID backup: pre-hook on CT 103 runs
pocket-id export+ scp keys → UNAS staging;services-backup-plansnapshots → jottacloud. Verified. - Stale NFS vaultwarden dir deleted:
/mnt/pve/unas/services/vaultwarden/removed. - Nextcloud CIFS → NFS: crash-loop fixed; CT 105 mp0 migrated from
/mnt/pve/unas_smbto/mnt/pve/unas; CIFS fstab entry removed. - Vaultwarden → local zfs:
/mnt/pve/unas/services/vaultwarden→/opt/stacks/vaultwarden/data; staledb.sqlite3deleted; DB already on CT 113 postgres. - WAL-G fixed:
immich_postgresandlobe-postgresbackups now succeed. Fixed: missing-u postgres, missing data path,set -eabort, deadshared-postgresreference. - ComfyUI
--async-offloadremoved: caused CPU fallback + pure-noise output on GGUF img2img with XPU. - apps/ decommissioned: dead duplicate compose files in
/opt/stacks/apps/renamed.DECOMMISSIONED-2026-05-22. - Docs consolidated: storage.md, volumes.md, homelab-architecture.md, proxmox-state.md, proxmox-memory-audit.md, ct-inventory.md, README.md all reconciled.
Key system state
| Host | IP | Role |
|---|---|---|
| nuc (PVE) | 192.168.1.20 | Proxmox host — SSH gateway to all CTs |
| CT 104 docker | 192.168.1.40 | Main Docker host — AI/ML, media, ~65 containers |
| CT 113 db | 192.168.1.6 | Shared Postgres 17 + WAL-G → Garage S3 (cron 02:00) |
| CT 110 id | 192.168.1.5 | Pocket-ID OIDC IdP |
| CT 103 backrest | 192.168.1.3 | Backrest — jottacloud via rclone; UNAS at /mnt/pve/unas |
| CT 111 dev | 192.168.1.42 | Coder + Gitea |
| CT 108 zoraxy | 192.168.1.4 | Reverse proxy + ACME — always confirm before changes |
Constraints to remember
- Zoraxy (CT 108): always get explicit user confirmation before any config change.
- Infisical (CT 112): LAN-only, no Zoraxy route — must not be internet-exposed.
- CT 104 has a separate agent reconciling compose files — rename decommissioned files (don't just stop), update
/opt/stacks/CLAUDE.md. - ComfyUI container limit: 24 G. Do not raise without measuring host impact.
- Pocket-ID is SQLite-only — no Postgres migration possible.