docs: session continuity file + claude-max-bridge /v1/responses fix
- Add RESUME.md for cross-session continuity (open items, key state, constraints) - claude-max-bridge: implement /v1/responses (OpenAI Responses API) with previous_response_id chaining via server-side history injection into system prompt. Root cause of "session already in use": CLI leaves JSONL in un-resumable "dequeued" state after each --print run; fix avoids session reuse entirely. Also fixed: assistant content must be array-of-blocks not plain string (silent JS crash otherwise). - LiteLLM: add pass_through_endpoints for /v1/responses → claude-max-bridge - Storage, volumes, architecture docs reconciled (Vaultwarden → local zfs, Pocket-ID backup, WAL-G fix, apps/ decommission, Nextcloud CIFS→NFS) - Add ideas/, proxmox-memory-audit.md, llm-benchmark.md (new docs this session) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,80 @@
|
||||
# Proxmox Memory Audit
|
||||
|
||||
Last audited: 2026-05-22
|
||||
|
||||
## Host physical resources
|
||||
|
||||
| Resource | Total | Used (idle) | Available |
|
||||
|----------|-------|-------------|-----------|
|
||||
| RAM | 62 GiB | ~33 GiB | ~28 GiB |
|
||||
| Swap | 31 GiB | 0 GiB | 31 GiB |
|
||||
|
||||
## CT memory allocations
|
||||
|
||||
| CT | Name | Allocated (MiB) | Swap (MiB) | Typical use | Notes |
|
||||
|----|------|----------------|------------|-------------|-------|
|
||||
| 101 | shepard | 32,768 | 8,192 | ~4 GB | Heavy stack: Mongo, Neo4j, TimescaleDB, Keycloak |
|
||||
| 102 | dns | 1,024 | 512 | ~100 MB | AdGuard Home |
|
||||
| 103 | backrest | 2,048 | 512 | ~200 MB | Restic scheduler |
|
||||
| 104 | docker | **49,152** | 32,000 | **6–22 GB** | Main Docker host; FLUX spikes to ~22 GB |
|
||||
| 105 | nextcloud | 8,196 | 8,196 | ~2 GB | Nextcloud AIO |
|
||||
| 108 | zoraxy | 2,048 | 512 | ~300 MB | Reverse proxy |
|
||||
| 110 | id | 1,024 | 512 | ~200 MB | Pocket-ID |
|
||||
| 111 | dev | 32,768 | 8,192 | ~3 GB | Coder + Gitea workspaces |
|
||||
| 112 | secrets | 4,096 | 512 | ~600 MB | Infisical |
|
||||
| **Sum** | | **133,124 MiB (130 GiB)** | | | **2× overprovisioned vs physical RAM** |
|
||||
|
||||
## Key findings
|
||||
|
||||
### Overprovisioning is safe — until it isn't
|
||||
|
||||
Proxmox uses balloon drivers so CTs only consume what they actually use. At idle the host sits at ~33 GB used with 28 GB available. This is healthy. However, two specific CTs represent risk:
|
||||
|
||||
- **CT 104** (49 GiB limit) + **CT 111** (32 GiB limit) together could claim 81 GiB — well over the 62 GiB physical. If both hit peak simultaneously the host would start swapping heavily.
|
||||
- CT 104's ComfyUI (FLUX generation) spikes from 6 GB baseline to ~22 GB. That 16 GB spike on top of the 33 GB idle baseline = 49 GB total, within physical RAM but tight.
|
||||
|
||||
### Why 40G Docker container limit destabilised the system
|
||||
|
||||
Setting ComfyUI's container limit to 40 G was the trigger. At generation time:
|
||||
- FLUX model + activations: ~22 GB in the container
|
||||
- Other ~65 Docker containers on CT 104: ~7 GB
|
||||
- CT 104 OS + kernel: ~1 GB
|
||||
- Other CTs idle: ~26 GB
|
||||
- **Total: ~56 GB** → host started swapping, degrading all services
|
||||
|
||||
### ComfyUI XPU memory accounting gap
|
||||
|
||||
ComfyUI's `get_free_memory()` for Intel XPU queries `torch.xpu.get_device_properties().total_memory` = **58 GB** (the full shared memory pool — Arc shares system RAM). It has no awareness of the Docker cgroup limit. Smart memory management (`free_memory()`) calculates `memory_required - get_free_memory()` which is always hugely negative → never evicts models. This makes `--disable-smart-memory` irrelevant for XPU; smart memory is already broken.
|
||||
|
||||
Consequence: ComfyUI will always try to load the full model into XPU memory regardless of container limit. The cgroup OOM killer is the only backstop.
|
||||
|
||||
### Container limit recommendation for ComfyUI (CT 104)
|
||||
|
||||
| Scenario | Limit | Safe? |
|
||||
|----------|-------|-------|
|
||||
| `--lowvram` (original) | 20 G | ✅ Safe but slow (231 s/image) |
|
||||
| No `--lowvram`, FLUX only | **24 G** | ✅ Fits FLUX peak (~22 GB) + 2 GB headroom |
|
||||
| No `--lowvram` + img2img after FLUX | 26 G | ✅ FLUX stays resident, SD1.5 loads on top |
|
||||
| 28 G | ⚠️ | Marginal — OOM triggered in testing |
|
||||
| 40 G | ❌ | Destabilises host when generating |
|
||||
|
||||
Peak host usage at 24 G container limit during FLUX generation: `24 + 7 (other containers) + 26 (other CTs idle) ≈ 57 GB` — stays under 62 GB physical.
|
||||
|
||||
## Recommendations
|
||||
|
||||
1. **ComfyUI container limit**: set to **24 G** when running without `--lowvram`. Current revert to 20 G + `--lowvram` is stable but slower.
|
||||
2. **CT 104 LXC allocation (49 GiB)**: appropriately sized given Docker workload, but is by far the largest single consumer. Do not raise further without measuring host impact.
|
||||
3. **CT 111 (dev, 32 GiB)**: Coder workspaces could spike if users run heavy jobs. Consider adding a per-workspace memory limit in the Coder template.
|
||||
4. **Watch list**: CT 101 (Shepard, 32 GiB) + CT 104 simultaneously at peak = 54 GB → host would need to swap. Unlikely in practice but possible during CI runs on CT 111 + FLUX generation on CT 104.
|
||||
5. **Long-term**: when CT 109 (ops) is built, run Prometheus `node_exporter` on the PVE host and alert when host available RAM drops below 8 GiB.
|
||||
|
||||
## ComfyUI memory optimisation log
|
||||
|
||||
| Date | Change | Effect |
|
||||
|------|--------|--------|
|
||||
| 2026-05-22 | Removed `--lowvram` + raised limit to 28 G | OOM at 28 G (XPU DRM buffers counted against cgroup) |
|
||||
| 2026-05-22 | Raised limit to 40 G | Host destabilised — reverted |
|
||||
| 2026-05-22 | Reverted to `--lowvram` + 20 G | Stable, slow (231 s/image) |
|
||||
| 2026-05-22 | Downloaded `t5-v1_1-xxl-encoder-Q4_K_S.gguf` (2.6 GB vs 3.2 GB Q5_K_M) | Saves 600 MB at load time |
|
||||
| 2026-05-22 | Kept Q5_K_M T5 (Q4_K_S degrades prompt following per city96); removed `--lowvram`; set limit to 24 G; added `--async-offload --force-fp16` | ~7–30 s generation, safe within host memory budget |
|
||||
| 2026-05-22 | Removed `--async-offload` | Flag incompatible with GGUF img2img on XPU — caused full CPU fallback (3.5 min/step) and pure-noise output. Removed; XPU generation now correct at ~7 s/step for txt2img. Current CLI: `--listen 0.0.0.0 --enable-cors-header --use-pytorch-cross-attention --disable-smart-memory --force-fp16` |
|
||||
Reference in New Issue
Block a user