baseline: docs as of pre-reorg audit (2026-05-20)
This commit is contained in:
@@ -0,0 +1,999 @@
|
||||
# Proxmox Host Optimization Inventory — `nuc`
|
||||
|
||||
Generated: 2026-05-20
|
||||
Host: `nuc` · PVE `9.1.11` · Kernel `6.17.13-4-pve` · Debian 13 (trixie)
|
||||
CPU: Intel Core Ultra 7 155H (16C / 22T, hybrid P+E+LP-E) · 1 socket · 1 NUMA
|
||||
RAM: 62 GiB physical · 31 GiB zram swap (50 % of RAM, zstd, prio 100)
|
||||
Storage: single Crucial P3 2 TB NVMe (QLC, DRAM-less) → `rpool` (ZFS, ashift=12, no redundancy)
|
||||
Workload: 1 VM (HAOS) + 6 LXCs (Docker, Adguard, Backrest, Nextcloud, Zoraxy, Dockge)
|
||||
|
||||
---
|
||||
|
||||
## TL;DR — top 5 actionable wins
|
||||
|
||||
1. **Memory overcommit is dangerous.** Allocated guest RAM (≈ 290 GiB) is ~4.7× physical (62 GiB). Right-size CT 101 (was 160 → done, now 32) and CT 104 (still 128) — see §2. ✅ *applied 2026-05-20*
|
||||
2. **ZFS ARC is artificially capped at 6.2 GiB.** Default would be ~31 GiB. After §1 settles, raise to 16 GiB. See §3.
|
||||
3. **No redundancy on a QLC SSD with 19 % wear and 59 TB written.** Single-disk `rpool` on a DRAM-less consumer QLC drive is a SPOF. Add a second NVMe and convert to mirror — biggest reliability win available. See §6.
|
||||
4. **Backups never prune.** Was configured `keep-all=1` — fixed to `keep-last=3,keep-daily=7,keep-weekly=4,keep-monthly=6`. See §7. ✅ *applied 2026-05-20*
|
||||
5. **`atime` and `autotrim`** on ZFS. ✅ *applied 2026-05-20*
|
||||
6. **No DNS rewrites** in AdGuard — every internal target is IP-only; add a split-horizon for `nuclide.systems` and a `.lan` shorthand set. See §11a.
|
||||
7. **Self-signed Proxmox web UI cert** — front via Zoraxy for free LE. See §11b.
|
||||
|
||||
---
|
||||
|
||||
## 1. System snapshot
|
||||
|
||||
| Resource | State | Notes |
|
||||
|---|---|---|
|
||||
| Load avg | normal | PSI: CPU some=2.4 % / IO some=1 % over 60 s |
|
||||
| Memory | 52 / 62 GiB used, 4.5 GiB free | tight; zram swap 15 GiB in use |
|
||||
| Swap | zram0 (zstd, 31 GiB) prio 100 | working as designed; just a symptom of §1 |
|
||||
| ARC | 6.0 / 6.2 GiB (capped) | hit ratio ~99 % but cap is far below default |
|
||||
| NVMe wear | **Percentage Used 19 %**, 59.2 TB written | ~5 % wear/year at current rate; healthy for now |
|
||||
| Temperature | 56–58 °C | well under the 95 °C critical threshold |
|
||||
| Uptime | (see `uptime`) | scrub clean, no checksum errors |
|
||||
| Cluster | standalone | quorum OK, no HA configured |
|
||||
|
||||
---
|
||||
|
||||
## 2. Memory & VM/CT sizing (measured numbers)
|
||||
|
||||
Read from `/sys/fs/cgroup/lxc/<id>/memory.{current,peak,max}` and `free -h` inside each guest:
|
||||
|
||||
| Guest | Cap | Current | Peak | Swap-in-use | Verdict |
|
||||
|---|---|---|---|---|---|
|
||||
| VM 100 (haos) | 16 384 MiB | 13 056 MiB | n/a | n/a | balloon disabled; HAOS actually uses what it has |
|
||||
| CT 101 (shepard) | 160 000 MiB | **7.2 GiB** | **15.6 GiB** | 706 MiB | **wildly over-sized** — peak is 10 % of cap |
|
||||
| CT 102 (adguard) | 512 MiB | 343 MiB | **509 MiB (99 %)** | 19 MiB | **under-sized** — at the cap, AdGuardHome alone is 350 MiB |
|
||||
| CT 103 (backrest) | 512 MiB | 89 MiB | 305 MiB | 16 MiB | fine |
|
||||
| CT 104 (docker/AI) | 128 000 MiB | 18.8 GiB | **29.3 GiB** | **9.3 GiB** | real workload, but currently swapping — likely starved by CT 101 |
|
||||
| CT 105 (nextcloud) | 8 192 MiB | 2.1 GiB | 3.5 GiB | 53 MiB | fine |
|
||||
| CT 108 (zoraxy) | 2 048 MiB | 271 MiB | 463 MiB | 25 MiB | fine; could halve |
|
||||
|
||||
**Sum of declared caps ≈ 290 GiB on a 62 GiB host.** Sum of actual peaks ≈ 49 GiB — totally fits. CT 101's 160 GB cap is the entire problem: it's a phantom that scares the scheduler without using anything close to that.
|
||||
|
||||
### Concrete CT 101 picture
|
||||
|
||||
12 cores, load avg 8.5, ~9 Docker containers (Shepard frontend/backend, Keycloak, Neo4j, MongoDB, MongoExpress, TimescaleDB, Caddy, home-showcase-collector). Peak RSS 15.6 GiB.
|
||||
|
||||
→ Drop memory cap to **32 GiB** (2× peak headroom). No reboot required for LXC memory changes.
|
||||
|
||||
### Concrete CT 104 picture
|
||||
|
||||
16 cores, load avg 8.0, ~65 Docker containers including Immich (with ML/vectorchord), ComfyUI (image-gen), LobeChat, n8n, Daytona, LiteLLM, Vaultwarden, Paperless-ngx+AI, Karakeep, Memos, Gotify, Garage S3, plus a forest of MCP servers, Speaches (OpenVINO using the Arc iGPU). 128 GiB of 200 GiB rootfs used.
|
||||
|
||||
Peak RSS 29.3 GiB, but **9.3 GiB sitting in swap** — under memory pressure. Two paths:
|
||||
1. **Recommended**: cut CT 101 first, then CT 104's pressure mostly disappears on its own. Re-measure peak after CT 101 is fixed. Likely safe to cap at **48 GiB** then.
|
||||
2. Leave the 128 GiB cap as a generous ceiling — harmless once CT 101 is sane.
|
||||
|
||||
### Other guests
|
||||
|
||||
- **CT 102 (AdGuard)**: hitting its cap. Bump to **1 GiB** to give DNS room to breathe. (DNS is the single point of failure for the LAN — don't squeeze it.) `pct set 102 -memory 1024`.
|
||||
- **CT 108 (Zoraxy)**: cap **1 GiB** is plenty.
|
||||
- **VM 100 (HAOS)**: 13 GiB actually used. Enable ballooning so the host can reclaim idle RAM:
|
||||
```bash
|
||||
qm set 100 -balloon 4096
|
||||
```
|
||||
This leaves memory=16384 as a ceiling but lets the host shrink it under pressure.
|
||||
- KSM is already active (10 255 shared pages) — fine.
|
||||
- zram swap: keep size at `ram / 2` until CT 101 is fixed; reduce to `ram / 4` afterwards.
|
||||
|
||||
---
|
||||
|
||||
## 3. ZFS tuning
|
||||
|
||||
### Pool
|
||||
|
||||
| Setting | Current | Recommend | Why |
|
||||
|---|---|---|---|
|
||||
| `autotrim` | **off** | `on` | QLC needs trim; weekly fstrim alone is OK but autotrim is "free" |
|
||||
| `ashift` | 12 | keep | correct for NVMe |
|
||||
| `atime` | **on** (relatime) | **off** | unused on a hypervisor; reduces write amp on QLC |
|
||||
| `xattr` | sa | keep | already optimal |
|
||||
| `compression` | on (lz4) | keep | helping (1.61× on HAOS disk) |
|
||||
| `dnodesize` | legacy | `auto` | minor; only matters with millions of small files |
|
||||
| `recordsize` (rpool) | 128 K | keep for general | tune per-dataset (see below) |
|
||||
|
||||
### ARC
|
||||
|
||||
`/etc/modprobe.d/zfs.conf` currently caps `zfs_arc_max=6669991936` (≈ 6.2 GiB).
|
||||
- After §2 sizing is done, raise this to 16 GiB: `options zfs zfs_arc_max=17179869184` and `zfs_arc_min=4294967296`.
|
||||
- Apply live without reboot: `echo 17179869184 > /sys/module/zfs/parameters/zfs_arc_max`.
|
||||
|
||||
### Per-dataset
|
||||
|
||||
- **VM zvols** (`rpool/data/vm-*`): default `volblocksize` is 16 K — fine. HAOS disk uses `cache=writethrough`; on ZFS, **switch to `cache=none`** (or unset) — writethrough doubles the sync cost on top of ZFS's own integrity guarantees.
|
||||
- **CT 104 (`subvol-104-disk-0`, Docker + image-gen)**: keep `recordsize=128K`. The workload is dominated by large model files and image outputs, not small-file DB traffic — shrinking the record size would hurt, not help.
|
||||
- **Nextcloud (`subvol-105-disk-1`)**: leave at 128 K (mixed sizes, mostly larger files).
|
||||
|
||||
### Pool features
|
||||
|
||||
`zpool upgrade rpool` was run during this audit and enabled `redaction_list_spill` + `raidz_expansion`. Other disabled features (`fast_dedup`, `longname`, `large_microzap`, `dynamic_gang_header`, `block_cloning_endian`, `physical_rewrite`) can be enabled with another `zpool upgrade rpool` — only do this if you do not need to roll back to an older ZFS.
|
||||
|
||||
### Commands
|
||||
|
||||
```bash
|
||||
zpool set autotrim=on rpool
|
||||
zfs set atime=off rpool
|
||||
# (optional, once memory is sane):
|
||||
echo 'options zfs zfs_arc_max=17179869184' > /etc/modprobe.d/zfs.conf
|
||||
update-initramfs -u -k all
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. Storage & VM disk options
|
||||
|
||||
### VM 100 (`haos`)
|
||||
|
||||
```diff
|
||||
- scsi0: local-zfs:vm-100-disk-1,cache=writethrough,discard=on,size=32G,ssd=1
|
||||
+ scsi0: local-zfs:vm-100-disk-1,cache=none,discard=on,iothread=1,size=32G,ssd=1
|
||||
```
|
||||
|
||||
- `cache=none` (or remove cache entirely) — let ZFS manage caching.
|
||||
- `iothread=1` with `virtio-scsi-pci` controller — already using `virtio-scsi-pci`, just add iothread.
|
||||
- Already has `discard=on` and `ssd=1` ✔
|
||||
|
||||
### LXC `local-zfs` storage
|
||||
|
||||
- `sparse 1` is set ✔ — thin-provisioned.
|
||||
- All CTs are on `local-zfs` rootfs; OK.
|
||||
|
||||
### UNAS share — current state (measured)
|
||||
|
||||
Backend: `192.168.1.31` (looks like a UniFi NAS — exports `/volume/.../.unifi-drive/storage/.data`, the only NFS export listed is restricted to four allowed clients: the host `.20`, CT 104 `.40`, plus `.60` and `172.30.33.1`).
|
||||
|
||||
Two parallel mounts on the host pointed at the **same backing data**:
|
||||
|
||||
| Mount | Type | Options (key bits) | Consumers |
|
||||
|---|---|---|---|
|
||||
| `/mnt/pve/unas` | NFS **v3** | `proto=tcp, mountproto=udp, rsize/wsize=1M, hard, relatime, timeo=600` | CT 103 (backrest), CT 104 (docker) — bind-mounted to `/mnt/pve/unas` inside |
|
||||
| `/mnt/pve/unas_smb` | CIFS **v3.1.1** | `cache=strict, actimeo=1, soft, rsize/wsize=4M, uid/gid=33` | CT 105 (nextcloud) — bind-mounted to `/mnt/pve/unas` inside |
|
||||
|
||||
**Issues:**
|
||||
|
||||
1. **CT 105 is on CIFS to the same data CT 104 uses via NFS.** Pure duplication. Nextcloud does massive amounts of `stat()` traffic; `actimeo=1` on the CIFS mount forces *every* metadata lookup to hit the wire, which is slow.
|
||||
2. **NFS is v3, not v4.x.** v4 is preferred unless the UDM doesn't export it. v4 fixes locking, removes the separate mountd dance, and supports session trunking.
|
||||
3. **`mountproto=udp`** under packet loss can intermittently fail to (re)mount. Set `mountproto=tcp`.
|
||||
4. **`hard` mount with no `intr` equivalent.** If UNAS goes away, anything blocked on it hangs the calling process indefinitely. For non-critical use cases (Nextcloud, but **not** backrest), `soft,timeo=100,retrans=3` is friendlier — Backrest backups should stay `hard`.
|
||||
5. **CT 105 cannot mount NFS directly** because the UNAS export only allows IPs `.20/.40/.60/.172.30.33.1` — `.41` (CT 105) is missing. So either keep the host-side bind-mount approach (correct) or have UNAS export to `.41` too.
|
||||
6. **The bind-mount approach is correct** for unprivileged CTs that can't run NFS/CIFS clients themselves. Don't change that pattern.
|
||||
|
||||
**Recommended consolidation:**
|
||||
|
||||
```bash
|
||||
# 1. Probe whether the NAS speaks NFSv4
|
||||
mount -t nfs -o vers=4.2,proto=tcp 192.168.1.31:/var/nfs/shared/storage /mnt/test
|
||||
# if it works:
|
||||
pvesm set unas --options vers=4.2,proto=tcp,hard,noatime
|
||||
# (this re-mounts on next access; or unmount/remount /mnt/pve/unas)
|
||||
|
||||
# 2. Switch CT 105 to the NFS bind-mount
|
||||
pct set 105 --mp0 /mnt/pve/unas,mp=/mnt/pve/unas
|
||||
# (CT 105 currently uses unas_smb → unas. New line bind-mounts the NFS mount.)
|
||||
# Then verify nextcloud-aio still sees uid/gid 33 properly — NFS uses host UIDs,
|
||||
# whereas CIFS was forcing uid=33. May need to chown on the NAS or add an idmap.
|
||||
|
||||
# 3. Drop the CIFS storage once CT 105 is migrated
|
||||
pvesm remove unas_smb # if it exists as PVE storage
|
||||
# or remove the entry from /etc/pve/storage.cfg
|
||||
```
|
||||
|
||||
**Notes on perf:**
|
||||
|
||||
- 1 GbE NIC ceiling is ~118 MB/s. NFS rsize/wsize=1M already saturates that.
|
||||
- If you have or add a 2.5 GbE / 10 GbE link on either side, also enable jumbo frames (MTU 9000) on `vmbr0`, the host NIC, and UNAS. That alone can ~double bulk-read throughput.
|
||||
- For Nextcloud's small-file stats: `actimeo=60` (NFS) or `cache=loose,actimeo=60` (CIFS, if you stay on it) — dramatically cuts roundtrips at the cost of slightly stale directory listings.
|
||||
|
||||
---
|
||||
|
||||
## 5. CPU / boot / kernel
|
||||
|
||||
| Item | State | Recommend |
|
||||
|---|---|---|
|
||||
| Governor | `performance` | keep |
|
||||
| HWP EPP | `default` | set to `balance_performance` if you want some idle savings without latency cost: `echo balance_performance > /sys/devices/system/cpu/cpu*/cpufreq/energy_performance_preference` |
|
||||
| `intel_iommu=on iommu=pt` | set ✔ | keep |
|
||||
| GPU passthrough (`i915.force_probe=!7dd5 xe.force_probe=7dd5`) | set for Arc Xe (Meteor Lake) | keep |
|
||||
| `nvme_core.default_ps_max_latency_us=0` | set ✔ | disables NVMe power-save — good for stability, costs ~1 W idle |
|
||||
| `kernel.numa_balancing` | 0 | correct for single socket |
|
||||
| Old kernels installed | 6.17.13-4 (current) + 7.0.0-3 | keep both for now; remove 7.0.0-3 once you've booted 7.0.2-5 successfully after the pending upgrade |
|
||||
|
||||
### Hybrid-core scheduling
|
||||
|
||||
The 155H has P-cores (cores 0–11), E-cores (12–17), LP-E cores (18–21). Linux 6.x with `intel_pstate=active` handles ITD/HWP well; no manual pinning is needed for current workloads. If a CT becomes latency-sensitive, you can pin it with `cpuset` via `lxc.cgroup2.cpuset.cpus` (P-cores only).
|
||||
|
||||
---
|
||||
|
||||
## 6. Reliability / SPOF
|
||||
|
||||
### Single disk is the biggest risk
|
||||
|
||||
- Current: 1× CT2000P3PSSD8 (QLC, DRAM-less). 19 % wear in (check `power_on_hours`).
|
||||
- **Recommendation:** add a second 2 TB NVMe (ideally a TLC drive with DRAM — Crucial T500, WD SN770, Samsung 990 Pro, Solidigm P44 Pro). Then:
|
||||
```bash
|
||||
zpool attach rpool nvme-CT2000P3PSSD8_2429E8BBCFB4-part3 /dev/disk/by-id/<new-disk>-part3
|
||||
```
|
||||
(requires partitioning the new disk to match — `sgdisk -R` from the existing). Pool becomes a mirror with full self-heal.
|
||||
- If a second internal slot isn't available, at minimum increase backup frequency for CT 101/104 (the irreplaceable data).
|
||||
|
||||
### Boot redundancy
|
||||
|
||||
`proxmox-boot-tool kernel list` shows one bootloader entry. After §6 mirror is set up, run `proxmox-boot-tool init /dev/<new-disk>-partN` so either disk can boot.
|
||||
|
||||
---
|
||||
|
||||
## 7. Backups (HIGH PRIORITY — silent risk)
|
||||
|
||||
`/etc/pve/storage.cfg`:
|
||||
```
|
||||
nfs: unas
|
||||
prune-backups keep-all=1
|
||||
```
|
||||
|
||||
**`keep-all=1` means backups are never deleted automatically.** UNAS already holds 2 TB. Set a real policy, e.g.:
|
||||
|
||||
```bash
|
||||
pvesm set unas --prune-backups keep-last=3,keep-daily=7,keep-weekly=4,keep-monthly=6
|
||||
```
|
||||
|
||||
Also: there is **no `vzdump` job configured** in `/etc/pve/jobs.cfg`. Backups are either manual or driven from CT 103 (Backrest). Recommend a scheduled `vzdump` job for at least VM 100 and CT 101/104 *in addition to* Backrest, so PVE-native restores remain trivial.
|
||||
|
||||
---
|
||||
|
||||
## 8. APT / repositories cleanup
|
||||
|
||||
State today:
|
||||
|
||||
```
|
||||
/etc/apt/sources.list.d/
|
||||
├── ceph.list # all lines commented — fine but consider deleting the file
|
||||
├── proxmox.sources # pve-no-subscription (modern deb822) ← keep
|
||||
├── pve-enterprise.list.bak # backup, safe to remove
|
||||
├── pve-enterprise.sources # Enabled: false ← keep as-is or remove
|
||||
├── pve-install-repo.list # pve-no-subscription duplicate
|
||||
└── pve-no-subscription.list # pve-no-subscription duplicate
|
||||
```
|
||||
|
||||
`pve-install-repo.list` and `pve-no-subscription.list` duplicate what `proxmox.sources` already declares. APT deduplicates fetches but the duplication is a foot-gun (one of them will go stale on the next PVE major version transition). Recommended cleanup:
|
||||
|
||||
```bash
|
||||
rm /etc/apt/sources.list.d/pve-install-repo.list
|
||||
rm /etc/apt/sources.list.d/pve-no-subscription.list
|
||||
rm /etc/apt/sources.list.d/pve-enterprise.list.bak
|
||||
# keep proxmox.sources and pve-enterprise.sources (already disabled)
|
||||
apt update
|
||||
```
|
||||
|
||||
Also: there are **9 pending upgrades** including `pve-manager 9.1.18` (you're on 9.1.11) and a kernel update. Run `apt update && apt full-upgrade` at a convenient window.
|
||||
|
||||
### Unattended-upgrades (configured 2026-05-20)
|
||||
|
||||
The host previously had a cron line `0 2 * * * apt-get update && apt-get upgrade -y` that was silently no-op'ing on every kernel / PVE point release — `apt-get upgrade` refuses to install new dependencies, which PVE updates always introduce.
|
||||
|
||||
Replaced with `unattended-upgrades` in a conservative profile:
|
||||
|
||||
| File | Purpose |
|
||||
|---|---|
|
||||
| `/etc/apt/apt.conf.d/52unattended-upgrades-pve` | local policy — origins allowlist + email + reboot policy |
|
||||
| `/etc/apt/apt.conf.d/20auto-upgrades` | enables the daily update-list + unattended-upgrade run |
|
||||
|
||||
**Auto-applied:**
|
||||
- `origin=Debian,codename=trixie,label=Debian` (stable main)
|
||||
- `origin=Debian,codename=trixie-security,label=Debian-Security`
|
||||
- `origin=Debian,codename=trixie-updates` (stable point updates)
|
||||
|
||||
**Held for manual `apt full-upgrade`** (intentionally — review release notes first):
|
||||
- `origin=Proxmox,...` — pve-manager, kernels, qemu-server, all PVE components
|
||||
|
||||
**Settings:**
|
||||
- `Automatic-Reboot "false"` — kernel updates require a manual reboot
|
||||
- `Remove-Unused-Dependencies "true"` — autoremove orphans after upgrades
|
||||
- `AutoFixInterruptedDpkg "true"` — resume after crash mid-upgrade
|
||||
- `Mail "notify@home.box"`, `MailReport "on-change"` — alerts on actual changes
|
||||
|
||||
**Triggered by:**
|
||||
- `apt-daily.timer` (daily ~07:00) — refresh package lists
|
||||
- `apt-daily-upgrade.timer` (daily ~06:00) — apply unattended upgrades
|
||||
|
||||
**Caveat: mail delivery isn't reaching you yet.** Postfix is up but has `relayhost = (none)` — change notifications get delivered locally to `/var/mail/notify` on the host, not to your inbox. Set up a smart-host relay (Gmail/Postmark/etc.) if you want the mails to actually land. Until then, check `/var/log/unattended-upgrades/unattended-upgrades.log` for history.
|
||||
|
||||
**Verify any time:**
|
||||
```bash
|
||||
unattended-upgrade --dry-run --debug 2>&1 | grep -E "Allowed origins|would be upgraded|pkgs that look"
|
||||
systemctl list-timers apt-daily-upgrade.timer
|
||||
tail /var/log/unattended-upgrades/unattended-upgrades.log
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 9. Networking
|
||||
|
||||
- Single bridge `vmbr0` on `enp86s0` — no VLAN aware (`bridge-vlan-aware yes`). If you ever want to segment guests by VLAN, add it now (no impact on existing guests as long as you don't tag them):
|
||||
```
|
||||
bridge-vlan-aware yes
|
||||
bridge-vids 2-4094
|
||||
```
|
||||
- `net.core.rmem_max` / `wmem_max` are at distro defaults (208 KiB). With a 1 GbE NIC the impact is small (link is already saturated at NFS rsize=1M), but with future 2.5/10 GbE bump to 16 MiB:
|
||||
```bash
|
||||
cat >/etc/sysctl.d/99-net.conf <<'EOF'
|
||||
net.core.rmem_max=16777216
|
||||
net.core.wmem_max=16777216
|
||||
net.ipv4.tcp_rmem=4096 87380 16777216
|
||||
net.ipv4.tcp_wmem=4096 65536 16777216
|
||||
EOF
|
||||
sysctl --system
|
||||
```
|
||||
- TCP congestion control is `cubic`. `bbr` is generally better for mixed workloads — change only if you measure a problem.
|
||||
- `wlo1` is present but unused — confirm and disable in BIOS or `iface wlo1 inet manual` (already done). No action.
|
||||
|
||||
---
|
||||
|
||||
## 10. Container-specific issues
|
||||
|
||||
### CT 104 (docker / AI / image-gen) — 128 GiB cap, 16 cores, GPU passthrough
|
||||
|
||||
Measured: 18.8 GiB current, peak 29.3 GiB, **9.3 GiB in swap**, load 8.0, ~65 Docker containers (Immich + ML, ComfyUI, LobeChat, n8n, Daytona, LiteLLM, Vaultwarden, Paperless+AI, many MCP servers, Speaches-OpenVINO).
|
||||
|
||||
- `cores: 16` is justified by the workload (load avg 8 across 16 = ~50 % avg). Don't drop.
|
||||
- 128 GiB cap is generous but currently 9 GiB is in swap — this is **memory pressure from CT 101's phantom cap**, not from CT 104's own needs. Fix CT 101 first, re-measure, then either lower CT 104 to 48 GiB or leave it.
|
||||
- GPU pass-through (`/dev/dri/card1`, `renderD128`) confirmed visible inside CT and being used by Speaches via OpenVINO ✔
|
||||
- Keep `recordsize=128K` (large files dominate).
|
||||
- `swap: 32000` is high — consider `swap: 8192`. Heavy CT swap-out on a QLC root SSD adds write amplification.
|
||||
- Bind-mounted `/mnt/pve/unas` (NFS) is the right choice ✔
|
||||
- Bonus: raise scheduler weight if image-gen latency matters: `pct set 104 -cpuunits 200`.
|
||||
|
||||
### CT 101 (shepard / docker) — currently 160 GiB cap, peak 15.6 GiB
|
||||
|
||||
Workload: ~9 containers — Shepard frontend/backend, Keycloak, Neo4j, MongoDB, MongoExpress, TimescaleDB, Caddy, home-showcase-collector.
|
||||
|
||||
- **Drop memory cap to 32 GiB**: `pct set 101 -memory 32768`. No restart needed.
|
||||
- `cores: 12` is fine (load 8.5 — close to fully loaded, real work).
|
||||
- `swap: 8192` ✔
|
||||
- This single change unblocks everything else.
|
||||
|
||||
### CT 102 (adguard / DNS) — undersized
|
||||
|
||||
Measured: 343 MiB used at the **512 MiB cap**, AdGuardHome alone is 350 MiB RSS, the CT is one OOM event from killing the LAN's DNS.
|
||||
|
||||
- `pct set 102 -memory 1024` (bump to 1 GiB).
|
||||
- Add `onboot: 1` (already set) + `startup: order=1` (boots first) + `protection: 1` (anti-fatfinger). This is the only DNS — treat it like infrastructure.
|
||||
|
||||
### CT 103 (backrest) — fine
|
||||
|
||||
89 MiB used, peak 305 MiB. No changes needed.
|
||||
|
||||
### CT 105 (nextcloud) — privileged + CIFS
|
||||
|
||||
- **Privileged container** (no `unprivileged: 1`). For a public-facing app this is the wrong tradeoff. Migration: stop CT, `vzdump` backup, restore as unprivileged. Be ready to fix file ownership on the bind-mount afterwards (privileged UID 33 → unprivileged needs `lxc.idmap`).
|
||||
- **Uses CIFS bind-mount to UNAS** with `actimeo=1` — see §4. Switch to NFS bind-mount, after confirming UID mapping (CIFS forces uid=33; NFS uses host UIDs as-is).
|
||||
- Memory cap 8 GiB, peak 3.5 GiB — fine as-is.
|
||||
- Runs Nextcloud AIO stack (mastercontainer-managed: apache, postgres, redis, imaginary, collabora, notify-push).
|
||||
|
||||
### CT 108 (zoraxy) — slight oversize
|
||||
|
||||
Peak 463 MiB on a 2 GiB cap. Lower to 1 GiB if desired (cosmetic).
|
||||
|
||||
### CT 104 — Docker stacks inventory (`/opt/stacks`)
|
||||
|
||||
CT 104 keeps its Docker workloads in a **git-tracked monorepo at `/opt/stacks/`** with one directory per stack, plus meta-docs (`PORTMAP.md`, `storage.md`, `volumes.md`, `docker-networks.md`, `todo.md`). Good practice — this is how to keep ~65 containers manageable. The other CTs (101, 105) don't have `/opt/stacks` — their compose files live elsewhere.
|
||||
|
||||
**Stack list (29 dirs):**
|
||||
|
||||
| Stack | Status | Notes |
|
||||
|---|---|---|
|
||||
| `ai/` | **active** — large subtree | comfyui, lobehub, litellm, speaches, mcp-gateway, mcp-servers (many MCP yml files), searxng. Custom `syncstack.py` to manage cross-file project names. |
|
||||
| `arr-stack/` | dormant (defined, not running) | rdtclient, prowlarr, audiobookshelf, shelfarr, flaresolverr |
|
||||
| `arcane/` | dormant | Docker dashboard |
|
||||
| `daytona/` | active (as `daytona-minimal`) | dev environments + runner + registry |
|
||||
| `dozzle/` | dormant | container log viewer |
|
||||
| `gotify/` | active | push notifications |
|
||||
| `homepage/` | dormant | dashboard |
|
||||
| `immich/` | active (5 containers) | photo platform + ML |
|
||||
| `karakeep/` | active (3 containers) | bookmark mgr + chrome + meilisearch |
|
||||
| `memos/` | active | notes |
|
||||
| `n8n/` | active | pinned `2.20.11` (good — there's an explicit version-drift comment in the compose) |
|
||||
| `nexa/` | dormant | (?) |
|
||||
| `paperless_ai/`, `paperless-ngx/` | active (5 containers between them) | OCR pipeline |
|
||||
| `pocketid/` | active | OIDC provider |
|
||||
| `proxy/` | dormant | (?) |
|
||||
| `qdrant/` | dormant | vector DB |
|
||||
| `shared-db/` | active | shared-postgres + garage (S3-compatible) + pgadmin (defined) |
|
||||
| `streamio/` | dormant | media |
|
||||
| `traccar/` | active | GPS tracker |
|
||||
| `vaultwarden/` | active | password mgr |
|
||||
| `vpn/` | dormant | gluetun (intended VPN egress wrapper?) |
|
||||
| `backups/`, `docs/`, `scripts/` | meta dirs (no compose) | |
|
||||
|
||||
**Observations:**
|
||||
|
||||
1. **~10 stacks are defined but dormant.** No RAM/CPU cost while down, but their images sit on disk and the git repo accrues dead code. Either run them, document why they're parked, or `git rm` them — repo drift is the silent killer of "I know what's running" confidence.
|
||||
2. **Stack name ≠ compose project name** for the `ai/` and `daytona/` trees (multiple compose files per dir, different project names). The `syncstack.py` helper exists for this; just be aware that `docker compose -f` lookups by directory name don't match.
|
||||
3. **Disk-reclaim potential (measured `docker system df`):**
|
||||
|
||||
| Asset | Total | Reclaimable |
|
||||
|---|---|---|
|
||||
| Images | 83 / 61.6 GB | **8.3 GB** |
|
||||
| Build cache | 105 entries / 9.0 GB | **4.6 GB** |
|
||||
| Volumes | 31 / 3.0 GB | 940 MB (21 dangling) |
|
||||
| Containers | 60 active | 0 |
|
||||
|
||||
```bash
|
||||
pct exec 104 -- docker system prune -a --volumes
|
||||
# or non-destructively just the build cache:
|
||||
pct exec 104 -- docker builder prune -a
|
||||
```
|
||||
~13 GB to recover. On a 200 GB rootfs that's 64 % full, this is meaningful.
|
||||
|
||||
4. **Stack→Zoraxy mapping (§11b):** when fronting via Zoraxy, the canonical service endpoints (per `PORTMAP.md`) are CT 104's IP `192.168.1.40` + port. Worth cross-referencing that file when setting up reverse-proxy entries.
|
||||
|
||||
### General LXC hygiene
|
||||
|
||||
- Strip the giant HTML banners from the top of every `/etc/pve/lxc/*.conf` — pure noise in `pct config`.
|
||||
- All CTs except 105 are unprivileged ✔ with `nesting=1,keyctl=1` (appropriate for Docker).
|
||||
|
||||
### Anti-fat-finger protection (`protection: 1`) + boot order
|
||||
|
||||
Applied across all critical guests (2026-05-20). `protection: 1` blocks `pct destroy` / "Remove" from the UI until manually unset — cheap insurance against the wrong-CT-deleted incident.
|
||||
|
||||
| Guest | protection | startup order | Rationale |
|
||||
|---|---|---|---|
|
||||
| CT 102 (AdGuard / DNS) | ✅ | **order=1** (boots first) | LAN-wide DNS — nothing resolves until this is up |
|
||||
| CT 108 (Zoraxy / LE proxy) | ✅ | **order=2** | Public-facing reverse proxy + LE; depends on DNS |
|
||||
| CT 103 (Backrest) | ✅ | default | Holds backup config and snapshot metadata |
|
||||
| CT 104 (Docker / AI / image-gen) | ✅ | default | Largest data footprint (200 G rootfs); 65 containers |
|
||||
| CT 105 (Nextcloud) | ✅ | default | User data |
|
||||
| VM 100 (HAOS) | ✅ | default | Home automation state |
|
||||
| CT 101 (shepard) | ❌ (left optional) | default | Currently a dev/iteration target; protect once stabilized: `pct set 101 -protection 1` |
|
||||
|
||||
The `onboot: 1` flag was already set on all guests ✔ — they all auto-start on host reboot. The two `startup` ordered ones now also boot in the right sequence: DNS → Zoraxy → everything else in parallel.
|
||||
|
||||
To remove protection on a guest later: `pct set <id> -protection 0` (or `qm set 100 -protection 0`).
|
||||
|
||||
---
|
||||
|
||||
## 11a. DNS — AdGuard rewrites & site-wide consistency
|
||||
|
||||
**Current state (measured):**
|
||||
|
||||
- AdGuard Home runs on CT 102 (`192.168.1.2`, hostname `dns`) — DNS on `0.0.0.0:53/tcp+udp`, admin UI on `:80`.
|
||||
- Upstream stack is well-configured: load-balanced mix of Telekom-DoT, NextDNS, DNS4EU, Quad9, Cloudflare, Google (DoH/DoQ where supported), bootstrap+fallback on Cloudflare.
|
||||
- All 6 LXCs resolve via `nameserver 192.168.1.2` ✔
|
||||
- Proxmox host (`192.168.1.20`) resolves via `192.168.1.2` ✔ — but its `search` domain is `box` (probably an install-time leftover; AdGuard's `local_domain_name` is `lan`).
|
||||
- HAOS VM is on `192.168.1.60` — DNS setting unknown without Home Assistant access; **verify**.
|
||||
- `rewrites: []` and **`rewrites_enabled: false`** — no internal name resolution is happening today.
|
||||
|
||||
### Why this matters
|
||||
|
||||
Without rewrites, you address everything by IP. That's brittle (IP changes break links), invisible in logs, and prevents nice tricks like split-horizon DNS for `nuclide.systems` (so the same name resolves to Zoraxy LAN-internally without going through your public IP / WAN hairpin).
|
||||
|
||||
### Recommended rewrite set
|
||||
|
||||
In AdGuard UI → Filters → DNS rewrites, **enable rewrites** and add:
|
||||
|
||||
```
|
||||
# Split-horizon: public domain → Zoraxy on LAN
|
||||
nuclide.systems → 192.168.1.4
|
||||
*.nuclide.systems → 192.168.1.4
|
||||
|
||||
# Service-name shortcuts under the local_domain_name (`.lan`)
|
||||
pve.lan → 192.168.1.20 # Proxmox UI
|
||||
nuc.lan → 192.168.1.20 # host shorthand
|
||||
dns.lan → 192.168.1.2 # AdGuard itself
|
||||
zoraxy.lan → 192.168.1.4 # reverse proxy
|
||||
shepard.lan → 192.168.1.49 # CT 101
|
||||
docker.lan → 192.168.1.40 # CT 104
|
||||
nextcloud.lan → 192.168.1.41 # CT 105
|
||||
haos.lan → 192.168.1.60 # VM 100
|
||||
unas.lan → 192.168.1.31 # NAS
|
||||
router.lan → 192.168.1.1 # UniFi gateway
|
||||
```
|
||||
|
||||
The split-horizon entries are the highest-value: once Zoraxy proxies `pve.nuclide.systems` (see §11b), the same URL works both from the public internet and from inside the LAN — with no NAT-loopback weirdness and with the LAN traffic never leaving the building.
|
||||
|
||||
**Edit the YAML directly if preferred** (`/opt/AdGuardHome/AdGuardHome.yaml` inside CT 102), then restart AdGuard. The line `rewrites_enabled: false` must flip to `true`.
|
||||
|
||||
### Verify clients are actually using AdGuard
|
||||
|
||||
After rewrites are in, walk the inventory:
|
||||
|
||||
| Client | Should use DNS | Check |
|
||||
|---|---|---|
|
||||
| All 6 LXCs | ✔ already at .2 | `pct exec <id> -- cat /etc/resolv.conf` |
|
||||
| Proxmox host | ✔ already at .2 | `cat /etc/resolv.conf` |
|
||||
| HAOS VM (192.168.1.60) | unknown | HAOS UI → Settings → System → Network → check DNS servers; should be `192.168.1.2` |
|
||||
| Router (192.168.1.1, UniFi) | **DHCP-hands-out DNS to clients** — must serve .2 as primary | UniFi: Settings → Networks → LAN → DHCP DNS: `192.168.1.2` |
|
||||
| IoT devices (Roborock at .64, others) | inherit via DHCP from router | once UniFi DHCP serves .2, every device that DHCP-renews picks it up. Force-renew or reboot stragglers. |
|
||||
| Anything with hard-coded `1.1.1.1` / `8.8.8.8` | bypassing the filter | grep service configs for upstream DNS — apps like Pi-hole-aware clients, some Smart TVs, Chromecasts |
|
||||
|
||||
### Optional hardening once the rewrites are stable
|
||||
|
||||
- **DNSSEC**: `enable_dnssec: true` (currently false). Most upstreams already validate, but flipping this on adds end-to-end checking.
|
||||
- **Block DoH bypass**: AdGuard has a "Blocked services" preset for Cloudflare/Google DoH endpoints — useful so clients can't sidestep the filter.
|
||||
- **Resolve `box` search-domain confusion**: edit `/etc/resolv.conf` (or set it via `/etc/network/interfaces`) on the host to `search lan` so it matches AdGuard's `local_domain_name`.
|
||||
- **Secondary DNS for resilience**: CT 102 is currently the sole DNS — if it crashes, the LAN goes blind. Either add `protection: 1` (done above) and rely on it, or stand up a tiny secondary AdGuard on a different CT and configure UniFi DHCP to hand out both. (Out of scope for low-hanging fruit, but worth knowing.)
|
||||
- **Conditional forwarding for UNAS**: if UNAS announces hostnames (mDNS or its own DNS), point AdGuard at it for the relevant zone via "Upstream DNS servers" using the `[/zone/]upstream` syntax.
|
||||
|
||||
### Action checklist
|
||||
|
||||
1. AdGuard UI → Filters → DNS rewrites: paste the table above.
|
||||
2. AdGuard UI → Settings → DNS settings → **enable** "DNS rewrites".
|
||||
3. UniFi: confirm DHCP option 6 = `192.168.1.2` (LAN clients get AdGuard).
|
||||
4. HAOS: confirm Home Assistant has `192.168.1.2` set as DNS.
|
||||
5. Force-renew DHCP leases on key clients (or just wait — most renew within 24 h).
|
||||
6. After §11b is done, the public `pve.nuclide.systems` resolves to `.4` from inside the LAN automatically.
|
||||
|
||||
---
|
||||
|
||||
## 11b. TLS certificates for the Proxmox web UI
|
||||
|
||||
**Current state:** the PVE web UI on `https://192.168.1.20:8006` uses the self-signed certificate generated at install (`/etc/pve/local/pveproxy-ssl.pem` is absent → falls back to `pve-ssl.pem`). Every login throws a browser warning.
|
||||
|
||||
**The wider setup:** Zoraxy (CT 108, `192.168.1.4`) already handles Let's Encrypt for `nuclide.systems` (the public domain for this host). So there are three sane options; pick **A** unless you have a reason not to.
|
||||
|
||||
### Option A — Reverse-proxy PVE through Zoraxy (recommended)
|
||||
|
||||
Pros: single source of LE truth (Zoraxy already renews); no DNS-plugin setup; no exposing the API; nice domain like `pve.nuclide.systems`. Cons: depends on Zoraxy being up (keep IP:8006 as fallback); needs WebSocket pass-through for the noVNC console and `xterm.js` shell.
|
||||
|
||||
1. **Zoraxy host entry**
|
||||
- Domain: `pve.nuclide.systems` (or whatever subdomain)
|
||||
- Target: `https://192.168.1.20:8006`
|
||||
- **Enable WebSocket support** (mandatory — noVNC, xterm.js, task log streaming all use it)
|
||||
- **Skip backend TLS verification** (PVE's cert is self-signed)
|
||||
- Enable HSTS once you've confirmed the setup works
|
||||
- Optionally restrict by source: only LAN + Cloudflare/Tailscale IPs
|
||||
2. **DNS**: add an A record `pve.nuclide.systems` → public IP (or split-horizon to `192.168.1.20` for LAN). Zoraxy will ACME-challenge via whichever method it's configured for (HTTP-01 or DNS-01).
|
||||
3. **Keep `https://192.168.1.20:8006` reachable on LAN** as an emergency fallback. Don't disable it.
|
||||
4. Set the PVE redirect-to-HTTPS rules in Zoraxy on for both `:80` and `:443`.
|
||||
|
||||
**Important caveat:** the PVE Mobile app and the `pvesh` / API clients may not love going through a reverse proxy (they're picky about TLS SNI and cookie domains). Keep direct IP access available for API tooling, or test thoroughly.
|
||||
|
||||
### Option B — PVE's built-in ACME with DNS-01
|
||||
|
||||
Pros: no reverse proxy in the path; PVE renews itself; works for the API too. Cons: requires a DNS provider plugin (your registrar's API token), and an LE-acceptable FQDN that resolves publicly.
|
||||
|
||||
1. Register an ACME account:
|
||||
```bash
|
||||
pvenode acme account register default you@nuclide.systems
|
||||
```
|
||||
2. Configure a DNS plugin. PVE supports `acme-dns`, `cloudflare`, `route53`, `desec`, etc. via the `acme.sh` plugin set. Example for Cloudflare:
|
||||
```bash
|
||||
pvenode acme plugin add dns cf --api cf --data CF_Token=XXXXXXXX
|
||||
```
|
||||
Replace `cf` plugin name to match whichever registrar you use for `nuclide.systems`.
|
||||
3. Tell the node which domain(s) and how:
|
||||
```bash
|
||||
pvenode config set --acme domains=nuc.nuclide.systems
|
||||
pvenode config set --acmedomain0 domain=nuc.nuclide.systems,plugin=cf
|
||||
```
|
||||
4. Order:
|
||||
```bash
|
||||
pvenode acme cert order
|
||||
```
|
||||
PVE drops the cert at `/etc/pve/nodes/nuc/pveproxy-ssl.pem` and renews ~30 days before expiry via the `pve-daily-update` timer.
|
||||
|
||||
### Option C — Push Zoraxy's cert into PVE
|
||||
|
||||
Only useful if A and B are off the table. Zoraxy stores its issued certs (location varies by Zoraxy version — typically under its data dir, e.g. `/opt/zoraxy/conf/certs/`). Cron a script that copies the active cert/key and concatenates them as `/etc/pve/local/pveproxy-ssl.pem` (cert + chain) and `/etc/pve/local/pveproxy-ssl.key`, then `systemctl reload pveproxy`. Brittle — only worth it if you must.
|
||||
|
||||
### Recommended path
|
||||
|
||||
Do **A** (reverse proxy through Zoraxy) for the web UI. It piggybacks on existing renewal. The mobile-app/API edge cases are usually fine if Zoraxy passes the WebSocket and preserves the `Host` header. If you later need full ACME on the node itself (e.g. you want valid TLS for `pvesh` and the API at the node FQDN too), layer **B** on top — they don't conflict.
|
||||
|
||||
### Other CTs with web UIs worth fronting via Zoraxy
|
||||
|
||||
While you're at it, route through Zoraxy for free LE:
|
||||
- Backrest (CT 103) — currently IP-only
|
||||
- AdGuard (CT 102) admin UI — `192.168.1.2:3000`
|
||||
- Nextcloud (CT 105) — almost certainly already exposed; verify it terminates LE in Zoraxy and not internally
|
||||
- Zoraxy itself (CT 108) — self-hosted, already TLS
|
||||
|
||||
For each, add a Zoraxy host entry, set a subdomain, and disable any local TLS / port-exposed listener that bypasses Zoraxy.
|
||||
|
||||
---
|
||||
|
||||
## 11. Maintenance / observability
|
||||
|
||||
| Item | State | Recommend |
|
||||
|---|---|---|
|
||||
| `lm-sensors` | not installed | `apt install lm-sensors && sensors-detect --auto` for CPU/NVMe temps in the UI |
|
||||
| Journal size | 1.5 GiB | OK; cap at 1 GiB if you want predictability: `journalctl --vacuum-size=1G` and `SystemMaxUse=1G` in `journald.conf` |
|
||||
| `fstrim.timer` | active (weekly) | OK; `zfs trim` runs separately when autotrim=on |
|
||||
| ZFS scrub | last run 2026-05-10, clean | default monthly timer is good |
|
||||
| Subscription nag | not removed | If desired, `pve-no-nag` patch or the `proxmox-helper-scripts` line — purely cosmetic |
|
||||
| Email alerts | (check `/etc/pve/user.cfg`) | configure `root@pam` email for failed scrub / failed backup notifications |
|
||||
|
||||
---
|
||||
|
||||
## 13. Update management — current model
|
||||
|
||||
The host previously had `0 2 * * * apt-get update && apt-get upgrade -y` (silently no-op'd on every PVE/kernel update) and a weekly `bash <(wget tteck/.../update-lxcs-cron.sh)` cron that ran `dist-upgrade` across every LXC. Both removed 2026-05-20 and replaced with the structure below.
|
||||
|
||||
### Layer 1 — Host packages (Debian + Proxmox)
|
||||
|
||||
- `unattended-upgrades 2.12` installed on the host.
|
||||
- `/etc/apt/apt.conf.d/52unattended-upgrades-pve` allows only `Debian`, `Debian-Security`, `trixie-updates` — **Proxmox origin held** for manual review.
|
||||
- Triggered by `apt-daily.timer` and `apt-daily-upgrade.timer` (ship with `apt`, both `active enabled`).
|
||||
- `Automatic-Reboot "false"` — kernel updates wait for a manual reboot.
|
||||
- `Mail "notify@home.box"`, `MailReport "on-change"` — Postfix is up but **`relayhost = (none)`**, so mail is delivered locally to `/var/mail/notify` (not your inbox until you wire a smart-host).
|
||||
|
||||
### Layer 2 — CT OS packages (Debian)
|
||||
|
||||
`unattended-upgrades` deployed inside every CT (CT 104 already had it; 101/102/103/105/108/110 added 2026-05-20):
|
||||
|
||||
| CT | u-u version | Distro | Status |
|
||||
|---|---|---|---|
|
||||
| 101 | 2.12 | trixie | active |
|
||||
| 102 | 2.9.1 | bookworm | active |
|
||||
| 103 | 2.9.1 | bookworm | active |
|
||||
| 104 | 2.12 | trixie | active |
|
||||
| 105 | 2.9.1 | bookworm | active |
|
||||
| 108 | 2.12 | trixie | active |
|
||||
| 110 | 2.12 | trixie | active |
|
||||
|
||||
Per-CT allowlist is Debian-only — third-party repos (`docker.com`, `jotta.cloud`, `claude.ai`, `cli.github.com`, `dl.k6.io`) are **excluded** because they ship breaking changes outside Debian's freeze. Upgrade those with explicit `apt upgrade <pkg>`.
|
||||
|
||||
### Layer 3 — Helper-script app binaries (AdGuard, Zoraxy)
|
||||
|
||||
Each helper-scripts CT ships `/usr/bin/update` that re-`curl|bash`'s the community-scripts installer. Replaced with proper systemd timers using the apps' own update mechanisms:
|
||||
|
||||
| CT | Timer | Schedule | Mechanism |
|
||||
|---|---|---|---|
|
||||
| 102 AdGuard | `adguard-update.timer` | Wed 03:30 (+15 m jitter) | native `AdGuardHome --update` flag |
|
||||
| 108 Zoraxy | `zoraxy-update.timer` | Wed 03:40 (+15 m jitter) | GitHub releases API, **stable semver only (skips RCs)**, binary swap + 30 s health check + auto-rollback |
|
||||
|
||||
Both log to `/var/log/{adguard,zoraxy}-update.log` and journal. Manual invoke: `systemctl start <name>-update.service`.
|
||||
|
||||
### Layer 4 — Docker engine inside CTs
|
||||
|
||||
`docker-ce` updates in CT 101, 104, 105, 110 — held by the Debian-only allowlist. Apply with `apt upgrade docker-ce docker-ce-cli containerd.io` when you want them. Add `origin=Docker` to the allowlist if you want to auto-apply (not recommended; engine updates occasionally break running containers).
|
||||
|
||||
### Layer 5 — Docker images (the ~75 containers)
|
||||
|
||||
**Plan: deploy Diun on CT 109** (observability LXC, see §16). Diun watches image tags on registries, posts to Gotify when a new image is available. Pulls remain manual (`docker compose pull && up -d`) — protects against latest-tag drift like the n8n incident pinned in `/opt/stacks/n8n/docker-compose.yaml`.
|
||||
|
||||
**Layer 6 — Nextcloud-AIO**: self-updates via the mastercontainer (CT 105). No external mechanism needed.
|
||||
|
||||
---
|
||||
|
||||
## 14. VM 100 (HAOS) auto-restart watchdog
|
||||
|
||||
Old approach: `*/5 * * * * /root/vm100.sh > /dev/null` in cron. Script archived to `/root/vm100.sh.bak` on 2026-05-20.
|
||||
|
||||
Replaced with a systemd timer + oneshot:
|
||||
|
||||
- `/usr/local/sbin/vm100-watchdog.sh` — only restarts on `status: stopped`; skips paused/prelaunch/transitional states; respects `/var/lock/qemu-server/lock-100.conf` so it doesn't race vzdump or migration
|
||||
- `vm100-watchdog.service` (Type=oneshot)
|
||||
- `vm100-watchdog.timer` (`OnUnitActiveSec=1min`, RandomizedDelaySec=15s)
|
||||
|
||||
Recovery latency improved from 5 min → 1 min; logging structured in `journalctl -u vm100-watchdog`.
|
||||
|
||||
---
|
||||
|
||||
## 15. Identity — Pocket-ID on its own CT
|
||||
|
||||
### State as of 2026-05-20
|
||||
|
||||
**CT 110 "id"** created at `192.168.1.5` as the dedicated IdP host. Pocket-ID was previously on CT 104 as one of ~65 docker containers; moved off because:
|
||||
|
||||
- Pocket-ID gates auth for everything else; sharing a host with image-gen workloads (9 GiB swap pressure on CT 104) means a single OOM/restart there knocks out SSO across the whole estate
|
||||
- Pattern match with CT 102 (DNS) and CT 108 (reverse proxy) — each infra-critical concern gets its own CT
|
||||
|
||||
| CT 110 setting | Value |
|
||||
|---|---|
|
||||
| Hostname / IP | `id` / `192.168.1.5` |
|
||||
| Cores / RAM / rootfs | 1 / 1 GB / 4 GB |
|
||||
| Privilege | unprivileged, nesting=1, keyctl=1 |
|
||||
| Boot order | `onboot=1`, `startup=order=3` (after DNS=1, Zoraxy=2) |
|
||||
| Protection | `protection: 1` |
|
||||
| Auto-updates | unattended-upgrades, Debian-only allowlist |
|
||||
| Docker | 29.5.1 + compose v5.1.3 |
|
||||
|
||||
### Duplication procedure used
|
||||
|
||||
1. `sqlite3 pocket-id.db ".backup /tmp/pi-snap/pocket-id.db"` on CT 104 (online, no downtime to id.nuclide.systems)
|
||||
2. tar everything except `*.db*`; restore the live snapshot as `pocket-id.db`
|
||||
3. `pct pull` → `pct push` to CT 110
|
||||
4. Adapted compose to drop the `shared_backend` external network reference (CT 110 uses default bridge)
|
||||
5. `docker compose up -d`
|
||||
6. Verified `http://192.168.1.5:11000/healthz` returns 200
|
||||
|
||||
### Zoraxy cutover
|
||||
|
||||
`id.nuclide.systems` upstream needs to change from `192.168.1.40:11000` → `192.168.1.5:11000`. Single-line config edit in Zoraxy + reload. Verified live in §11b once executed.
|
||||
|
||||
### Secrets-rotation list (deferred to cutover day)
|
||||
|
||||
- `OIDC_CLIENT_SECRET` for the Arcane registration (exposed in chat transcript): rotate in Pocket-ID UI, update Arcane env, restart Arcane
|
||||
- `ENCRYPTION_KEY` and `JWT_SECRET` in `/opt/stacks/arcane/docker-compose.yml`: move to `.env` (currently empty), regenerate, restart Arcane. Existing user sessions get invalidated — fine, ask everyone to log in again
|
||||
|
||||
---
|
||||
|
||||
## 16. CT 109 "observe" — planned observability + ops LXC
|
||||
|
||||
Single LXC holding everything monitoring/ops-shaped. Sizing target: **4 cores / 6 GiB RAM / 30 GiB rootfs**, unprivileged, nesting=1, IP `192.168.1.6` (next free in the infra range after CT 110).
|
||||
|
||||
### Stack to deploy on CT 109
|
||||
|
||||
| Service | Purpose |
|
||||
|---|---|
|
||||
| Arcane Manager | central docker management UI; edge agents on CT 101 + CT 104 + CT 110 (mTLS, agent-dialed-out) |
|
||||
| Dozzle UI | aggregated container logs; Dozzle agents on CT 101 + CT 104 + CT 110 |
|
||||
| Prometheus | metrics TSDB, 30 d retention |
|
||||
| Grafana | dashboards over Prometheus |
|
||||
| Alertmanager + alertmanager-gotify-bridge | alert routing → Gotify |
|
||||
| **Homarr** (selected over Homepage) | unified dashboard, native OIDC with Pocket-ID, native Prometheus widget + Grafana iframe support |
|
||||
| Diun | docker image update notifier → Gotify |
|
||||
| Tinyauth | forward-auth gate for non-OIDC apps (Backrest, raw Dozzle, raw Prometheus, raw Grafana). Speaks OIDC client to Pocket-ID, returns Basic-Auth/Remote-User header to upstream |
|
||||
| docker-socket-proxy | local + remote (CT 101/104/110) — hardened read-only docker.sock for Homarr discovery |
|
||||
|
||||
### Sidecars deployed on each Docker host
|
||||
|
||||
| Host | Sidecars |
|
||||
|---|---|
|
||||
| Host (PVE) | node-exporter, smartctl-exporter, pve-exporter (host metrics into Prometheus) |
|
||||
| CT 101 | node-exporter, cAdvisor, Dozzle agent, Arcane edge agent, docker-socket-proxy |
|
||||
| CT 104 | node-exporter, cAdvisor, Dozzle agent, Arcane edge agent, docker-socket-proxy |
|
||||
| CT 110 | node-exporter, Dozzle agent, Arcane edge agent (test target — small/clean) |
|
||||
|
||||
### Services that stay where they are (NOT on CT 109)
|
||||
|
||||
- AdGuard (CT 102), Zoraxy (CT 108), Backrest (CT 103), Pocket-ID (CT 110) — each has its own CT for blast-radius reasons
|
||||
- All app stacks (Immich, Nextcloud, Shepard, n8n, ComfyUI, etc.)
|
||||
|
||||
### Migration of Gotify (Tier B — schedule when CT 109 is otherwise stable)
|
||||
|
||||
Gotify currently runs in CT 104 docker stack at `gotify.nuclide.systems`. Moving to CT 109 isolates alerting from CT 104 outages but means updating env vars / webhook targets in ~10 places (MCP servers, Backrest webhooks). Plan: copy DB and app tokens via volume tar, deploy on CT 109, update Zoraxy upstream, then sweep dependents.
|
||||
|
||||
### Migration of MCP Gateway (Tier B — split control / workload)
|
||||
|
||||
`/opt/stacks/ai/mcp-gateway/` on CT 104 is the OIDC-gated control plane for the ~20 MCP child containers. Decision (2026-05-20): move **only the gateway** to CT 109; **the child MCP containers stay on CT 104** (they're workload, not control plane).
|
||||
|
||||
Mechanics:
|
||||
- Gateway on CT 109 uses `DOCKER_HOST=tcp://<CT 104 socket-proxy>:2375` (the docker-socket-proxy already planned for CT 104) instead of the bind-mounted `/var/run/docker.sock`. socket-proxy ACL must allow `containers, exec, images` (read+write).
|
||||
- Migrate state files via volume tar: `config.json`, `agents.json`, `prompts/`, `usage.db`, `gateway_tokens.json`, `nc_user_creds.json`. Keep them on CT 109 local zfs, not UNAS (per-request latency matters).
|
||||
- Pocket-ID redirect URI stays `https://mcp.nuclide.systems/sso/callback` — only Zoraxy's upstream flips from `192.168.1.40:8080` to the CT 109 IP.
|
||||
- Don't touch the children; gateway still spawns them by name against CT 104's daemon.
|
||||
|
||||
Build-order slot: after Arcane + socket-proxy land in §16's checklist, before Tinyauth.
|
||||
|
||||
### Arcane specifics
|
||||
|
||||
- Existing `/opt/stacks/arcane/` on CT 104 has the Manager 80 % built: 41 MB SQLite DB carrying Pocket-ID OIDC client + admin user `fkrebs@nucli.de`
|
||||
- Container currently stopped (~since May 19)
|
||||
- **Edge mode** chosen for agents — agent dials out to Manager over HTTPS/gRPC with mTLS. Manager auto-generates the edge CA on first start; each agent enrolls with a token and gets a 1 yr client cert with auto-renewal
|
||||
- Same `ghcr.io/getarcaneapp/arcane:latest` image for Manager and agents; mode is env-driven (`ARCANE_EDGE_AGENT=true` on agents)
|
||||
- Intel iGPU is not in Arcane's GPU monitoring catalog (NVIDIA + AMD only). LXC-level passthrough to CT 104 stays untouched; only Arcane's GPU widgets will be blank for the Arc
|
||||
|
||||
### Build order
|
||||
|
||||
1. Create CT 109 (specs above)
|
||||
2. Deploy node-exporter on host + Prometheus + Grafana first (start collecting baselines)
|
||||
3. Deploy Arcane Manager with the existing DB restored from CT 104
|
||||
4. Issue agent tokens, deploy Edge agents on CT 101/104/110
|
||||
5. Deploy Dozzle Manager + agents
|
||||
6. Deploy Diun
|
||||
7. Deploy Tinyauth (configure Pocket-ID client first)
|
||||
8. Deploy Homarr
|
||||
9. Front the lot via Zoraxy: `arcane.`, `dozzle.`, `grafana.`, `prom.`, `home.` `.nuclide.systems`
|
||||
10. Verify each through the Tinyauth gate where applicable
|
||||
11. Add Diun watchlist + Alertmanager routing → Gotify
|
||||
|
||||
---
|
||||
|
||||
## 17. LLDP / UniFi topology visibility
|
||||
|
||||
Investigated 2026-05-20:
|
||||
|
||||
- `lldpd` is **installed and active** on the host (defaults: advertises + listens)
|
||||
- Existing `/usr/local/bin/update_interface_desc.sh` hourly cron consumes LLDP from neighbors and writes `# PortDescr:` comments into `/etc/network/interfaces`
|
||||
- Host is on switch **port 10** of the D-Link DGS-1210-28P (`SysName: dgs1210`, FW `6.32.008`).
|
||||
- `/usr/local/bin/update_interface_desc.sh` was silently failing — `lldpcli` not in cron's PATH **and** the original logic only appended new lines, never replaced stale ones. Rewritten 2026-05-20:
|
||||
- PATH set inside the script (`/usr/sbin` included)
|
||||
- Uses `lldpcli show neighbors -f keyvalue` for machine-readable parsing
|
||||
- Single `# LLDP: <chassis> :: <port-descr>` line per interface; legacy `# PortDescr:` lines stripped
|
||||
- Idempotent (second run = no-op, no backup churn)
|
||||
- Daily-keyed backup `/etc/network/interfaces.bak.YYYYMMDD`
|
||||
- Only logs on actual change
|
||||
- Old script preserved at `/usr/local/bin/update_interface_desc.sh.bak-2026-05-20`
|
||||
|
||||
**The UDM Pro won't see "nuc" in its topology view** because LLDP frames use the Nearest-Bridge multicast (`01:80:c2:00:00:0e`) which any 802.1D-compliant switch terminates by spec — and the D-Link DGS-1210-28P sits between the host and the UDM. LLDP-MED is **not** a fix for this; it's for endpoint (VoIP/MFP) discovery, not transparent LLDP forwarding.
|
||||
|
||||
Remediation paths in order of effort:
|
||||
|
||||
1. **Enable SNMP v2c/v3 on the DGS-1210** + add as a Generic SNMP device in UniFi → UDM sees the switch and can map port↔MAC. Most practical for this stack.
|
||||
2. **Add `snmpd` to the host** + generic SNMP device in UniFi → CPU/mem/iface stats from Proxmox visible in UniFi (not in topology, but in monitoring).
|
||||
3. **UniFi-managed switch between host and UDM** — clean answer; requires hardware.
|
||||
|
||||
D-Link DGS-1210 admin UI lives at `http://192.168.1.10/` (login page, field names `pelican` / `pinkpanther`). Verify the admin password is non-default — DGS-1210 ships with `admin`/blank or `admin`/`admin` on most firmware revisions. A flat-LAN switch with default creds is one of the easier vectors.
|
||||
|
||||
Useful inspection from the host any time: `lldpcli show neighbors`.
|
||||
|
||||
---
|
||||
|
||||
## 18a. UNAS access — UID consistency model (post-NFSv4 investigation)
|
||||
|
||||
**Investigation result (2026-05-20):** UNAS Pro advertises NFSv4 in `rpcinfo` but has no v4 export tree configured. Every v4 mount attempt returns `No such file or directory`. Ubiquiti has **not** announced NFSv4 support and the [community thread asking for it](https://community.ui.com/questions/When-will-NFSv4-be-available-for-UNAS-Pro/b8f8b00b-c5ac-4e19-adc4-9e5cae06d35f) has no ETA. The official help center also confirms: *"UniFi Drive does not support certain NFS export options, such as `no_root_squash`"*. So root_squash + v3-only is the long-term reality.
|
||||
|
||||
### Universal UID landscape on UNAS
|
||||
|
||||
Every NFS client write lands as **uid 977 / gid 988** (UNAS's `all_squash` + `anon_uid=977`/`anon_gid=988`). The chown probe confirmed no client can change this from the host side. Files written via the legacy CIFS mount appear as **uid 33** to the CIFS client *but* are stored differently on UNAS — the CIFS `forceuid=33` mount option lies about ownership client-side.
|
||||
|
||||
### Per-CT access pattern (canonical)
|
||||
|
||||
| CT | Mount | Container uid | Effective on disk | Status |
|
||||
|---|---|---|---|---|
|
||||
| 103 backrest | NFS bind from host | root | squashes to 977 | ✓ consistent |
|
||||
| 104 docker | NFS bind from host | mostly root, n8n=1000 (latent) | squashes to 977 | ✓ for root containers; n8n latent if it ever writes to UNAS |
|
||||
| 105 nextcloud | CIFS today (forceuid=33) → **NFS + bindfs** target | uid 33 (www-data) inside Nextcloud, bindfs translates to 977 on disk | needs migration | 🚨 still CIFS |
|
||||
|
||||
### Convention for new containers
|
||||
|
||||
Set `PUID=977 PGID=988` on any container that writes to UNAS. This pre-aligns with UNAS's enforced mapping and avoids latent permission issues (the n8n class). For images that don't support PUID/PGID, run them as root inside the container — root squashes to 977 cleanly.
|
||||
|
||||
### Why bindfs for CT 105 specifically
|
||||
|
||||
Nextcloud's PHP code hard-checks file ownership against `www-data` (uid 33). Without remap, NFS reads return uid 977 and Nextcloud refuses to operate normally. CIFS hides this with `forceuid=33`. NFS+bindfs achieves the same lie with the much faster NFS rail underneath — verified ~5× speed-up on metadata-heavy ops in the non-destructive test on 2026-05-20.
|
||||
|
||||
### Trigger event to revisit
|
||||
|
||||
Watch [community.ui.com/RELEASES](https://community.ui.com/RELEASES) for a UniFi Drive release that adds:
|
||||
- NFSv4 export option (would enable idmap)
|
||||
- `no_root_squash` support (would enable server-side chown to specific uids)
|
||||
- Configurable `anonuid`/`anongid` (would let us match a real uid)
|
||||
|
||||
Any of these would let us simplify the CT 105 stack.
|
||||
|
||||
---
|
||||
|
||||
## 18. Homarr inventory — services to include on the dashboard
|
||||
|
||||
Captured here so the eventual Homarr config can be assembled in one pass. Groups follow the existing `homepage.*` label convention used in compose files.
|
||||
|
||||
### Group: `infrastructure`
|
||||
|
||||
| Service | URL | Notes |
|
||||
|---|---|---|
|
||||
| Proxmox UI | `https://192.168.1.20:8006` | until LE via Zoraxy lands, see §11b |
|
||||
| AdGuard Home (CT 102) | `http://192.168.1.2/` (UI on `:80`) | DNS + admin |
|
||||
| Zoraxy (CT 108) | `http://192.168.1.4:8000/` | reverse proxy admin |
|
||||
| Backrest (CT 103) | `http://192.168.1.3:9898/` | backup orchestration, will be fronted via Tinyauth + Zoraxy |
|
||||
| **Pocket-ID (CT 110)** | `https://id.nuclide.systems/` | new home, 2026-05-20 |
|
||||
|
||||
### Group: `network`
|
||||
|
||||
| Service | URL | Notes |
|
||||
|---|---|---|
|
||||
| UDM Pro | `https://192.168.1.1/` | UniFi controller |
|
||||
| **D-Link DGS-1210-28P** | `http://192.168.1.10/` | core L2 switch; host on port 10 |
|
||||
| UNAS Pro | `https://192.168.1.31/` | UniFi NAS |
|
||||
|
||||
### Group: `observability` (to populate when CT 109 lands)
|
||||
|
||||
| Service | URL |
|
||||
|---|---|
|
||||
| Grafana | `https://grafana.nuclide.systems` |
|
||||
| Arcane | `https://arcane.nuclide.systems` |
|
||||
| Dozzle | `https://dozzle.nuclide.systems` |
|
||||
| Prometheus | `https://prom.nuclide.systems` (gated by Tinyauth) |
|
||||
| Alertmanager | `https://alerts.nuclide.systems` (gated by Tinyauth) |
|
||||
|
||||
### Group: `apps` (subset — long list, fill from existing `homepage.*` labels in `/opt/stacks/*/`)
|
||||
|
||||
Immich, Nextcloud, Vaultwarden, Karakeep, Memos, Paperless-ngx, n8n, ComfyUI, LobeChat, LiteLLM, Traccar, Gotify, Speaches, Daytona, Searxng, Kroki, etc. Pull display labels and icons from the existing `homepage.name=` / `homepage.icon=` values per compose.
|
||||
|
||||
---
|
||||
|
||||
## 12. Suggested action order
|
||||
|
||||
1. **Done 2026-05-20** ✅:
|
||||
- **Memory / sizing**: CT 101 160 → 32 GiB; CT 102 512 MiB → 1 GiB; HAOS balloon = 4 GiB
|
||||
- **Protection**: CT 102 startup=1+protection; CT 108 startup=2+protection; CT 103/104/105/110/VM100 protection
|
||||
- **CT 110 (id)** built at 192.168.1.5, Pocket-ID duplicated (online SQLite snapshot)
|
||||
- **ZFS**: `autotrim=on`, `atime=off rpool`
|
||||
- **Backups**: retention set on `unas`
|
||||
- **APT**: duplicate sources removed; host `unattended-upgrades` deployed (Debian-only)
|
||||
- **Per-CT u-u**: deployed to all 7 CTs with Debian-only allowlist
|
||||
- **Cron cleanup**: removed weekly tteck-LXC-update curl-pipe-bash; removed daily broken `apt-get upgrade -y`; replaced `/root/vm100.sh` cron with `vm100-watchdog.timer` (1 min, lock-aware)
|
||||
- **Self-updaters**: AdGuard (`--update` flag) Wed 03:30; Zoraxy (GitHub stable releases + rollback) Wed 03:40
|
||||
- **`zpool upgrade rpool`** ran during the audit (enabled `redaction_list_spill`, `raidz_expansion`)
|
||||
2. **Today with a maintenance window**:
|
||||
- **Cut Zoraxy over to CT 110** for `id.nuclide.systems` (single upstream edit; rollback path = revert one line)
|
||||
- `apt full-upgrade` (kernel 7.0.0-3 → 7.0.2-5, pve-manager 9.1.11 → 9.1.18) + reboot
|
||||
- VM 100 disk options (`cache=none`, `iothread=1`) — §4. Requires VM stop/start.
|
||||
- Switch CT 105 from CIFS → NFS bind-mount (§4). Test UID mapping.
|
||||
- Drop the `unas_smb` storage once CT 105 is migrated.
|
||||
3. **This week**:
|
||||
- Re-measure CT 104 peak RSS after fixing CT 101 — likely safe to drop to 48 GiB.
|
||||
- Convert CT 105 to unprivileged (backup → restore as unprivileged).
|
||||
- Raise ARC cap to 16 GiB.
|
||||
- Probe NFSv4 against UNAS; switch if supported.
|
||||
- **Wire Postfix relayhost** (Gmail/Postmark/your SMTP) so unattended-upgrades + zfs-zed + cron failures actually mail you.
|
||||
- **Rotate Backrest plan** to back up real data (currently still pointed at `/media/data-dir` — a 50 KB test file from August 2025); see §7.
|
||||
4. **Medium term**:
|
||||
- **Build CT 109** observability LXC (§16) — Prometheus + Grafana + Arcane Manager + Dozzle + Homarr + Diun + Tinyauth
|
||||
- Migrate Gotify from CT 104 to CT 109 (~30 min of env-var updates)
|
||||
- **Rotate exposed secrets** that appeared in this transcript: Arcane `OIDC_CLIENT_SECRET`, `ENCRYPTION_KEY`, `JWT_SECRET`; Immich `IMMICH_API_KEY`
|
||||
- Backrest: enable auth, redesign plans to cover all data tiers (§7), front via Tinyauth for OIDC
|
||||
5. **Next purchase window**:
|
||||
- Second NVMe → mirror `rpool` (§6)
|
||||
- `proxmox-boot-tool init` on the new disk
|
||||
- 2.5 GbE NIC + matching switch port to UNAS for image-gen / backup speed
|
||||
- UniFi-managed switch between host and UDM (or accept SNMP-only visibility from UniFi)
|
||||
|
||||
---
|
||||
|
||||
---
|
||||
|
||||
## 19. Changes applied 2026-05-20 (session 2)
|
||||
|
||||
### Optimizations executed
|
||||
|
||||
| # | Item | Command / action | Result |
|
||||
|---|------|-----------------|--------|
|
||||
| 1 | CT 101 protection + boot order | `pct set 101 -protection 1 -startup order=10` | ✅ |
|
||||
| 4 | CT 104 memory cap 128→48 GiB | `pct set 104 -memory 49152` | ✅ (swap 32→8 deferred: still 10.4 GB in use) |
|
||||
| 5 | CT 102 rootfs 2→4 GiB | `pct resize 102 rootfs 4G` | ✅ now 28% used |
|
||||
| A | VM 100 disk: cache=none + iothread=1 | `qm set 100 -scsi0 ...,cache=none,iothread=1` + `scsihw virtio-scsi-single` | ✅ HAOS healthy |
|
||||
| C | Host apt full-upgrade | kernel 7.0.0→7.0.2-5, pve-manager 9.1.11→9.1.18 | ✅ installed; **reboot pending** |
|
||||
|
||||
### Pocket-ID migration completed
|
||||
|
||||
- `id.nuclide.systems` Zoraxy proxy cutover confirmed: `192.168.1.40:11000` → `192.168.1.5:11000`
|
||||
- CT 104 `/opt/stacks/pocketid/` directory fully removed (data migrated to CT 110 2026-05-20)
|
||||
- Backup: `/opt/zoraxy/conf/proxy/id.nuclide.systems.config.bak-pre-ct110` (keep as rollback)
|
||||
|
||||
### Proxmox OIDC via Pocket-ID
|
||||
|
||||
Realm `pocket-id` added; user `fkrebs@nucli.de@pocket-id` mapped to Administrator role.
|
||||
|
||||
```bash
|
||||
pveum realm add pocket-id \
|
||||
--type openid \
|
||||
--issuer-url https://id.nuclide.systems \
|
||||
--client-id 38469e7e-1fff-4841-83a9-74bf38d847eb \
|
||||
--client-key <secret> \
|
||||
--username-claim email \
|
||||
--comment "Pocket-ID OIDC"
|
||||
|
||||
pveum user add fkrebs@nucli.de@pocket-id
|
||||
pveum aclmod / --users fkrebs@nucli.de@pocket-id --roles Administrator
|
||||
```
|
||||
|
||||
OIDC client inserted directly into Pocket-ID SQLite (API key stored as SHA-256 hash — not reversible):
|
||||
|
||||
```
|
||||
DB: /opt/stacks/pocketid/data/pocket-id.db on CT 110
|
||||
Table: oidc_clients
|
||||
client_id: 38469e7e-1fff-4841-83a9-74bf38d847eb
|
||||
name: Proxmox VE
|
||||
callback_urls: ["https://192.168.1.20:8006"]
|
||||
```
|
||||
|
||||
**To add future OIDC clients without UI access:**
|
||||
```bash
|
||||
python3 -c "
|
||||
import uuid, secrets, bcrypt, json, datetime
|
||||
client_id = str(uuid.uuid4())
|
||||
secret_plain = secrets.token_urlsafe(32)
|
||||
secret_hash = bcrypt.hashpw(secret_plain.encode(), bcrypt.gensalt(rounds=10)).decode()
|
||||
print(f'id={client_id}')
|
||||
print(f'secret={secret_plain}')
|
||||
print(f'hash={secret_hash}')
|
||||
"
|
||||
# Then INSERT into oidc_clients with the hash, use secret_plain in the app config
|
||||
# callback_urls and logout_callback_urls are JSON arrays stored as BLOB
|
||||
# credentials field is '{}' for standard clients
|
||||
```
|
||||
|
||||
**Note on Pocket-ID API keys:** The `key` column in `api_keys` stores a SHA-256 hash of the real key (64-char hex). The plaintext key is only shown once at creation time in the UI. If lost, create a new one — there is no recovery path.
|
||||
|
||||
**Login flow:** In PVE web UI, select realm `pocket-id` at login. You will be redirected to `https://id.nuclide.systems` for authentication and returned to PVE. The `email` claim is used as the PVE username.
|
||||
|
||||
---
|
||||
|
||||
## Audit footnote — side effects of this run
|
||||
|
||||
- `zpool upgrade rpool` was executed (not `-n`). Enabled features: `redaction_list_spill`, `raidz_expansion`. Safe on current ZFS version; the pool can no longer be imported by ZFS releases that pre-date these features. No data risk.
|
||||
- No other state changes were made by this audit.
|
||||
Reference in New Issue
Block a user