commit 04d54ac9b057fca182f5e081ae5433085978df64 Author: fkrebs (via Claude) Date: Wed May 20 23:23:31 2026 +0200 baseline: docs as of pre-reorg audit (2026-05-20) diff --git a/README.md b/README.md new file mode 100644 index 0000000..0d8d37b --- /dev/null +++ b/README.md @@ -0,0 +1,85 @@ +# nuclide.systems — Infrastructure Docs + +Central documentation for the `nuc` Proxmox homelab and all hosted services. +**Source of truth:** `/docs/` on the Proxmox host (`192.168.1.20`). +This directory is the entry point — docs generated anywhere else should land here. + +--- + +## Infrastructure + +| File | Contents | +|------|----------| +| [proxmox-optimizations.md](proxmox-optimizations.md) | Full audit, sizing decisions, action log, pending tasks | + +--- + +## Docker Stacks (CT 104 / 192.168.1.40) + +### Service docs (`stacks/docs/`) + +| File | Contents | +|------|----------| +| [homelab-architecture.md](stacks/docs/homelab-architecture.md) | Overall architecture overview | +| [mcp-gateway.md](stacks/docs/mcp-gateway.md) | MCP gateway design and endpoints | +| [mcp-gateway-requirements.md](stacks/docs/mcp-gateway-requirements.md) | MCP gateway requirements | +| [zoraxy-setup.md](stacks/docs/zoraxy-setup.md) | Zoraxy reverse proxy setup notes | +| [AdGuard-DNS-writes.md](stacks/docs/AdGuard-DNS-writes.md) | AdGuard DNS rewrite rules | +| [comfyui-mcp.md](stacks/docs/comfyui-mcp.md) | ComfyUI MCP integration | +| [dev-environment.md](stacks/docs/dev-environment.md) | CT 111 Coder + Gitea dev platform | +| [traefik-migration.md](stacks/docs/traefik-migration.md) | Traefik → Zoraxy migration notes | +| [traefik-migration-docker-labels.md](stacks/docs/traefik-migration-docker-labels.md) | Label mapping reference | +| [cloud-gpu-extension.md](stacks/docs/cloud-gpu-extension.md) | Cloud GPU burst setup | +| [scrubbing-list.md](stacks/docs/scrubbing-list.md) | Services to clean up / decommission | +| [stack-ideas.md](stacks/docs/stack-ideas.md) | Backlog of stack ideas | +| [case-study.md](stacks/docs/case-study.md) | Case study writeup | + +### Operational (`stacks/`) + +| File | Contents | +|------|----------| +| [PORTMAP.md](stacks/PORTMAP.md) | Port assignments for all services | +| [storage.md](stacks/storage.md) | Volume and storage layout | +| [volumes.md](stacks/volumes.md) | Docker volume inventory | +| [docker-networks.md](stacks/docker-networks.md) | Docker network topology | +| [todo.md](stacks/todo.md) | Outstanding stack tasks | +| [CLAUDE.md](stacks/CLAUDE.md) | Agent breadcrumbs (migrated stacks, dormant services) | + +--- + +## Nexa Project (`stacks/nexa/`) + +| File | Contents | +|------|----------| +| [README.md](stacks/nexa/README.md) | Project overview | +| [docs/01-vision-and-scope.md](stacks/nexa/docs/01-vision-and-scope.md) | Vision and scope | +| [docs/02-roadmap.md](stacks/nexa/docs/02-roadmap.md) | Roadmap | +| [docs/03-architecture.md](stacks/nexa/docs/03-architecture.md) | Architecture | +| [docs/04-integration-matrix.md](stacks/nexa/docs/04-integration-matrix.md) | Integration matrix | +| [docs/05-command-system.md](stacks/nexa/docs/05-command-system.md) | Command system | +| [docs/06-classification-logic.md](stacks/nexa/docs/06-classification-logic.md) | Classification logic | +| [docs/07-workflow-spec.md](stacks/nexa/docs/07-workflow-spec.md) | Workflow spec | +| [docs/08-graphrag-architecture.md](stacks/nexa/docs/08-graphrag-architecture.md) | GraphRAG architecture | +| [docs/09-deployment.md](stacks/nexa/docs/09-deployment.md) | Deployment | +| [docs/10-operations.md](stacks/nexa/docs/10-operations.md) | Operations | +| [docs/11-open-questions.md](stacks/nexa/docs/11-open-questions.md) | Open questions | +| [docs/12-optimization-opportunities.md](stacks/nexa/docs/12-optimization-opportunities.md) | Optimization opportunities | +| [docs/13-information-wishlist.md](stacks/nexa/docs/13-information-wishlist.md) | Information wishlist | + +--- + +## AI Agents (`stacks/ai/`) + +| File | Contents | +|------|----------| +| [home-info.md](stacks/ai/home-info.md) | Home context for agents | +| [morning-briefing-agent.md](stacks/ai/morning-briefing-agent.md) | Morning briefing agent config | +| [PROMPTING-2026.md](stacks/ai/PROMPTING-2026.md) | Prompting guide 2026 | + +--- + +## Contributing + +- Drop new docs in the appropriate subdirectory +- Update this README when adding a new section +- When CT 109 (management LXC) is built, the docs platform will serve this directory diff --git a/proxmox-optimizations.md b/proxmox-optimizations.md new file mode 100644 index 0000000..1802c6d --- /dev/null +++ b/proxmox-optimizations.md @@ -0,0 +1,999 @@ +# Proxmox Host Optimization Inventory — `nuc` + +Generated: 2026-05-20 +Host: `nuc` · PVE `9.1.11` · Kernel `6.17.13-4-pve` · Debian 13 (trixie) +CPU: Intel Core Ultra 7 155H (16C / 22T, hybrid P+E+LP-E) · 1 socket · 1 NUMA +RAM: 62 GiB physical · 31 GiB zram swap (50 % of RAM, zstd, prio 100) +Storage: single Crucial P3 2 TB NVMe (QLC, DRAM-less) → `rpool` (ZFS, ashift=12, no redundancy) +Workload: 1 VM (HAOS) + 6 LXCs (Docker, Adguard, Backrest, Nextcloud, Zoraxy, Dockge) + +--- + +## TL;DR — top 5 actionable wins + +1. **Memory overcommit is dangerous.** Allocated guest RAM (≈ 290 GiB) is ~4.7× physical (62 GiB). Right-size CT 101 (was 160 → done, now 32) and CT 104 (still 128) — see §2. ✅ *applied 2026-05-20* +2. **ZFS ARC is artificially capped at 6.2 GiB.** Default would be ~31 GiB. After §1 settles, raise to 16 GiB. See §3. +3. **No redundancy on a QLC SSD with 19 % wear and 59 TB written.** Single-disk `rpool` on a DRAM-less consumer QLC drive is a SPOF. Add a second NVMe and convert to mirror — biggest reliability win available. See §6. +4. **Backups never prune.** Was configured `keep-all=1` — fixed to `keep-last=3,keep-daily=7,keep-weekly=4,keep-monthly=6`. See §7. ✅ *applied 2026-05-20* +5. **`atime` and `autotrim`** on ZFS. ✅ *applied 2026-05-20* +6. **No DNS rewrites** in AdGuard — every internal target is IP-only; add a split-horizon for `nuclide.systems` and a `.lan` shorthand set. See §11a. +7. **Self-signed Proxmox web UI cert** — front via Zoraxy for free LE. See §11b. + +--- + +## 1. System snapshot + +| Resource | State | Notes | +|---|---|---| +| Load avg | normal | PSI: CPU some=2.4 % / IO some=1 % over 60 s | +| Memory | 52 / 62 GiB used, 4.5 GiB free | tight; zram swap 15 GiB in use | +| Swap | zram0 (zstd, 31 GiB) prio 100 | working as designed; just a symptom of §1 | +| ARC | 6.0 / 6.2 GiB (capped) | hit ratio ~99 % but cap is far below default | +| NVMe wear | **Percentage Used 19 %**, 59.2 TB written | ~5 % wear/year at current rate; healthy for now | +| Temperature | 56–58 °C | well under the 95 °C critical threshold | +| Uptime | (see `uptime`) | scrub clean, no checksum errors | +| Cluster | standalone | quorum OK, no HA configured | + +--- + +## 2. Memory & VM/CT sizing (measured numbers) + +Read from `/sys/fs/cgroup/lxc//memory.{current,peak,max}` and `free -h` inside each guest: + +| Guest | Cap | Current | Peak | Swap-in-use | Verdict | +|---|---|---|---|---|---| +| VM 100 (haos) | 16 384 MiB | 13 056 MiB | n/a | n/a | balloon disabled; HAOS actually uses what it has | +| CT 101 (shepard) | 160 000 MiB | **7.2 GiB** | **15.6 GiB** | 706 MiB | **wildly over-sized** — peak is 10 % of cap | +| CT 102 (adguard) | 512 MiB | 343 MiB | **509 MiB (99 %)** | 19 MiB | **under-sized** — at the cap, AdGuardHome alone is 350 MiB | +| CT 103 (backrest) | 512 MiB | 89 MiB | 305 MiB | 16 MiB | fine | +| CT 104 (docker/AI) | 128 000 MiB | 18.8 GiB | **29.3 GiB** | **9.3 GiB** | real workload, but currently swapping — likely starved by CT 101 | +| CT 105 (nextcloud) | 8 192 MiB | 2.1 GiB | 3.5 GiB | 53 MiB | fine | +| CT 108 (zoraxy) | 2 048 MiB | 271 MiB | 463 MiB | 25 MiB | fine; could halve | + +**Sum of declared caps ≈ 290 GiB on a 62 GiB host.** Sum of actual peaks ≈ 49 GiB — totally fits. CT 101's 160 GB cap is the entire problem: it's a phantom that scares the scheduler without using anything close to that. + +### Concrete CT 101 picture + +12 cores, load avg 8.5, ~9 Docker containers (Shepard frontend/backend, Keycloak, Neo4j, MongoDB, MongoExpress, TimescaleDB, Caddy, home-showcase-collector). Peak RSS 15.6 GiB. + +→ Drop memory cap to **32 GiB** (2× peak headroom). No reboot required for LXC memory changes. + +### Concrete CT 104 picture + +16 cores, load avg 8.0, ~65 Docker containers including Immich (with ML/vectorchord), ComfyUI (image-gen), LobeChat, n8n, Daytona, LiteLLM, Vaultwarden, Paperless-ngx+AI, Karakeep, Memos, Gotify, Garage S3, plus a forest of MCP servers, Speaches (OpenVINO using the Arc iGPU). 128 GiB of 200 GiB rootfs used. + +Peak RSS 29.3 GiB, but **9.3 GiB sitting in swap** — under memory pressure. Two paths: +1. **Recommended**: cut CT 101 first, then CT 104's pressure mostly disappears on its own. Re-measure peak after CT 101 is fixed. Likely safe to cap at **48 GiB** then. +2. Leave the 128 GiB cap as a generous ceiling — harmless once CT 101 is sane. + +### Other guests + +- **CT 102 (AdGuard)**: hitting its cap. Bump to **1 GiB** to give DNS room to breathe. (DNS is the single point of failure for the LAN — don't squeeze it.) `pct set 102 -memory 1024`. +- **CT 108 (Zoraxy)**: cap **1 GiB** is plenty. +- **VM 100 (HAOS)**: 13 GiB actually used. Enable ballooning so the host can reclaim idle RAM: + ```bash + qm set 100 -balloon 4096 + ``` + This leaves memory=16384 as a ceiling but lets the host shrink it under pressure. +- KSM is already active (10 255 shared pages) — fine. +- zram swap: keep size at `ram / 2` until CT 101 is fixed; reduce to `ram / 4` afterwards. + +--- + +## 3. ZFS tuning + +### Pool + +| Setting | Current | Recommend | Why | +|---|---|---|---| +| `autotrim` | **off** | `on` | QLC needs trim; weekly fstrim alone is OK but autotrim is "free" | +| `ashift` | 12 | keep | correct for NVMe | +| `atime` | **on** (relatime) | **off** | unused on a hypervisor; reduces write amp on QLC | +| `xattr` | sa | keep | already optimal | +| `compression` | on (lz4) | keep | helping (1.61× on HAOS disk) | +| `dnodesize` | legacy | `auto` | minor; only matters with millions of small files | +| `recordsize` (rpool) | 128 K | keep for general | tune per-dataset (see below) | + +### ARC + +`/etc/modprobe.d/zfs.conf` currently caps `zfs_arc_max=6669991936` (≈ 6.2 GiB). +- After §2 sizing is done, raise this to 16 GiB: `options zfs zfs_arc_max=17179869184` and `zfs_arc_min=4294967296`. +- Apply live without reboot: `echo 17179869184 > /sys/module/zfs/parameters/zfs_arc_max`. + +### Per-dataset + +- **VM zvols** (`rpool/data/vm-*`): default `volblocksize` is 16 K — fine. HAOS disk uses `cache=writethrough`; on ZFS, **switch to `cache=none`** (or unset) — writethrough doubles the sync cost on top of ZFS's own integrity guarantees. +- **CT 104 (`subvol-104-disk-0`, Docker + image-gen)**: keep `recordsize=128K`. The workload is dominated by large model files and image outputs, not small-file DB traffic — shrinking the record size would hurt, not help. +- **Nextcloud (`subvol-105-disk-1`)**: leave at 128 K (mixed sizes, mostly larger files). + +### Pool features + +`zpool upgrade rpool` was run during this audit and enabled `redaction_list_spill` + `raidz_expansion`. Other disabled features (`fast_dedup`, `longname`, `large_microzap`, `dynamic_gang_header`, `block_cloning_endian`, `physical_rewrite`) can be enabled with another `zpool upgrade rpool` — only do this if you do not need to roll back to an older ZFS. + +### Commands + +```bash +zpool set autotrim=on rpool +zfs set atime=off rpool +# (optional, once memory is sane): +echo 'options zfs zfs_arc_max=17179869184' > /etc/modprobe.d/zfs.conf +update-initramfs -u -k all +``` + +--- + +## 4. Storage & VM disk options + +### VM 100 (`haos`) + +```diff +- scsi0: local-zfs:vm-100-disk-1,cache=writethrough,discard=on,size=32G,ssd=1 ++ scsi0: local-zfs:vm-100-disk-1,cache=none,discard=on,iothread=1,size=32G,ssd=1 +``` + +- `cache=none` (or remove cache entirely) — let ZFS manage caching. +- `iothread=1` with `virtio-scsi-pci` controller — already using `virtio-scsi-pci`, just add iothread. +- Already has `discard=on` and `ssd=1` ✔ + +### LXC `local-zfs` storage + +- `sparse 1` is set ✔ — thin-provisioned. +- All CTs are on `local-zfs` rootfs; OK. + +### UNAS share — current state (measured) + +Backend: `192.168.1.31` (looks like a UniFi NAS — exports `/volume/.../.unifi-drive/storage/.data`, the only NFS export listed is restricted to four allowed clients: the host `.20`, CT 104 `.40`, plus `.60` and `172.30.33.1`). + +Two parallel mounts on the host pointed at the **same backing data**: + +| Mount | Type | Options (key bits) | Consumers | +|---|---|---|---| +| `/mnt/pve/unas` | NFS **v3** | `proto=tcp, mountproto=udp, rsize/wsize=1M, hard, relatime, timeo=600` | CT 103 (backrest), CT 104 (docker) — bind-mounted to `/mnt/pve/unas` inside | +| `/mnt/pve/unas_smb` | CIFS **v3.1.1** | `cache=strict, actimeo=1, soft, rsize/wsize=4M, uid/gid=33` | CT 105 (nextcloud) — bind-mounted to `/mnt/pve/unas` inside | + +**Issues:** + +1. **CT 105 is on CIFS to the same data CT 104 uses via NFS.** Pure duplication. Nextcloud does massive amounts of `stat()` traffic; `actimeo=1` on the CIFS mount forces *every* metadata lookup to hit the wire, which is slow. +2. **NFS is v3, not v4.x.** v4 is preferred unless the UDM doesn't export it. v4 fixes locking, removes the separate mountd dance, and supports session trunking. +3. **`mountproto=udp`** under packet loss can intermittently fail to (re)mount. Set `mountproto=tcp`. +4. **`hard` mount with no `intr` equivalent.** If UNAS goes away, anything blocked on it hangs the calling process indefinitely. For non-critical use cases (Nextcloud, but **not** backrest), `soft,timeo=100,retrans=3` is friendlier — Backrest backups should stay `hard`. +5. **CT 105 cannot mount NFS directly** because the UNAS export only allows IPs `.20/.40/.60/.172.30.33.1` — `.41` (CT 105) is missing. So either keep the host-side bind-mount approach (correct) or have UNAS export to `.41` too. +6. **The bind-mount approach is correct** for unprivileged CTs that can't run NFS/CIFS clients themselves. Don't change that pattern. + +**Recommended consolidation:** + +```bash +# 1. Probe whether the NAS speaks NFSv4 +mount -t nfs -o vers=4.2,proto=tcp 192.168.1.31:/var/nfs/shared/storage /mnt/test +# if it works: +pvesm set unas --options vers=4.2,proto=tcp,hard,noatime +# (this re-mounts on next access; or unmount/remount /mnt/pve/unas) + +# 2. Switch CT 105 to the NFS bind-mount +pct set 105 --mp0 /mnt/pve/unas,mp=/mnt/pve/unas +# (CT 105 currently uses unas_smb → unas. New line bind-mounts the NFS mount.) +# Then verify nextcloud-aio still sees uid/gid 33 properly — NFS uses host UIDs, +# whereas CIFS was forcing uid=33. May need to chown on the NAS or add an idmap. + +# 3. Drop the CIFS storage once CT 105 is migrated +pvesm remove unas_smb # if it exists as PVE storage +# or remove the entry from /etc/pve/storage.cfg +``` + +**Notes on perf:** + +- 1 GbE NIC ceiling is ~118 MB/s. NFS rsize/wsize=1M already saturates that. +- If you have or add a 2.5 GbE / 10 GbE link on either side, also enable jumbo frames (MTU 9000) on `vmbr0`, the host NIC, and UNAS. That alone can ~double bulk-read throughput. +- For Nextcloud's small-file stats: `actimeo=60` (NFS) or `cache=loose,actimeo=60` (CIFS, if you stay on it) — dramatically cuts roundtrips at the cost of slightly stale directory listings. + +--- + +## 5. CPU / boot / kernel + +| Item | State | Recommend | +|---|---|---| +| Governor | `performance` | keep | +| HWP EPP | `default` | set to `balance_performance` if you want some idle savings without latency cost: `echo balance_performance > /sys/devices/system/cpu/cpu*/cpufreq/energy_performance_preference` | +| `intel_iommu=on iommu=pt` | set ✔ | keep | +| GPU passthrough (`i915.force_probe=!7dd5 xe.force_probe=7dd5`) | set for Arc Xe (Meteor Lake) | keep | +| `nvme_core.default_ps_max_latency_us=0` | set ✔ | disables NVMe power-save — good for stability, costs ~1 W idle | +| `kernel.numa_balancing` | 0 | correct for single socket | +| Old kernels installed | 6.17.13-4 (current) + 7.0.0-3 | keep both for now; remove 7.0.0-3 once you've booted 7.0.2-5 successfully after the pending upgrade | + +### Hybrid-core scheduling + +The 155H has P-cores (cores 0–11), E-cores (12–17), LP-E cores (18–21). Linux 6.x with `intel_pstate=active` handles ITD/HWP well; no manual pinning is needed for current workloads. If a CT becomes latency-sensitive, you can pin it with `cpuset` via `lxc.cgroup2.cpuset.cpus` (P-cores only). + +--- + +## 6. Reliability / SPOF + +### Single disk is the biggest risk + +- Current: 1× CT2000P3PSSD8 (QLC, DRAM-less). 19 % wear in (check `power_on_hours`). +- **Recommendation:** add a second 2 TB NVMe (ideally a TLC drive with DRAM — Crucial T500, WD SN770, Samsung 990 Pro, Solidigm P44 Pro). Then: + ```bash + zpool attach rpool nvme-CT2000P3PSSD8_2429E8BBCFB4-part3 /dev/disk/by-id/-part3 + ``` + (requires partitioning the new disk to match — `sgdisk -R` from the existing). Pool becomes a mirror with full self-heal. +- If a second internal slot isn't available, at minimum increase backup frequency for CT 101/104 (the irreplaceable data). + +### Boot redundancy + +`proxmox-boot-tool kernel list` shows one bootloader entry. After §6 mirror is set up, run `proxmox-boot-tool init /dev/-partN` so either disk can boot. + +--- + +## 7. Backups (HIGH PRIORITY — silent risk) + +`/etc/pve/storage.cfg`: +``` +nfs: unas + prune-backups keep-all=1 +``` + +**`keep-all=1` means backups are never deleted automatically.** UNAS already holds 2 TB. Set a real policy, e.g.: + +```bash +pvesm set unas --prune-backups keep-last=3,keep-daily=7,keep-weekly=4,keep-monthly=6 +``` + +Also: there is **no `vzdump` job configured** in `/etc/pve/jobs.cfg`. Backups are either manual or driven from CT 103 (Backrest). Recommend a scheduled `vzdump` job for at least VM 100 and CT 101/104 *in addition to* Backrest, so PVE-native restores remain trivial. + +--- + +## 8. APT / repositories cleanup + +State today: + +``` +/etc/apt/sources.list.d/ +├── ceph.list # all lines commented — fine but consider deleting the file +├── proxmox.sources # pve-no-subscription (modern deb822) ← keep +├── pve-enterprise.list.bak # backup, safe to remove +├── pve-enterprise.sources # Enabled: false ← keep as-is or remove +├── pve-install-repo.list # pve-no-subscription duplicate +└── pve-no-subscription.list # pve-no-subscription duplicate +``` + +`pve-install-repo.list` and `pve-no-subscription.list` duplicate what `proxmox.sources` already declares. APT deduplicates fetches but the duplication is a foot-gun (one of them will go stale on the next PVE major version transition). Recommended cleanup: + +```bash +rm /etc/apt/sources.list.d/pve-install-repo.list +rm /etc/apt/sources.list.d/pve-no-subscription.list +rm /etc/apt/sources.list.d/pve-enterprise.list.bak +# keep proxmox.sources and pve-enterprise.sources (already disabled) +apt update +``` + +Also: there are **9 pending upgrades** including `pve-manager 9.1.18` (you're on 9.1.11) and a kernel update. Run `apt update && apt full-upgrade` at a convenient window. + +### Unattended-upgrades (configured 2026-05-20) + +The host previously had a cron line `0 2 * * * apt-get update && apt-get upgrade -y` that was silently no-op'ing on every kernel / PVE point release — `apt-get upgrade` refuses to install new dependencies, which PVE updates always introduce. + +Replaced with `unattended-upgrades` in a conservative profile: + +| File | Purpose | +|---|---| +| `/etc/apt/apt.conf.d/52unattended-upgrades-pve` | local policy — origins allowlist + email + reboot policy | +| `/etc/apt/apt.conf.d/20auto-upgrades` | enables the daily update-list + unattended-upgrade run | + +**Auto-applied:** +- `origin=Debian,codename=trixie,label=Debian` (stable main) +- `origin=Debian,codename=trixie-security,label=Debian-Security` +- `origin=Debian,codename=trixie-updates` (stable point updates) + +**Held for manual `apt full-upgrade`** (intentionally — review release notes first): +- `origin=Proxmox,...` — pve-manager, kernels, qemu-server, all PVE components + +**Settings:** +- `Automatic-Reboot "false"` — kernel updates require a manual reboot +- `Remove-Unused-Dependencies "true"` — autoremove orphans after upgrades +- `AutoFixInterruptedDpkg "true"` — resume after crash mid-upgrade +- `Mail "notify@home.box"`, `MailReport "on-change"` — alerts on actual changes + +**Triggered by:** +- `apt-daily.timer` (daily ~07:00) — refresh package lists +- `apt-daily-upgrade.timer` (daily ~06:00) — apply unattended upgrades + +**Caveat: mail delivery isn't reaching you yet.** Postfix is up but has `relayhost = (none)` — change notifications get delivered locally to `/var/mail/notify` on the host, not to your inbox. Set up a smart-host relay (Gmail/Postmark/etc.) if you want the mails to actually land. Until then, check `/var/log/unattended-upgrades/unattended-upgrades.log` for history. + +**Verify any time:** +```bash +unattended-upgrade --dry-run --debug 2>&1 | grep -E "Allowed origins|would be upgraded|pkgs that look" +systemctl list-timers apt-daily-upgrade.timer +tail /var/log/unattended-upgrades/unattended-upgrades.log +``` + +--- + +## 9. Networking + +- Single bridge `vmbr0` on `enp86s0` — no VLAN aware (`bridge-vlan-aware yes`). If you ever want to segment guests by VLAN, add it now (no impact on existing guests as long as you don't tag them): + ``` + bridge-vlan-aware yes + bridge-vids 2-4094 + ``` +- `net.core.rmem_max` / `wmem_max` are at distro defaults (208 KiB). With a 1 GbE NIC the impact is small (link is already saturated at NFS rsize=1M), but with future 2.5/10 GbE bump to 16 MiB: + ```bash + cat >/etc/sysctl.d/99-net.conf <<'EOF' + net.core.rmem_max=16777216 + net.core.wmem_max=16777216 + net.ipv4.tcp_rmem=4096 87380 16777216 + net.ipv4.tcp_wmem=4096 65536 16777216 + EOF + sysctl --system + ``` +- TCP congestion control is `cubic`. `bbr` is generally better for mixed workloads — change only if you measure a problem. +- `wlo1` is present but unused — confirm and disable in BIOS or `iface wlo1 inet manual` (already done). No action. + +--- + +## 10. Container-specific issues + +### CT 104 (docker / AI / image-gen) — 128 GiB cap, 16 cores, GPU passthrough + +Measured: 18.8 GiB current, peak 29.3 GiB, **9.3 GiB in swap**, load 8.0, ~65 Docker containers (Immich + ML, ComfyUI, LobeChat, n8n, Daytona, LiteLLM, Vaultwarden, Paperless+AI, many MCP servers, Speaches-OpenVINO). + +- `cores: 16` is justified by the workload (load avg 8 across 16 = ~50 % avg). Don't drop. +- 128 GiB cap is generous but currently 9 GiB is in swap — this is **memory pressure from CT 101's phantom cap**, not from CT 104's own needs. Fix CT 101 first, re-measure, then either lower CT 104 to 48 GiB or leave it. +- GPU pass-through (`/dev/dri/card1`, `renderD128`) confirmed visible inside CT and being used by Speaches via OpenVINO ✔ +- Keep `recordsize=128K` (large files dominate). +- `swap: 32000` is high — consider `swap: 8192`. Heavy CT swap-out on a QLC root SSD adds write amplification. +- Bind-mounted `/mnt/pve/unas` (NFS) is the right choice ✔ +- Bonus: raise scheduler weight if image-gen latency matters: `pct set 104 -cpuunits 200`. + +### CT 101 (shepard / docker) — currently 160 GiB cap, peak 15.6 GiB + +Workload: ~9 containers — Shepard frontend/backend, Keycloak, Neo4j, MongoDB, MongoExpress, TimescaleDB, Caddy, home-showcase-collector. + +- **Drop memory cap to 32 GiB**: `pct set 101 -memory 32768`. No restart needed. +- `cores: 12` is fine (load 8.5 — close to fully loaded, real work). +- `swap: 8192` ✔ +- This single change unblocks everything else. + +### CT 102 (adguard / DNS) — undersized + +Measured: 343 MiB used at the **512 MiB cap**, AdGuardHome alone is 350 MiB RSS, the CT is one OOM event from killing the LAN's DNS. + +- `pct set 102 -memory 1024` (bump to 1 GiB). +- Add `onboot: 1` (already set) + `startup: order=1` (boots first) + `protection: 1` (anti-fatfinger). This is the only DNS — treat it like infrastructure. + +### CT 103 (backrest) — fine + +89 MiB used, peak 305 MiB. No changes needed. + +### CT 105 (nextcloud) — privileged + CIFS + +- **Privileged container** (no `unprivileged: 1`). For a public-facing app this is the wrong tradeoff. Migration: stop CT, `vzdump` backup, restore as unprivileged. Be ready to fix file ownership on the bind-mount afterwards (privileged UID 33 → unprivileged needs `lxc.idmap`). +- **Uses CIFS bind-mount to UNAS** with `actimeo=1` — see §4. Switch to NFS bind-mount, after confirming UID mapping (CIFS forces uid=33; NFS uses host UIDs as-is). +- Memory cap 8 GiB, peak 3.5 GiB — fine as-is. +- Runs Nextcloud AIO stack (mastercontainer-managed: apache, postgres, redis, imaginary, collabora, notify-push). + +### CT 108 (zoraxy) — slight oversize + +Peak 463 MiB on a 2 GiB cap. Lower to 1 GiB if desired (cosmetic). + +### CT 104 — Docker stacks inventory (`/opt/stacks`) + +CT 104 keeps its Docker workloads in a **git-tracked monorepo at `/opt/stacks/`** with one directory per stack, plus meta-docs (`PORTMAP.md`, `storage.md`, `volumes.md`, `docker-networks.md`, `todo.md`). Good practice — this is how to keep ~65 containers manageable. The other CTs (101, 105) don't have `/opt/stacks` — their compose files live elsewhere. + +**Stack list (29 dirs):** + +| Stack | Status | Notes | +|---|---|---| +| `ai/` | **active** — large subtree | comfyui, lobehub, litellm, speaches, mcp-gateway, mcp-servers (many MCP yml files), searxng. Custom `syncstack.py` to manage cross-file project names. | +| `arr-stack/` | dormant (defined, not running) | rdtclient, prowlarr, audiobookshelf, shelfarr, flaresolverr | +| `arcane/` | dormant | Docker dashboard | +| `daytona/` | active (as `daytona-minimal`) | dev environments + runner + registry | +| `dozzle/` | dormant | container log viewer | +| `gotify/` | active | push notifications | +| `homepage/` | dormant | dashboard | +| `immich/` | active (5 containers) | photo platform + ML | +| `karakeep/` | active (3 containers) | bookmark mgr + chrome + meilisearch | +| `memos/` | active | notes | +| `n8n/` | active | pinned `2.20.11` (good — there's an explicit version-drift comment in the compose) | +| `nexa/` | dormant | (?) | +| `paperless_ai/`, `paperless-ngx/` | active (5 containers between them) | OCR pipeline | +| `pocketid/` | active | OIDC provider | +| `proxy/` | dormant | (?) | +| `qdrant/` | dormant | vector DB | +| `shared-db/` | active | shared-postgres + garage (S3-compatible) + pgadmin (defined) | +| `streamio/` | dormant | media | +| `traccar/` | active | GPS tracker | +| `vaultwarden/` | active | password mgr | +| `vpn/` | dormant | gluetun (intended VPN egress wrapper?) | +| `backups/`, `docs/`, `scripts/` | meta dirs (no compose) | | + +**Observations:** + +1. **~10 stacks are defined but dormant.** No RAM/CPU cost while down, but their images sit on disk and the git repo accrues dead code. Either run them, document why they're parked, or `git rm` them — repo drift is the silent killer of "I know what's running" confidence. +2. **Stack name ≠ compose project name** for the `ai/` and `daytona/` trees (multiple compose files per dir, different project names). The `syncstack.py` helper exists for this; just be aware that `docker compose -f` lookups by directory name don't match. +3. **Disk-reclaim potential (measured `docker system df`):** + + | Asset | Total | Reclaimable | + |---|---|---| + | Images | 83 / 61.6 GB | **8.3 GB** | + | Build cache | 105 entries / 9.0 GB | **4.6 GB** | + | Volumes | 31 / 3.0 GB | 940 MB (21 dangling) | + | Containers | 60 active | 0 | + + ```bash + pct exec 104 -- docker system prune -a --volumes + # or non-destructively just the build cache: + pct exec 104 -- docker builder prune -a + ``` + ~13 GB to recover. On a 200 GB rootfs that's 64 % full, this is meaningful. + +4. **Stack→Zoraxy mapping (§11b):** when fronting via Zoraxy, the canonical service endpoints (per `PORTMAP.md`) are CT 104's IP `192.168.1.40` + port. Worth cross-referencing that file when setting up reverse-proxy entries. + +### General LXC hygiene + +- Strip the giant HTML banners from the top of every `/etc/pve/lxc/*.conf` — pure noise in `pct config`. +- All CTs except 105 are unprivileged ✔ with `nesting=1,keyctl=1` (appropriate for Docker). + +### Anti-fat-finger protection (`protection: 1`) + boot order + +Applied across all critical guests (2026-05-20). `protection: 1` blocks `pct destroy` / "Remove" from the UI until manually unset — cheap insurance against the wrong-CT-deleted incident. + +| Guest | protection | startup order | Rationale | +|---|---|---|---| +| CT 102 (AdGuard / DNS) | ✅ | **order=1** (boots first) | LAN-wide DNS — nothing resolves until this is up | +| CT 108 (Zoraxy / LE proxy) | ✅ | **order=2** | Public-facing reverse proxy + LE; depends on DNS | +| CT 103 (Backrest) | ✅ | default | Holds backup config and snapshot metadata | +| CT 104 (Docker / AI / image-gen) | ✅ | default | Largest data footprint (200 G rootfs); 65 containers | +| CT 105 (Nextcloud) | ✅ | default | User data | +| VM 100 (HAOS) | ✅ | default | Home automation state | +| CT 101 (shepard) | ❌ (left optional) | default | Currently a dev/iteration target; protect once stabilized: `pct set 101 -protection 1` | + +The `onboot: 1` flag was already set on all guests ✔ — they all auto-start on host reboot. The two `startup` ordered ones now also boot in the right sequence: DNS → Zoraxy → everything else in parallel. + +To remove protection on a guest later: `pct set -protection 0` (or `qm set 100 -protection 0`). + +--- + +## 11a. DNS — AdGuard rewrites & site-wide consistency + +**Current state (measured):** + +- AdGuard Home runs on CT 102 (`192.168.1.2`, hostname `dns`) — DNS on `0.0.0.0:53/tcp+udp`, admin UI on `:80`. +- Upstream stack is well-configured: load-balanced mix of Telekom-DoT, NextDNS, DNS4EU, Quad9, Cloudflare, Google (DoH/DoQ where supported), bootstrap+fallback on Cloudflare. +- All 6 LXCs resolve via `nameserver 192.168.1.2` ✔ +- Proxmox host (`192.168.1.20`) resolves via `192.168.1.2` ✔ — but its `search` domain is `box` (probably an install-time leftover; AdGuard's `local_domain_name` is `lan`). +- HAOS VM is on `192.168.1.60` — DNS setting unknown without Home Assistant access; **verify**. +- `rewrites: []` and **`rewrites_enabled: false`** — no internal name resolution is happening today. + +### Why this matters + +Without rewrites, you address everything by IP. That's brittle (IP changes break links), invisible in logs, and prevents nice tricks like split-horizon DNS for `nuclide.systems` (so the same name resolves to Zoraxy LAN-internally without going through your public IP / WAN hairpin). + +### Recommended rewrite set + +In AdGuard UI → Filters → DNS rewrites, **enable rewrites** and add: + +``` +# Split-horizon: public domain → Zoraxy on LAN +nuclide.systems → 192.168.1.4 +*.nuclide.systems → 192.168.1.4 + +# Service-name shortcuts under the local_domain_name (`.lan`) +pve.lan → 192.168.1.20 # Proxmox UI +nuc.lan → 192.168.1.20 # host shorthand +dns.lan → 192.168.1.2 # AdGuard itself +zoraxy.lan → 192.168.1.4 # reverse proxy +shepard.lan → 192.168.1.49 # CT 101 +docker.lan → 192.168.1.40 # CT 104 +nextcloud.lan → 192.168.1.41 # CT 105 +haos.lan → 192.168.1.60 # VM 100 +unas.lan → 192.168.1.31 # NAS +router.lan → 192.168.1.1 # UniFi gateway +``` + +The split-horizon entries are the highest-value: once Zoraxy proxies `pve.nuclide.systems` (see §11b), the same URL works both from the public internet and from inside the LAN — with no NAT-loopback weirdness and with the LAN traffic never leaving the building. + +**Edit the YAML directly if preferred** (`/opt/AdGuardHome/AdGuardHome.yaml` inside CT 102), then restart AdGuard. The line `rewrites_enabled: false` must flip to `true`. + +### Verify clients are actually using AdGuard + +After rewrites are in, walk the inventory: + +| Client | Should use DNS | Check | +|---|---|---| +| All 6 LXCs | ✔ already at .2 | `pct exec -- cat /etc/resolv.conf` | +| Proxmox host | ✔ already at .2 | `cat /etc/resolv.conf` | +| HAOS VM (192.168.1.60) | unknown | HAOS UI → Settings → System → Network → check DNS servers; should be `192.168.1.2` | +| Router (192.168.1.1, UniFi) | **DHCP-hands-out DNS to clients** — must serve .2 as primary | UniFi: Settings → Networks → LAN → DHCP DNS: `192.168.1.2` | +| IoT devices (Roborock at .64, others) | inherit via DHCP from router | once UniFi DHCP serves .2, every device that DHCP-renews picks it up. Force-renew or reboot stragglers. | +| Anything with hard-coded `1.1.1.1` / `8.8.8.8` | bypassing the filter | grep service configs for upstream DNS — apps like Pi-hole-aware clients, some Smart TVs, Chromecasts | + +### Optional hardening once the rewrites are stable + +- **DNSSEC**: `enable_dnssec: true` (currently false). Most upstreams already validate, but flipping this on adds end-to-end checking. +- **Block DoH bypass**: AdGuard has a "Blocked services" preset for Cloudflare/Google DoH endpoints — useful so clients can't sidestep the filter. +- **Resolve `box` search-domain confusion**: edit `/etc/resolv.conf` (or set it via `/etc/network/interfaces`) on the host to `search lan` so it matches AdGuard's `local_domain_name`. +- **Secondary DNS for resilience**: CT 102 is currently the sole DNS — if it crashes, the LAN goes blind. Either add `protection: 1` (done above) and rely on it, or stand up a tiny secondary AdGuard on a different CT and configure UniFi DHCP to hand out both. (Out of scope for low-hanging fruit, but worth knowing.) +- **Conditional forwarding for UNAS**: if UNAS announces hostnames (mDNS or its own DNS), point AdGuard at it for the relevant zone via "Upstream DNS servers" using the `[/zone/]upstream` syntax. + +### Action checklist + +1. AdGuard UI → Filters → DNS rewrites: paste the table above. +2. AdGuard UI → Settings → DNS settings → **enable** "DNS rewrites". +3. UniFi: confirm DHCP option 6 = `192.168.1.2` (LAN clients get AdGuard). +4. HAOS: confirm Home Assistant has `192.168.1.2` set as DNS. +5. Force-renew DHCP leases on key clients (or just wait — most renew within 24 h). +6. After §11b is done, the public `pve.nuclide.systems` resolves to `.4` from inside the LAN automatically. + +--- + +## 11b. TLS certificates for the Proxmox web UI + +**Current state:** the PVE web UI on `https://192.168.1.20:8006` uses the self-signed certificate generated at install (`/etc/pve/local/pveproxy-ssl.pem` is absent → falls back to `pve-ssl.pem`). Every login throws a browser warning. + +**The wider setup:** Zoraxy (CT 108, `192.168.1.4`) already handles Let's Encrypt for `nuclide.systems` (the public domain for this host). So there are three sane options; pick **A** unless you have a reason not to. + +### Option A — Reverse-proxy PVE through Zoraxy (recommended) + +Pros: single source of LE truth (Zoraxy already renews); no DNS-plugin setup; no exposing the API; nice domain like `pve.nuclide.systems`. Cons: depends on Zoraxy being up (keep IP:8006 as fallback); needs WebSocket pass-through for the noVNC console and `xterm.js` shell. + +1. **Zoraxy host entry** + - Domain: `pve.nuclide.systems` (or whatever subdomain) + - Target: `https://192.168.1.20:8006` + - **Enable WebSocket support** (mandatory — noVNC, xterm.js, task log streaming all use it) + - **Skip backend TLS verification** (PVE's cert is self-signed) + - Enable HSTS once you've confirmed the setup works + - Optionally restrict by source: only LAN + Cloudflare/Tailscale IPs +2. **DNS**: add an A record `pve.nuclide.systems` → public IP (or split-horizon to `192.168.1.20` for LAN). Zoraxy will ACME-challenge via whichever method it's configured for (HTTP-01 or DNS-01). +3. **Keep `https://192.168.1.20:8006` reachable on LAN** as an emergency fallback. Don't disable it. +4. Set the PVE redirect-to-HTTPS rules in Zoraxy on for both `:80` and `:443`. + +**Important caveat:** the PVE Mobile app and the `pvesh` / API clients may not love going through a reverse proxy (they're picky about TLS SNI and cookie domains). Keep direct IP access available for API tooling, or test thoroughly. + +### Option B — PVE's built-in ACME with DNS-01 + +Pros: no reverse proxy in the path; PVE renews itself; works for the API too. Cons: requires a DNS provider plugin (your registrar's API token), and an LE-acceptable FQDN that resolves publicly. + +1. Register an ACME account: + ```bash + pvenode acme account register default you@nuclide.systems + ``` +2. Configure a DNS plugin. PVE supports `acme-dns`, `cloudflare`, `route53`, `desec`, etc. via the `acme.sh` plugin set. Example for Cloudflare: + ```bash + pvenode acme plugin add dns cf --api cf --data CF_Token=XXXXXXXX + ``` + Replace `cf` plugin name to match whichever registrar you use for `nuclide.systems`. +3. Tell the node which domain(s) and how: + ```bash + pvenode config set --acme domains=nuc.nuclide.systems + pvenode config set --acmedomain0 domain=nuc.nuclide.systems,plugin=cf + ``` +4. Order: + ```bash + pvenode acme cert order + ``` + PVE drops the cert at `/etc/pve/nodes/nuc/pveproxy-ssl.pem` and renews ~30 days before expiry via the `pve-daily-update` timer. + +### Option C — Push Zoraxy's cert into PVE + +Only useful if A and B are off the table. Zoraxy stores its issued certs (location varies by Zoraxy version — typically under its data dir, e.g. `/opt/zoraxy/conf/certs/`). Cron a script that copies the active cert/key and concatenates them as `/etc/pve/local/pveproxy-ssl.pem` (cert + chain) and `/etc/pve/local/pveproxy-ssl.key`, then `systemctl reload pveproxy`. Brittle — only worth it if you must. + +### Recommended path + +Do **A** (reverse proxy through Zoraxy) for the web UI. It piggybacks on existing renewal. The mobile-app/API edge cases are usually fine if Zoraxy passes the WebSocket and preserves the `Host` header. If you later need full ACME on the node itself (e.g. you want valid TLS for `pvesh` and the API at the node FQDN too), layer **B** on top — they don't conflict. + +### Other CTs with web UIs worth fronting via Zoraxy + +While you're at it, route through Zoraxy for free LE: +- Backrest (CT 103) — currently IP-only +- AdGuard (CT 102) admin UI — `192.168.1.2:3000` +- Nextcloud (CT 105) — almost certainly already exposed; verify it terminates LE in Zoraxy and not internally +- Zoraxy itself (CT 108) — self-hosted, already TLS + +For each, add a Zoraxy host entry, set a subdomain, and disable any local TLS / port-exposed listener that bypasses Zoraxy. + +--- + +## 11. Maintenance / observability + +| Item | State | Recommend | +|---|---|---| +| `lm-sensors` | not installed | `apt install lm-sensors && sensors-detect --auto` for CPU/NVMe temps in the UI | +| Journal size | 1.5 GiB | OK; cap at 1 GiB if you want predictability: `journalctl --vacuum-size=1G` and `SystemMaxUse=1G` in `journald.conf` | +| `fstrim.timer` | active (weekly) | OK; `zfs trim` runs separately when autotrim=on | +| ZFS scrub | last run 2026-05-10, clean | default monthly timer is good | +| Subscription nag | not removed | If desired, `pve-no-nag` patch or the `proxmox-helper-scripts` line — purely cosmetic | +| Email alerts | (check `/etc/pve/user.cfg`) | configure `root@pam` email for failed scrub / failed backup notifications | + +--- + +## 13. Update management — current model + +The host previously had `0 2 * * * apt-get update && apt-get upgrade -y` (silently no-op'd on every PVE/kernel update) and a weekly `bash <(wget tteck/.../update-lxcs-cron.sh)` cron that ran `dist-upgrade` across every LXC. Both removed 2026-05-20 and replaced with the structure below. + +### Layer 1 — Host packages (Debian + Proxmox) + +- `unattended-upgrades 2.12` installed on the host. +- `/etc/apt/apt.conf.d/52unattended-upgrades-pve` allows only `Debian`, `Debian-Security`, `trixie-updates` — **Proxmox origin held** for manual review. +- Triggered by `apt-daily.timer` and `apt-daily-upgrade.timer` (ship with `apt`, both `active enabled`). +- `Automatic-Reboot "false"` — kernel updates wait for a manual reboot. +- `Mail "notify@home.box"`, `MailReport "on-change"` — Postfix is up but **`relayhost = (none)`**, so mail is delivered locally to `/var/mail/notify` (not your inbox until you wire a smart-host). + +### Layer 2 — CT OS packages (Debian) + +`unattended-upgrades` deployed inside every CT (CT 104 already had it; 101/102/103/105/108/110 added 2026-05-20): + +| CT | u-u version | Distro | Status | +|---|---|---|---| +| 101 | 2.12 | trixie | active | +| 102 | 2.9.1 | bookworm | active | +| 103 | 2.9.1 | bookworm | active | +| 104 | 2.12 | trixie | active | +| 105 | 2.9.1 | bookworm | active | +| 108 | 2.12 | trixie | active | +| 110 | 2.12 | trixie | active | + +Per-CT allowlist is Debian-only — third-party repos (`docker.com`, `jotta.cloud`, `claude.ai`, `cli.github.com`, `dl.k6.io`) are **excluded** because they ship breaking changes outside Debian's freeze. Upgrade those with explicit `apt upgrade `. + +### Layer 3 — Helper-script app binaries (AdGuard, Zoraxy) + +Each helper-scripts CT ships `/usr/bin/update` that re-`curl|bash`'s the community-scripts installer. Replaced with proper systemd timers using the apps' own update mechanisms: + +| CT | Timer | Schedule | Mechanism | +|---|---|---|---| +| 102 AdGuard | `adguard-update.timer` | Wed 03:30 (+15 m jitter) | native `AdGuardHome --update` flag | +| 108 Zoraxy | `zoraxy-update.timer` | Wed 03:40 (+15 m jitter) | GitHub releases API, **stable semver only (skips RCs)**, binary swap + 30 s health check + auto-rollback | + +Both log to `/var/log/{adguard,zoraxy}-update.log` and journal. Manual invoke: `systemctl start -update.service`. + +### Layer 4 — Docker engine inside CTs + +`docker-ce` updates in CT 101, 104, 105, 110 — held by the Debian-only allowlist. Apply with `apt upgrade docker-ce docker-ce-cli containerd.io` when you want them. Add `origin=Docker` to the allowlist if you want to auto-apply (not recommended; engine updates occasionally break running containers). + +### Layer 5 — Docker images (the ~75 containers) + +**Plan: deploy Diun on CT 109** (observability LXC, see §16). Diun watches image tags on registries, posts to Gotify when a new image is available. Pulls remain manual (`docker compose pull && up -d`) — protects against latest-tag drift like the n8n incident pinned in `/opt/stacks/n8n/docker-compose.yaml`. + +**Layer 6 — Nextcloud-AIO**: self-updates via the mastercontainer (CT 105). No external mechanism needed. + +--- + +## 14. VM 100 (HAOS) auto-restart watchdog + +Old approach: `*/5 * * * * /root/vm100.sh > /dev/null` in cron. Script archived to `/root/vm100.sh.bak` on 2026-05-20. + +Replaced with a systemd timer + oneshot: + +- `/usr/local/sbin/vm100-watchdog.sh` — only restarts on `status: stopped`; skips paused/prelaunch/transitional states; respects `/var/lock/qemu-server/lock-100.conf` so it doesn't race vzdump or migration +- `vm100-watchdog.service` (Type=oneshot) +- `vm100-watchdog.timer` (`OnUnitActiveSec=1min`, RandomizedDelaySec=15s) + +Recovery latency improved from 5 min → 1 min; logging structured in `journalctl -u vm100-watchdog`. + +--- + +## 15. Identity — Pocket-ID on its own CT + +### State as of 2026-05-20 + +**CT 110 "id"** created at `192.168.1.5` as the dedicated IdP host. Pocket-ID was previously on CT 104 as one of ~65 docker containers; moved off because: + +- Pocket-ID gates auth for everything else; sharing a host with image-gen workloads (9 GiB swap pressure on CT 104) means a single OOM/restart there knocks out SSO across the whole estate +- Pattern match with CT 102 (DNS) and CT 108 (reverse proxy) — each infra-critical concern gets its own CT + +| CT 110 setting | Value | +|---|---| +| Hostname / IP | `id` / `192.168.1.5` | +| Cores / RAM / rootfs | 1 / 1 GB / 4 GB | +| Privilege | unprivileged, nesting=1, keyctl=1 | +| Boot order | `onboot=1`, `startup=order=3` (after DNS=1, Zoraxy=2) | +| Protection | `protection: 1` | +| Auto-updates | unattended-upgrades, Debian-only allowlist | +| Docker | 29.5.1 + compose v5.1.3 | + +### Duplication procedure used + +1. `sqlite3 pocket-id.db ".backup /tmp/pi-snap/pocket-id.db"` on CT 104 (online, no downtime to id.nuclide.systems) +2. tar everything except `*.db*`; restore the live snapshot as `pocket-id.db` +3. `pct pull` → `pct push` to CT 110 +4. Adapted compose to drop the `shared_backend` external network reference (CT 110 uses default bridge) +5. `docker compose up -d` +6. Verified `http://192.168.1.5:11000/healthz` returns 200 + +### Zoraxy cutover + +`id.nuclide.systems` upstream needs to change from `192.168.1.40:11000` → `192.168.1.5:11000`. Single-line config edit in Zoraxy + reload. Verified live in §11b once executed. + +### Secrets-rotation list (deferred to cutover day) + +- `OIDC_CLIENT_SECRET` for the Arcane registration (exposed in chat transcript): rotate in Pocket-ID UI, update Arcane env, restart Arcane +- `ENCRYPTION_KEY` and `JWT_SECRET` in `/opt/stacks/arcane/docker-compose.yml`: move to `.env` (currently empty), regenerate, restart Arcane. Existing user sessions get invalidated — fine, ask everyone to log in again + +--- + +## 16. CT 109 "observe" — planned observability + ops LXC + +Single LXC holding everything monitoring/ops-shaped. Sizing target: **4 cores / 6 GiB RAM / 30 GiB rootfs**, unprivileged, nesting=1, IP `192.168.1.6` (next free in the infra range after CT 110). + +### Stack to deploy on CT 109 + +| Service | Purpose | +|---|---| +| Arcane Manager | central docker management UI; edge agents on CT 101 + CT 104 + CT 110 (mTLS, agent-dialed-out) | +| Dozzle UI | aggregated container logs; Dozzle agents on CT 101 + CT 104 + CT 110 | +| Prometheus | metrics TSDB, 30 d retention | +| Grafana | dashboards over Prometheus | +| Alertmanager + alertmanager-gotify-bridge | alert routing → Gotify | +| **Homarr** (selected over Homepage) | unified dashboard, native OIDC with Pocket-ID, native Prometheus widget + Grafana iframe support | +| Diun | docker image update notifier → Gotify | +| Tinyauth | forward-auth gate for non-OIDC apps (Backrest, raw Dozzle, raw Prometheus, raw Grafana). Speaks OIDC client to Pocket-ID, returns Basic-Auth/Remote-User header to upstream | +| docker-socket-proxy | local + remote (CT 101/104/110) — hardened read-only docker.sock for Homarr discovery | + +### Sidecars deployed on each Docker host + +| Host | Sidecars | +|---|---| +| Host (PVE) | node-exporter, smartctl-exporter, pve-exporter (host metrics into Prometheus) | +| CT 101 | node-exporter, cAdvisor, Dozzle agent, Arcane edge agent, docker-socket-proxy | +| CT 104 | node-exporter, cAdvisor, Dozzle agent, Arcane edge agent, docker-socket-proxy | +| CT 110 | node-exporter, Dozzle agent, Arcane edge agent (test target — small/clean) | + +### Services that stay where they are (NOT on CT 109) + +- AdGuard (CT 102), Zoraxy (CT 108), Backrest (CT 103), Pocket-ID (CT 110) — each has its own CT for blast-radius reasons +- All app stacks (Immich, Nextcloud, Shepard, n8n, ComfyUI, etc.) + +### Migration of Gotify (Tier B — schedule when CT 109 is otherwise stable) + +Gotify currently runs in CT 104 docker stack at `gotify.nuclide.systems`. Moving to CT 109 isolates alerting from CT 104 outages but means updating env vars / webhook targets in ~10 places (MCP servers, Backrest webhooks). Plan: copy DB and app tokens via volume tar, deploy on CT 109, update Zoraxy upstream, then sweep dependents. + +### Migration of MCP Gateway (Tier B — split control / workload) + +`/opt/stacks/ai/mcp-gateway/` on CT 104 is the OIDC-gated control plane for the ~20 MCP child containers. Decision (2026-05-20): move **only the gateway** to CT 109; **the child MCP containers stay on CT 104** (they're workload, not control plane). + +Mechanics: +- Gateway on CT 109 uses `DOCKER_HOST=tcp://:2375` (the docker-socket-proxy already planned for CT 104) instead of the bind-mounted `/var/run/docker.sock`. socket-proxy ACL must allow `containers, exec, images` (read+write). +- Migrate state files via volume tar: `config.json`, `agents.json`, `prompts/`, `usage.db`, `gateway_tokens.json`, `nc_user_creds.json`. Keep them on CT 109 local zfs, not UNAS (per-request latency matters). +- Pocket-ID redirect URI stays `https://mcp.nuclide.systems/sso/callback` — only Zoraxy's upstream flips from `192.168.1.40:8080` to the CT 109 IP. +- Don't touch the children; gateway still spawns them by name against CT 104's daemon. + +Build-order slot: after Arcane + socket-proxy land in §16's checklist, before Tinyauth. + +### Arcane specifics + +- Existing `/opt/stacks/arcane/` on CT 104 has the Manager 80 % built: 41 MB SQLite DB carrying Pocket-ID OIDC client + admin user `fkrebs@nucli.de` +- Container currently stopped (~since May 19) +- **Edge mode** chosen for agents — agent dials out to Manager over HTTPS/gRPC with mTLS. Manager auto-generates the edge CA on first start; each agent enrolls with a token and gets a 1 yr client cert with auto-renewal +- Same `ghcr.io/getarcaneapp/arcane:latest` image for Manager and agents; mode is env-driven (`ARCANE_EDGE_AGENT=true` on agents) +- Intel iGPU is not in Arcane's GPU monitoring catalog (NVIDIA + AMD only). LXC-level passthrough to CT 104 stays untouched; only Arcane's GPU widgets will be blank for the Arc + +### Build order + +1. Create CT 109 (specs above) +2. Deploy node-exporter on host + Prometheus + Grafana first (start collecting baselines) +3. Deploy Arcane Manager with the existing DB restored from CT 104 +4. Issue agent tokens, deploy Edge agents on CT 101/104/110 +5. Deploy Dozzle Manager + agents +6. Deploy Diun +7. Deploy Tinyauth (configure Pocket-ID client first) +8. Deploy Homarr +9. Front the lot via Zoraxy: `arcane.`, `dozzle.`, `grafana.`, `prom.`, `home.` `.nuclide.systems` +10. Verify each through the Tinyauth gate where applicable +11. Add Diun watchlist + Alertmanager routing → Gotify + +--- + +## 17. LLDP / UniFi topology visibility + +Investigated 2026-05-20: + +- `lldpd` is **installed and active** on the host (defaults: advertises + listens) +- Existing `/usr/local/bin/update_interface_desc.sh` hourly cron consumes LLDP from neighbors and writes `# PortDescr:` comments into `/etc/network/interfaces` +- Host is on switch **port 10** of the D-Link DGS-1210-28P (`SysName: dgs1210`, FW `6.32.008`). +- `/usr/local/bin/update_interface_desc.sh` was silently failing — `lldpcli` not in cron's PATH **and** the original logic only appended new lines, never replaced stale ones. Rewritten 2026-05-20: + - PATH set inside the script (`/usr/sbin` included) + - Uses `lldpcli show neighbors -f keyvalue` for machine-readable parsing + - Single `# LLDP: :: ` line per interface; legacy `# PortDescr:` lines stripped + - Idempotent (second run = no-op, no backup churn) + - Daily-keyed backup `/etc/network/interfaces.bak.YYYYMMDD` + - Only logs on actual change + - Old script preserved at `/usr/local/bin/update_interface_desc.sh.bak-2026-05-20` + +**The UDM Pro won't see "nuc" in its topology view** because LLDP frames use the Nearest-Bridge multicast (`01:80:c2:00:00:0e`) which any 802.1D-compliant switch terminates by spec — and the D-Link DGS-1210-28P sits between the host and the UDM. LLDP-MED is **not** a fix for this; it's for endpoint (VoIP/MFP) discovery, not transparent LLDP forwarding. + +Remediation paths in order of effort: + +1. **Enable SNMP v2c/v3 on the DGS-1210** + add as a Generic SNMP device in UniFi → UDM sees the switch and can map port↔MAC. Most practical for this stack. +2. **Add `snmpd` to the host** + generic SNMP device in UniFi → CPU/mem/iface stats from Proxmox visible in UniFi (not in topology, but in monitoring). +3. **UniFi-managed switch between host and UDM** — clean answer; requires hardware. + +D-Link DGS-1210 admin UI lives at `http://192.168.1.10/` (login page, field names `pelican` / `pinkpanther`). Verify the admin password is non-default — DGS-1210 ships with `admin`/blank or `admin`/`admin` on most firmware revisions. A flat-LAN switch with default creds is one of the easier vectors. + +Useful inspection from the host any time: `lldpcli show neighbors`. + +--- + +## 18a. UNAS access — UID consistency model (post-NFSv4 investigation) + +**Investigation result (2026-05-20):** UNAS Pro advertises NFSv4 in `rpcinfo` but has no v4 export tree configured. Every v4 mount attempt returns `No such file or directory`. Ubiquiti has **not** announced NFSv4 support and the [community thread asking for it](https://community.ui.com/questions/When-will-NFSv4-be-available-for-UNAS-Pro/b8f8b00b-c5ac-4e19-adc4-9e5cae06d35f) has no ETA. The official help center also confirms: *"UniFi Drive does not support certain NFS export options, such as `no_root_squash`"*. So root_squash + v3-only is the long-term reality. + +### Universal UID landscape on UNAS + +Every NFS client write lands as **uid 977 / gid 988** (UNAS's `all_squash` + `anon_uid=977`/`anon_gid=988`). The chown probe confirmed no client can change this from the host side. Files written via the legacy CIFS mount appear as **uid 33** to the CIFS client *but* are stored differently on UNAS — the CIFS `forceuid=33` mount option lies about ownership client-side. + +### Per-CT access pattern (canonical) + +| CT | Mount | Container uid | Effective on disk | Status | +|---|---|---|---|---| +| 103 backrest | NFS bind from host | root | squashes to 977 | ✓ consistent | +| 104 docker | NFS bind from host | mostly root, n8n=1000 (latent) | squashes to 977 | ✓ for root containers; n8n latent if it ever writes to UNAS | +| 105 nextcloud | CIFS today (forceuid=33) → **NFS + bindfs** target | uid 33 (www-data) inside Nextcloud, bindfs translates to 977 on disk | needs migration | 🚨 still CIFS | + +### Convention for new containers + +Set `PUID=977 PGID=988` on any container that writes to UNAS. This pre-aligns with UNAS's enforced mapping and avoids latent permission issues (the n8n class). For images that don't support PUID/PGID, run them as root inside the container — root squashes to 977 cleanly. + +### Why bindfs for CT 105 specifically + +Nextcloud's PHP code hard-checks file ownership against `www-data` (uid 33). Without remap, NFS reads return uid 977 and Nextcloud refuses to operate normally. CIFS hides this with `forceuid=33`. NFS+bindfs achieves the same lie with the much faster NFS rail underneath — verified ~5× speed-up on metadata-heavy ops in the non-destructive test on 2026-05-20. + +### Trigger event to revisit + +Watch [community.ui.com/RELEASES](https://community.ui.com/RELEASES) for a UniFi Drive release that adds: +- NFSv4 export option (would enable idmap) +- `no_root_squash` support (would enable server-side chown to specific uids) +- Configurable `anonuid`/`anongid` (would let us match a real uid) + +Any of these would let us simplify the CT 105 stack. + +--- + +## 18. Homarr inventory — services to include on the dashboard + +Captured here so the eventual Homarr config can be assembled in one pass. Groups follow the existing `homepage.*` label convention used in compose files. + +### Group: `infrastructure` + +| Service | URL | Notes | +|---|---|---| +| Proxmox UI | `https://192.168.1.20:8006` | until LE via Zoraxy lands, see §11b | +| AdGuard Home (CT 102) | `http://192.168.1.2/` (UI on `:80`) | DNS + admin | +| Zoraxy (CT 108) | `http://192.168.1.4:8000/` | reverse proxy admin | +| Backrest (CT 103) | `http://192.168.1.3:9898/` | backup orchestration, will be fronted via Tinyauth + Zoraxy | +| **Pocket-ID (CT 110)** | `https://id.nuclide.systems/` | new home, 2026-05-20 | + +### Group: `network` + +| Service | URL | Notes | +|---|---|---| +| UDM Pro | `https://192.168.1.1/` | UniFi controller | +| **D-Link DGS-1210-28P** | `http://192.168.1.10/` | core L2 switch; host on port 10 | +| UNAS Pro | `https://192.168.1.31/` | UniFi NAS | + +### Group: `observability` (to populate when CT 109 lands) + +| Service | URL | +|---|---| +| Grafana | `https://grafana.nuclide.systems` | +| Arcane | `https://arcane.nuclide.systems` | +| Dozzle | `https://dozzle.nuclide.systems` | +| Prometheus | `https://prom.nuclide.systems` (gated by Tinyauth) | +| Alertmanager | `https://alerts.nuclide.systems` (gated by Tinyauth) | + +### Group: `apps` (subset — long list, fill from existing `homepage.*` labels in `/opt/stacks/*/`) + +Immich, Nextcloud, Vaultwarden, Karakeep, Memos, Paperless-ngx, n8n, ComfyUI, LobeChat, LiteLLM, Traccar, Gotify, Speaches, Daytona, Searxng, Kroki, etc. Pull display labels and icons from the existing `homepage.name=` / `homepage.icon=` values per compose. + +--- + +## 12. Suggested action order + +1. **Done 2026-05-20** ✅: + - **Memory / sizing**: CT 101 160 → 32 GiB; CT 102 512 MiB → 1 GiB; HAOS balloon = 4 GiB + - **Protection**: CT 102 startup=1+protection; CT 108 startup=2+protection; CT 103/104/105/110/VM100 protection + - **CT 110 (id)** built at 192.168.1.5, Pocket-ID duplicated (online SQLite snapshot) + - **ZFS**: `autotrim=on`, `atime=off rpool` + - **Backups**: retention set on `unas` + - **APT**: duplicate sources removed; host `unattended-upgrades` deployed (Debian-only) + - **Per-CT u-u**: deployed to all 7 CTs with Debian-only allowlist + - **Cron cleanup**: removed weekly tteck-LXC-update curl-pipe-bash; removed daily broken `apt-get upgrade -y`; replaced `/root/vm100.sh` cron with `vm100-watchdog.timer` (1 min, lock-aware) + - **Self-updaters**: AdGuard (`--update` flag) Wed 03:30; Zoraxy (GitHub stable releases + rollback) Wed 03:40 + - **`zpool upgrade rpool`** ran during the audit (enabled `redaction_list_spill`, `raidz_expansion`) +2. **Today with a maintenance window**: + - **Cut Zoraxy over to CT 110** for `id.nuclide.systems` (single upstream edit; rollback path = revert one line) + - `apt full-upgrade` (kernel 7.0.0-3 → 7.0.2-5, pve-manager 9.1.11 → 9.1.18) + reboot + - VM 100 disk options (`cache=none`, `iothread=1`) — §4. Requires VM stop/start. + - Switch CT 105 from CIFS → NFS bind-mount (§4). Test UID mapping. + - Drop the `unas_smb` storage once CT 105 is migrated. +3. **This week**: + - Re-measure CT 104 peak RSS after fixing CT 101 — likely safe to drop to 48 GiB. + - Convert CT 105 to unprivileged (backup → restore as unprivileged). + - Raise ARC cap to 16 GiB. + - Probe NFSv4 against UNAS; switch if supported. + - **Wire Postfix relayhost** (Gmail/Postmark/your SMTP) so unattended-upgrades + zfs-zed + cron failures actually mail you. + - **Rotate Backrest plan** to back up real data (currently still pointed at `/media/data-dir` — a 50 KB test file from August 2025); see §7. +4. **Medium term**: + - **Build CT 109** observability LXC (§16) — Prometheus + Grafana + Arcane Manager + Dozzle + Homarr + Diun + Tinyauth + - Migrate Gotify from CT 104 to CT 109 (~30 min of env-var updates) + - **Rotate exposed secrets** that appeared in this transcript: Arcane `OIDC_CLIENT_SECRET`, `ENCRYPTION_KEY`, `JWT_SECRET`; Immich `IMMICH_API_KEY` + - Backrest: enable auth, redesign plans to cover all data tiers (§7), front via Tinyauth for OIDC +5. **Next purchase window**: + - Second NVMe → mirror `rpool` (§6) + - `proxmox-boot-tool init` on the new disk + - 2.5 GbE NIC + matching switch port to UNAS for image-gen / backup speed + - UniFi-managed switch between host and UDM (or accept SNMP-only visibility from UniFi) + +--- + +--- + +## 19. Changes applied 2026-05-20 (session 2) + +### Optimizations executed + +| # | Item | Command / action | Result | +|---|------|-----------------|--------| +| 1 | CT 101 protection + boot order | `pct set 101 -protection 1 -startup order=10` | ✅ | +| 4 | CT 104 memory cap 128→48 GiB | `pct set 104 -memory 49152` | ✅ (swap 32→8 deferred: still 10.4 GB in use) | +| 5 | CT 102 rootfs 2→4 GiB | `pct resize 102 rootfs 4G` | ✅ now 28% used | +| A | VM 100 disk: cache=none + iothread=1 | `qm set 100 -scsi0 ...,cache=none,iothread=1` + `scsihw virtio-scsi-single` | ✅ HAOS healthy | +| C | Host apt full-upgrade | kernel 7.0.0→7.0.2-5, pve-manager 9.1.11→9.1.18 | ✅ installed; **reboot pending** | + +### Pocket-ID migration completed + +- `id.nuclide.systems` Zoraxy proxy cutover confirmed: `192.168.1.40:11000` → `192.168.1.5:11000` +- CT 104 `/opt/stacks/pocketid/` directory fully removed (data migrated to CT 110 2026-05-20) +- Backup: `/opt/zoraxy/conf/proxy/id.nuclide.systems.config.bak-pre-ct110` (keep as rollback) + +### Proxmox OIDC via Pocket-ID + +Realm `pocket-id` added; user `fkrebs@nucli.de@pocket-id` mapped to Administrator role. + +```bash +pveum realm add pocket-id \ + --type openid \ + --issuer-url https://id.nuclide.systems \ + --client-id 38469e7e-1fff-4841-83a9-74bf38d847eb \ + --client-key \ + --username-claim email \ + --comment "Pocket-ID OIDC" + +pveum user add fkrebs@nucli.de@pocket-id +pveum aclmod / --users fkrebs@nucli.de@pocket-id --roles Administrator +``` + +OIDC client inserted directly into Pocket-ID SQLite (API key stored as SHA-256 hash — not reversible): + +``` +DB: /opt/stacks/pocketid/data/pocket-id.db on CT 110 +Table: oidc_clients +client_id: 38469e7e-1fff-4841-83a9-74bf38d847eb +name: Proxmox VE +callback_urls: ["https://192.168.1.20:8006"] +``` + +**To add future OIDC clients without UI access:** +```bash +python3 -c " +import uuid, secrets, bcrypt, json, datetime +client_id = str(uuid.uuid4()) +secret_plain = secrets.token_urlsafe(32) +secret_hash = bcrypt.hashpw(secret_plain.encode(), bcrypt.gensalt(rounds=10)).decode() +print(f'id={client_id}') +print(f'secret={secret_plain}') +print(f'hash={secret_hash}') +" +# Then INSERT into oidc_clients with the hash, use secret_plain in the app config +# callback_urls and logout_callback_urls are JSON arrays stored as BLOB +# credentials field is '{}' for standard clients +``` + +**Note on Pocket-ID API keys:** The `key` column in `api_keys` stores a SHA-256 hash of the real key (64-char hex). The plaintext key is only shown once at creation time in the UI. If lost, create a new one — there is no recovery path. + +**Login flow:** In PVE web UI, select realm `pocket-id` at login. You will be redirected to `https://id.nuclide.systems` for authentication and returned to PVE. The `email` claim is used as the PVE username. + +--- + +## Audit footnote — side effects of this run + +- `zpool upgrade rpool` was executed (not `-n`). Enabled features: `redaction_list_spill`, `raidz_expansion`. Safe on current ZFS version; the pool can no longer be imported by ZFS releases that pre-date these features. No data risk. +- No other state changes were made by this audit. diff --git a/stacks/CLAUDE.md b/stacks/CLAUDE.md new file mode 100644 index 0000000..f0536da --- /dev/null +++ b/stacks/CLAUDE.md @@ -0,0 +1,44 @@ +# Notes for agents working in /opt/stacks (CT 104 "docker") + +Most stacks here are intentional and active. Two things deserve explicit awareness so +no agent "helpfully" reverts them: + +## Pocket-ID is gone from this CT (migrated 2026-05-20) + +- Pocket-ID **moved off CT 104** to dedicated **CT 110 ("id", 192.168.1.5)**. +- The compose file `/opt/stacks/pocketid/docker-compose.yml` was deliberately renamed to + `docker-compose.yml.MIGRATED-TO-CT110-2026-05-20` so accidental `docker compose up` in + that directory is a no-op. +- The data dir `/opt/stacks/pocketid/data/` is a migration snapshot only. It is **stale**; + authoritative writes happen on CT 110 since the cutover. +- Zoraxy upstream for `id.nuclide.systems` now points at `192.168.1.5:11000`. +- **Do not** restart, `compose up`, or re-deploy pocket-id here. If something appears + broken in OIDC, investigate CT 110 instead. SSH/Console access via `pct enter 110` on + the host, or `https://192.168.1.5:11000/`. + +## Dormant stacks — defined but not running on purpose + +Per the audit on 2026-05-20, the following stack directories exist but their containers +are intentionally not running. Some are planned to move to **CT 109 "observe"** when that +LXC is built; others are legacy experiments waiting on a decision: + +- `/opt/stacks/arcane/` — Manager will move to CT 109; edge agent eventually deploys here +- `/opt/stacks/dozzle/` — UI will move to CT 109; agent eventually deploys here +- `/opt/stacks/homepage/` — replaced by Homarr on CT 109; can be removed once CT 109 lands +- `/opt/stacks/arr-stack/`, `streamio/`, `nexa/`, `qdrant/`, `proxy/`, `vpn/` — legacy experiments + +Do not start these without checking with the operator first. + +## MCP gateway leaves, MCP child containers stay (planned) + +When CT 109 "observe" lands, **`/opt/stacks/ai/mcp-gateway/` moves to CT 109** (it's a control plane: OIDC, agent scheduling, token store, usage stats). The ~20 MCP child server containers in this CT (`mcp-searxng`, `mcp-gotify`, `mcp-immich`, `mcp-fetch`, `mcp-time`, `comfyui-mcp`, `daytona-mcp`, `kroki-mcp`, etc.) **stay here on CT 104**. + +The migrated gateway will reach this CT's docker daemon over the planned docker-socket-proxy (see `/docs/proxmox-optimizations.md` §16). Until CT 109 exists, the gateway runs locally and uses the bind-mounted socket. Do not piecemeal-move the gateway; it's tied to the CT 109 build. + +## What is safe to do here + +- Anything to do with the actively-running ~65 containers (Immich, n8n, Karakeep, + ComfyUI, Daytona, Vaultwarden, Paperless, LobeChat, LiteLLM, MCP servers, etc.). +- Image pulls + rebuilds for the active stacks (mind the n8n version pin — see its + `docker-compose.yaml` comment). +- Rotate exposed secrets in `.env` files (compose YAMLs themselves are git-tracked). diff --git a/stacks/PORTMAP.md b/stacks/PORTMAP.md new file mode 100644 index 0000000..769ef70 --- /dev/null +++ b/stacks/PORTMAP.md @@ -0,0 +1,207 @@ +# Port Map — NUC 14 Docker Stacks + +> **Reverse proxy**: [Zoraxy v3.3.2](http://192.168.1.4:8000) on `192.168.1.4:8000` (LXC 108) +> Wildcard cert `*.nuclide.systems` · source of truth: [`proxy/zoraxy/routes.json`](proxy/zoraxy/routes.json) +> Manage routes: `uv run scripts/zoraxy_sync.py [--apply|--prune|--list]` + +## Port Scheme + +| Range | Category | +|---|---| +| 10000–10999 | Infrastructure | +| 11000–11999 | Security & Auth | +| 12000–12999 | Media – Immich | +| 13000–13999 | Media – Downloads / Arr | +| 14000–14999 | AI Stack | +| 15000–15999 | Documents | +| 16000–16999 | Automation | +| 17000–17999 | Notes & Bookmarks | +| 18000–18999 | DevOps / Image Gen | +| 19000–19999 | Tracking | +| 20000–20999 | Storage Admin | +| 30000–30999 | VPN Control | + +--- + +## Infrastructure (10000–10999) + +| Port | Service | Container | Public URL | Notes | +|---|---|---|---|---| +| [10000](http://192.168.1.40:10000) | Homepage | `homepage` | — | Dashboard, LAN only | +| [10001](http://192.168.1.40:10001) | Dozzle | `dozzle` | [dozzle.nuclide.systems](https://dozzle.nuclide.systems) | Log viewer | +| [10002](http://192.168.1.40:10002) | Arcane | `arcane` | [arcane.nuclide.systems](https://arcane.nuclide.systems) | Web IDE | +| [10003](http://192.168.1.40:10003) | Gotify | `gotify` | [gotify.nuclide.systems](https://gotify.nuclide.systems) | Push notifications | +| [10004](http://192.168.1.40:10004) | Garage S3 API | `garage` | s3.nuclide.systems (non-responsive, see notes) | Internal: `http://garage:3900` | +| 10005 (127.0.0.1 only) | Garage Admin | `garage` | — | localhost only | + +--- + +## Security & Auth (11000–11999) + +| Port | Service | Container | Public URL | Notes | +|---|---|---|---|---| +| [11000](http://192.168.1.40:11000) | Pocket ID | `pocket-id` | [id.nuclide.systems](https://id.nuclide.systems) | OIDC provider | +| [11001](http://192.168.1.40:11001) | Vaultwarden | `vaultwarden` | [vault.nuclide.systems](https://vault.nuclide.systems) | Password manager | + +--- + +## Media – Immich (12000–12999) + +| Port | Service | Container | Public URL | Notes | +|---|---|---|---|---| +| [12000](http://192.168.1.40:12000) | Immich | `immich_server` | [immich.nuclide.systems](https://immich.nuclide.systems) | Photos/videos | +| — | Immich Power Tools | `immich_power_tools` | [immich-tools.nuclide.systems](https://immich-tools.nuclide.systems) | Container-internal :3000, Zoraxy proxy | + +--- + +## Media – Downloads / Arr Stack (13000–13999) + +All arr-stack services run behind `vpn_gluetun` container network. + +| Port | Service | Container | Public URL | Notes | +|---|---|---|---|---| +| [13001](http://192.168.1.40:13001) | RDTClient | `rdtclient` | — | LAN only | +| [13002](http://192.168.1.40:13002) | Prowlarr | `prowlarr` | — | LAN only | +| [13003](http://192.168.1.40:13003) | Audiobookshelf | `audiobookshelf` | [abs.nuclide.systems](https://abs.nuclide.systems) | | +| [13004](http://192.168.1.40:13004) | ShelfArr | `shelfarr` | — | LAN only | +| [13005](http://192.168.1.40:13005) | Flaresolverr | `flaresolverr` | — | Internal only | +| [30000](http://192.168.1.40:30000) | Gluetun VPN control | `vpn_gluetun` | — | HTTP control API | + +--- + +## AI Stack (14000–14999) + +| Port | Service | Container | Public URL | Notes | +|---|---|---|---|---| +| [14000](http://192.168.1.40:14000) | LiteLLM | `litellm` | [ai.nuclide.systems](https://ai.nuclide.systems) | LLM proxy gateway | +| [14001](http://192.168.1.40:14001) | LobeHub | `lobehub` | [chat.nuclide.systems](https://chat.nuclide.systems) | Chat UI | +| [14002](http://192.168.1.40:14002) | Qdrant | `qdrant_scientific` | — | Vector DB, internal only | +| [8080](http://192.168.1.40:8080) | MCP Gateway | `mcp-gateway` | [mcp.nuclide.systems](https://mcp.nuclide.systems) | Pocket ID gated; spawns MCP servers on `ai-internal` | +| [18002](http://192.168.1.40:18002) | ComfyUI | `comfyui` | — | LAN only (Intel Arc iGPU, FLUX.1-schnell GGUF) | +| [18003](http://192.168.1.40:18003) | ComfyUI MCP | `comfyui-mcp` | — | FastMCP; reachable via gateway at `mcp.nuclide.systems/comfyui/mcp` | +| — | SearXNG | `searxng` | — | Internal, `shared_backend`, used by LobeHub | + +--- + +## Documents (15000–15999) + +| Port | Service | Container | Public URL | Notes | +|---|---|---|---|---| +| [15000](http://192.168.1.40:15000) | Traccar HTTP | `traccar` | [traccar.nuclide.systems](https://traccar.nuclide.systems) | GPS tracking UI | +| [15001](http://192.168.1.40:15001) | Traccar GPS | `traccar` | — | TCP+UDP watch protocol | +| [15002](http://192.168.1.40:15002) | Paperless AI | `paperless-ai` | [paperless-ai.nuclide.systems](https://paperless-ai.nuclide.systems) | Internal-only Zoraxy policy | +| [15003](http://192.168.1.40:15003) | Paperless-ngx | `paperless-ngx-webserver-1` | [paperless.nuclide.systems](https://paperless.nuclide.systems) | Internal-only Zoraxy policy | + +--- + +## Automation (16000–16999) + +| Port | Service | Container | Public URL | Notes | +|---|---|---|---|---| +| [16000](http://192.168.1.40:16000) | n8n | `n8n` | [n8n.nuclide.systems](https://n8n.nuclide.systems) | Workflow automation | + +--- + +## Notes & Bookmarks (17000–17999) + +| Port | Service | Container | Public URL | Notes | +|---|---|---|---|---| +| [17000](http://192.168.1.40:17000) | Memos | `memos` | [memos.nuclide.systems](https://memos.nuclide.systems) | Notes | +| [17001](http://192.168.1.40:17001) | Karakeep | `karakeep` | [hoarder.nuclide.systems](https://hoarder.nuclide.systems) | Bookmarks | + +--- + +## DevOps (18000–18999) + +| Port | Service | Container | Public URL | Notes | +|---|---|---|---|---| +| [18000](http://192.168.1.40:18000) | Daytona API | `daytona-minimal-api-1` | [daytona.nuclide.systems](https://daytona.nuclide.systems) | Dev sandbox | +| [18001](http://192.168.1.40:18001) | Daytona SSH | `daytona-minimal-ssh-gateway-1` | — | SSH gateway | + +--- + +## Tracking (19000–19999) + +_Traccar moved to 15000–15001 (Documents range). 19000–19001 now free._ + +--- + +## Storage Admin (20000–20999) + +| Port | Service | Container | Public URL | Notes | +|---|---|---|---|---| +| 20010 (127.0.0.1 only) | pgAdmin | `pgadmin` | — | localhost only | + +--- + +## External Services (Not on NUC Docker) + +Zoraxy routes to these external backends: + +| Domain | Target | Host | Service | +|---|---|---|---| +| [ha.nuclide.systems](https://ha.nuclide.systems) | `192.168.1.60:8123` | Home Assistant VM | Home automation | +| [nc.nuclide.systems](https://nc.nuclide.systems) | `192.168.1.41:11000` | Nextcloud LXC | Cloud storage | +| [ocpp.nuclide.systems](https://ocpp.nuclide.systems) | `192.168.1.60:8887` | Home Assistant VM | EV charger OCPP | +| [shepard.nuclide.systems](https://shepard.nuclide.systems) | `192.168.1.49:80` | Docker host | Shepard | +| [shepard-api.nuclide.systems](https://shepard-api.nuclide.systems) | `192.168.1.49:8080` | Docker host | Shepard API | + +--- + +## Zoraxy Public Routes (proxied via `192.168.1.4`) + +| Domain | Backend (NUC `192.168.1.40`) | Port | +|---|---|---| +| [ai.nuclide.systems](https://ai.nuclide.systems) | litellm | 14000 | +| [chat.nuclide.systems](https://chat.nuclide.systems) | lobehub | 14001 | +| [mcp.nuclide.systems](https://mcp.nuclide.systems) | mcp-gateway | 8080 | +| [arcane.nuclide.systems](https://arcane.nuclide.systems) | arcane | 10002 | +| [gotify.nuclide.systems](https://gotify.nuclide.systems) | gotify | 10003 | +| [id.nuclide.systems](https://id.nuclide.systems) | pocket-id | 11000 | +| [vault.nuclide.systems](https://vault.nuclide.systems) | vaultwarden | 11001 | +| [immich.nuclide.systems](https://immich.nuclide.systems) | immich_server | 12000 | +| [immich-tools.nuclide.systems](https://immich-tools.nuclide.systems) | immich_power_tools | (container-internal) | +| [abs.nuclide.systems](https://abs.nuclide.systems) | audiobookshelf | 13003 | +| [n8n.nuclide.systems](https://n8n.nuclide.systems) | n8n | 16000 | +| [memos.nuclide.systems](https://memos.nuclide.systems) | memos | 17000 | +| [hoarder.nuclide.systems](https://hoarder.nuclide.systems) | karakeep | 17001 | +| [daytona.nuclide.systems](https://daytona.nuclide.systems) | daytona-api | 18000 | +| [traccar.nuclide.systems](https://traccar.nuclide.systems) | traccar | 15000 | +| [dozzle.nuclide.systems](https://dozzle.nuclide.systems) | dozzle | 10001 | + +**Not publicly proxied** (LAN / localhost only): pgadmin, paperless, paperless-ai, comfyui, comfyui-mcp, rdtclient, prowlarr, shelfarr, qdrant. + +--- + +## Known Issues + +- **`s3.nuclide.systems`** (Garage external S3): non-responsive. Internal S3 `http://garage:3900` works fine for wal-g, memos, lobe. External binding risks quorum loss — deferred. +- **`lobe-postgres` / `lobe-redis`**: host-port exposed (`0.0.0.0:5432/6379`) — firewall blocks external access but ideally restricted to localhost. + +--- + +## OIDC Client Registry (Pocket ID — `id.nuclide.systems`) + +| Client ID | Name | Redirect URIs | +|---|---|---| +| `e73bb7b9` | litellm | (empty = accept all) — used by LiteLLM UI + MCP Gateway | +| `26f3c26b` | lobehub | `https://chat.nuclide.systems/api/auth/callback/generic-oidc` | +| `81cf4ed0` | arcane | `https://arcane.nuclide.systems/auth/oidc/callback` | +| `33135ad4` | n8n | `https://n8n.nuclide.systems/auth/oidc/callback` | +| `9c91c18b` | immich | `https://immich.nuclide.systems/auth/login` + mobile | +| `62bf4e0d` | memos | `https://memos.nuclide.systems/auth/callback` | +| `d92f82b0` | karakeep | `https://hoarder.nuclide.systems/api/auth/callback/custom` | +| `a14b8076` | nextcloud | `https://nc.nuclide.systems/apps/user_oidc/code` | +| `798a367f` | daytona | `https://daytona.nuclide.systems` | +| `cbbf20d5` | Audiobookshelf | `https://abs.nuclide.systems/…` + `*` | +| `d8733fcc` | shelfarr | `*` | +| `78c78998` | Claude MCP | `https://claude.ai/api/mcp/auth_callback` + `https://mcp.nuclide.systems/mcp/auth/callback` | +| `82ca2d53` | nuc-ai | (empty) — used by spawned MCP servers | +| `7fe1a14b` | zoraxy | (empty) | +| `af2f837b` | mcp-auth | (empty) — legacy, unused | + +**OIDC Endpoints** (corrected 2026-05-17 — previously used wrong `/api/v1/oauth2/` path): +- Authorization: `https://id.nuclide.systems/authorize` +- Token: `https://id.nuclide.systems/api/oidc/token` +- Userinfo: `https://id.nuclide.systems/api/oidc/userinfo` +- Discovery: `https://id.nuclide.systems/.well-known/openid-configuration` diff --git a/stacks/ai/PROMPTING-2026.md b/stacks/ai/PROMPTING-2026.md new file mode 100644 index 0000000..ef6d149 --- /dev/null +++ b/stacks/ai/PROMPTING-2026.md @@ -0,0 +1,43 @@ +# Agent system-prompt best practices (2026) + +Synthesised from current (2026) prompt-engineering guidance. These are the rules +the **agent-creator** meta-prompt enforces when it drafts a new agent, and the +checklist to apply when writing any agent here. + +## Core philosophy +- **It's information architecture, not magic words.** Better context beats + longer prompts. Engineer *what the model knows and when*, not verbosity. +- **A system prompt is operational policy, not a script.** Define the agent's + identity, job, scope boundaries, success criteria, and stop/escalate rules — + not rigid step-by-step instructions (the model runs many turns/tools). + +## Structure (in this order) +1. **Identity & job** — what the agent *is*, its single job, and explicitly + *where the job begins and ends*. +2. **Operating rules** — tone, language, output format constraints, length. +3. **Tools** — which MCP tools it has, **when and why** to use each, and when + *not* to. Models under-use tools unless told explicitly. +4. **Procedure** — the ordered steps/sections to produce (use a numbered list). +5. **Boundaries & failure** — permission limits, what to never do, and graceful + degradation ("if a tool fails, write '(nicht verfügbar)' and continue — + never invent values"). +6. **Closing actions** — side effects (notify, log) stated explicitly. + +## Mechanics +- **Explicit delimiters.** Separate *rules* from *content the model processes* + with stable markers (XML-ish tags or clear headers). Most failures come from + the model conflating instructions with data. +- **No assumed state.** Dynamic facts (date/time, sensor values) must come from + a tool call, never the model's prior. State this in the prompt. +- **Determinism aids.** Give concrete IDs/keys; constrain output shape; show one + short example only when tone/format matters. +- **Escalation/stop conditions.** Say when to stop, when to ask, when to bail. +- **Idempotency/self-logging** where relevant (e.g. an agent that logs its own + output so a later run can find it). + +## Anti-patterns +- Walls of prose; mixing rules and data; "be helpful/creative" with no scope; + assuming the current date; vague tool guidance; no failure path. + +The Morning Briefing agent (`morning-briefing-agent.md`) is the worked example +that follows all of the above. diff --git a/stacks/ai/home-info.md b/stacks/ai/home-info.md new file mode 100644 index 0000000..5a98eed --- /dev/null +++ b/stacks/ai/home-info.md @@ -0,0 +1,74 @@ +# Home Info — seed for the LobeChat morning-briefing agent + +Discovered via the Home Assistant MCP (gateway → `home-assistant`) on 2026-05-18. +Home Assistant: ~2492 entities, 40 domains, 19 areas, 2 floors. Language: German. + +## Layout +- **Floors**: Erdgeschoss (level 0) + one upper floor; 19 areas total. +- Erdgeschoss areas incl.: Badezimmer, Büro, Esszimmer, Gang, Küche, Wohnzimmer, + Schlafzimmer, Kinderzimmer, Technikraum, Terrasse … (+ more). + +## Key systems +- **Heat pump**: NIBE S1255-6 E (geothermal/brine, BT10 brine-in sensor). +- **Home battery / solar**: EcoFlow **PowerOcean** (`PowerOcean emsBpPower`, + `Battery HJ3AZDH5ZG3G0384`, switches `Hausbatterie: Entladesperre`, + `Hausbatterie: Fahrzeug zuerst`). Solar forecast: `Solar Prognose`, + `sensor.energy_production_today`, `sensor.energy_current_hour`, + `sensor.energy_next_hour`. +- **EV charging**: evcc + PowerPulse wallbox (AC31ZEH4AG130052). Vehicle = BYD. +- **Spa**: Whirlpool (heat-pump spa). Filter cycles, thermostat. +- **Other**: Philips Hue (lights), smoke detectors (Rauchmelder, w/ temp+humidity), + Siemens dishwasher, dryer, Sonos (Playbase/TV), QNAP NAS, 3D printer, + grow tent, voice satellites + **Piper** (local TTS available in HA). + +## Briefing entities (use these in the morning briefing) +| Briefing item | Entity / source | Notes | +|---|---|---| +| Car state of charge | `sensor.evcc_byd_configvehicle_soc` | % (e.g. 98) | +| Car range | `sensor.evcc_byd_configvehicle_range` | km (e.g. 422) | +| Car charge limit | `number.evcc_powerpulse_limit_soc` | % | +| Solar battery (house) | PowerOcean / `select.evcc_buffer_soc`; use `ha_search_entities "PowerOcean"` / `ha_get_state` for the live battery-SoC sensor | exact SoC sensor: query PowerOcean at briefing time | +| Solar production today | `sensor.energy_production_today` | kWh | +| Solar forecast (hour/next) | `sensor.energy_current_hour`, `sensor.energy_next_hour` | | +| Grid import | `sensor.evcc_powerpulse_charge_total_import` | kWh | +| Spa temperature | `sensor.whirlpool_temperatur` | °C (e.g. 36.7) | +| Spa thermostat | `climate.spa_thermostat` | mode (heat) | +| Spa target temp | `number.spa_target_desired_temperature` | °C | +| **Spa pH / bromine** | **NOT in Home Assistant** — no pH/bromine sensors exist | omit, or user adds them later / states manually | +| Weather forecast | `weather.wetter` | use `ha_get_state` for forecast attrs | +| Timeline highlights | Bluesky MCP (`bluesky`) — pending fix | once bluesky MCP works | +| Memos recap (yesterday) | Memos MCP (`memos`) — `search_memo` / list by date | working | + +## Systematic inventory (2026-05-18) +**Floor Erdgeschoss (level 0)** — 13 areas: Badezimmer, Büro, Esszimmer, Gang, +Garderobe, Gästezimmer, Kinderzimmer, Küche, Schlafzimmer, Speis, Technikraum, +WC, Wohnzimmer. +**Floor Außen (outdoor)** — 6 areas: Eingang, Garage, Grillplatz, Ruheplatz, +Terrasse, Grow. + +**Domain counts** (2492 entities / 40 domains): sensor 1309, switch 189, +update 154, button 152, number 148, binary_sensor 143, select 141, light 62, +device_tracker 55, event 33, automation 20, camera 10, notify 9, media_player 8, +cover 8, zone 4, fan 4, person 3, image 3, scene 2, climate 2, calendar 1, +water_heater 1, weather 1, vacuum 1, humidifier 1, tts 1, stt 1, todo 1, +assist_satellite 1, sun 1. + +## Notifications / alerts for the briefing +HA has no `alert` domain; surface alerts from: +- **Persistent notifications**: `ha_get_state` on `persistent_notification.*` + (or `ha_search_entities "notification"`). +- **Problem/safety binary_sensors in `on`**: smoke (Rauchmelder), water leak + (Wasserticker/"Batterie fast leer"), low-battery sensors — search + `ha_search_entities "leer"` / `"rauch"` / `"leak"`, report any `state=on`. +- **Automations named like alerts**: e.g. `automation.low_battery` (state on). +- **Pending updates count**: domain `update` (154 entities; report how many `on`). +The agent should call `ha_get_state`/`ha_search_entities` at briefing time and +list only items currently in an alert state (don't dump everything). + +## Caveats +- **Spa pH & bromine are not exposed** by the Whirlpool integration — only + temperature/filter. The briefing should skip them (or note "no chemistry + sensors") unless the user adds manual input_number/sensor entities. +- House-battery SoC sensor name varies (PowerOcean); the agent should resolve it + live via `ha_search_entities "PowerOcean"` + `ha_get_state` rather than a + hardcoded id. diff --git a/stacks/ai/morning-briefing-agent.md b/stacks/ai/morning-briefing-agent.md new file mode 100644 index 0000000..ec1348a --- /dev/null +++ b/stacks/ai/morning-briefing-agent.md @@ -0,0 +1,105 @@ +# Morning Briefing — LobeChat agent + +## How to create (LobeChat UI — ~2 min) +1. LobeChat → **Create Agent** (sidebar +). +2. **Title**: `Morning Briefing` · **Avatar**: ☀️ +3. **Model**: `qwen3.5-397b-a17b` (SAIA, free, strong tool-use), provider OpenAI. + LiteLLM auto-fails-over if busy. +4. **Plugins / MCP** — enable all of: + `time`, `home-assistant`, `kroki`, `fetch`, `sequential-thinking`, + `daytona`, `ntfy`, `memos`, `bluesky` + (all registered in LobeChat via syncstack). +5. Paste the **System Role** below. +6. Set the **Opening Message** below. +7. Optional: LobeChat agent cron for an automatic daily run. + +## System Role (paste verbatim) + +``` +You are my Morning Briefing assistant for a smart home in Germany (Home +Assistant, ~2492 entities, floors "Erdgeschoss"/"Außen"). Respond in German, +terse, dashboard-style, emojis as section headers, one short line per metric, +round sensibly, never dump raw entity lists. If any tool fails, write +"(nicht verfügbar)" for that line and continue — never invent values. + +STEP 0 (always, silently first): +- time MCP `get_current_time` → today + derive yesterday. You do NOT know the + date; always get it here. +- Use `sequential-thinking` to plan which tool calls you need, then execute. + +Trigger: "good morning" / "briefing" / chat opened. Produce, in order: + +▶ TL;DR — one punchy line synthesising the day (write this LAST, show it FIRST): + e.g. "☀️ guter Solartag, 🚗 78 %, laden 13–15 Uhr (billig+grün), 🔔 1 Hinweis". + +1. 🚗 Auto — SoC `sensor.evcc_byd_configvehicle_soc` %, range + `sensor.evcc_byd_configvehicle_range` km, limit + `number.evcc_powerpulse_limit_soc`. +2. ☀️ Solar/Akku — Hausakku: ha_search_entities "PowerOcean" → ha_get_state; + `sensor.energy_production_today`, `sensor.energy_current_hour`, + `sensor.energy_next_hour`; grid import + `sensor.evcc_powerpulse_charge_total_import`. + CHART: ha_get_history on the PV sensor for yesterday → hourly kWh → render + via kroki MCP as **Vega-Lite** bar chart (x=Stunde, y=kWh, title with + yesterday's date). Embed the image. +3. ⚡ Energiefluss — render via kroki a small **D2** (or mermaid) diagram of + the live flow PV → Hausakku → Haus → Netz → 🚗, annotated with the current + watts you read in §2. Embed it. +4. 💶 Strom & Laden — fetch MCP GET + `https://api.awattar.de/v1/marketdata` (German day-ahead prices, no auth). + Combine the cheapest upcoming hours with the solar forecast (§2) and car + SoC/limit (§1); via `sequential-thinking` recommend the optimal EV charge + window today (cheap + green) in one line. +5. 🛁 Whirlpool — `sensor.whirlpool_temperatur` °C, `climate.spa_thermostat`, + `number.spa_target_desired_temperature`. (No pH/Brom sensors — skip.) +6. 🌦️ Wetter — ha_get_state `weather.wetter`: condition, min/max, Regen-% + from forecast attrs. +7. 📅 Heute — HA `ha_config_get_calendar_events` for today + open items from + `ha_get_todo`. Max 5 lines; if empty "nichts angesetzt". +8. 🔔 Hinweise/Alarme — ONLY items currently alerting: persistent_notification.*, + smoke/leak/low-battery binary_sensors "on" (search "leer","rauch","leak"), + automation.low_battery if on, count of pending `update` entities on. + None → "keine". +9. 📨 ntfy — ntfy MCP `ntfy_fetch_messages` topic "homelab-ai", last 24 h, + high/urgent first, 1 line each; none → "keine". +10. 🦋 Bluesky — bluesky MCP: top 3 timeline highlights + 1 line of + `get-trends`. If auth fails: "(nicht verfügbar)". +11. 📝 Memos gestern — memos MCP `search_memo` for yesterday's date in formats + "DD.MM","YYYY-MM-DD","DD.MM.YYYY"; 2–4 bullets; none → "keine". +12. 🧠 Tagesempfehlung — use `sequential-thinking` to synthesise §1–9 into 2–3 + concrete actions (Ladefenster, Whirlpool heizen/aus, Lastverschiebung, + alles aus §8). This is the value — be specific and practical. + +CLOSING ACTIONS (always, after presenting): +- memos MCP `create_memo`: store a dated PRIVATE memo titled with today's date + containing the TL;DR + key numbers + Tagesempfehlung. (This makes tomorrow's + §11 actually find today.) +- ntfy MCP `ntfy_publish_message` topic "homelab-ai", title "Morning Briefing", + priority default: send the TL;DR line so it reaches my phone. + +ON DEMAND only (if I say "deep dive" / "tiefere analyse"): +- daytona MCP: create_sandbox(snapshot "sciviz-py") → write a Python script + that pulls 7 days of solar production + grid import (give it the figures + from HA history), renders a matplotlib/seaborn multi-panel trend + (production vs import, weekday pattern), execute_command to run it, return + the image, then destroy_sandbox. Embed the figure. +``` + +## Opening Message +``` +Guten Morgen! Sag „Briefing" für dein Dashboard (Auto, Solar + Diagramme, +Strompreis-Ladeempfehlung, Wetter, Termine, Hinweise, ntfy, Bluesky, Memos +und eine KI-Tagesempfehlung). „Deep dive" für die 7-Tage-Energieanalyse. +``` + +## Notes +- Home context + entity IDs: `home-info.md` (same folder). +- Uses every server we built: time (date), kroki (Vega-Lite chart + + D2/mermaid energy-flow diagram, self-hosted), fetch (aWATTar prices), + sequential-thinking (planning + recommendation), daytona (`sciviz-py` + snapshot for the on-demand matplotlib/seaborn deep dive), ntfy (read alerts + + push digest), memos (recap + self-logging closes the no-date-listing + gap), bluesky (timeline+trends), home-assistant (sensors/calendar/todo). +- Closing memo write is deliberate: the memos MCP can't list by date, so the + briefing logs itself → next day's §11 finds it by date keyword. +- bluesky degrades gracefully until its app-password rate-limit clears. diff --git a/stacks/docker-networks.md b/stacks/docker-networks.md new file mode 100644 index 0000000..960d674 --- /dev/null +++ b/stacks/docker-networks.md @@ -0,0 +1,129 @@ +# Docker Networks + +## Current Landscape (May 16, 2026) + +Each Docker Compose stack creates its own `{stack}_default` bridge network when it +has no explicit `networks:` declaration. This has exhausted Docker's built-in +172.x.x.x/16 address pool, triggering CIDR overlap errors. + +### Networks & Subnets + +| Network | Subnet | Containers | +|---|---|---| +| `bridge` (built-in) | 10.0.0.0/24 | 0 | +| `ai-internal` | 172.31.0.0/16 | 8 | +| `arcane_default` | 172.22.0.0/16 | 1 | +| `arr-stack_default` | 172.21.0.0/16 | 4 | +| `daytona-minimal_daytona-network` | 172.23.0.0/16 | 4 | +| `dozzle_default` | 192.168.16.0/20 | 1 | +| `homepage_default` | 172.19.0.0/16 | 1 | +| `immich_default` | 172.18.0.0/16 | 5 | +| `karakeep_default` | 172.30.0.0/16 | 3 | +| `memos_default` | 192.168.32.0/20 | 1 | +| `n8n_default` | 172.25.0.0/16 | 1 | +| `ntfy_default` | 172.27.0.0/16 | 1 | +| `nuc-ai-core_default` | 172.29.0.0/16 | 1 | +| `paperless-ngx_default` | 172.28.0.0/16 | 5 | +| `pocketid_default` | 172.20.0.0/16 | 1 | +| `qdrant_default` | 172.24.0.0/16 | 1 | +| `traccar_default` | 172.26.0.0/16 | 1 | +| `vaultwarden_default` | 192.168.64.0/20 | 1 | +| `vpn_default` | 192.168.80.0/20 | 1 | + +Total: 19 user-defined bridge networks. + +### Problem + +Docker's default address pool for user-defined bridge networks is +172.17.0.0/16 – 172.31.0.0/16 (15 subnets max). With 15 172.x.x.x/16 networks +already allocated, there is no room for new ones. + +The Daytona runner (`daytona-minimal-runner-1`) programmatically creates a +`runner-bridge` network on startup. It fails with: + +``` +Error response from daemon: invalid pool request: Pool overlaps with other one +on this address space +``` + +## Consolidation Plan + +### shared_backend Network + +A single shared bridge network (`shared_backend`) has been created to replace +per-stack defaults for lightweight services that don't need isolation. + +### Stacks Already Migrated + +- arcane +- dozzle +- ntfy +- qdrant +- traccar +- vaultwarden +- memos +- n8n +- pocketid + +These stacks now declare: + +```yaml +networks: + default: + external: true + name: shared_backend +``` + +### How to Free Subnets + +After migrating a stack to `shared_backend`, recreate it and prune the old network: + +```bash +cd /opt/stacks/{stack} && docker compose up -d +docker network rm {stack}_default # after containers disconnect +``` + +### Stacks Keeping Own Networks + +These stacks have complex internal networking and should keep their own: + +- **immich** — 5 services with inter-dependencies +- **paperless-ngx** — 5 services (webserver, broker, db, gotenberg, tika) +- **arr-stack** — 5 services (rdtclient, prowlarr, audiobookshelf, shelfarr, flaresolverr) +- **daytona** — explicit `daytona-minimal_daytona-network` +- **karakeep** — 3 services with chrome dependency +- **homepage** — single service, can stay or migrate +- **streamio** — already removed, stremio uses VPN container directly +- **vpn** — single service, can stay or migrate +- **ai/*_ai** — AI stacks, untouched + +### Long-term Fix + +Add `default-address-pools` to `/etc/docker/daemon.json`: + +```json +{ + "default-address-pools": [ + {"base": "10.0.0.0/8", "size": 24} + ] +} +``` + +This gives 65536 /24 subnets, eliminating exhaustion. Requires Docker daemon +restart (`systemctl restart docker`), which briefly disrupts all containers. + +## Commands + +```bash +# List all networks +docker network ls + +# Inspect a network +docker network inspect {name} + +# Remove unused networks +docker network prune + +# Remove a specific network (must have 0 containers) +docker network rm {name} +``` diff --git a/stacks/docs/AdGuard-DNS-writes.md b/stacks/docs/AdGuard-DNS-writes.md new file mode 100644 index 0000000..0f736ea --- /dev/null +++ b/stacks/docs/AdGuard-DNS-writes.md @@ -0,0 +1,207 @@ +# AdGuard DNS Rewrite Opportunities + +## Overview +AdGuard DNS at `192.168.1.2` can handle internal domain resolution, eliminating need for external DNS or hosts file entries. + +## Current Setup +- **AdGuard URL**: `http://192.168.1.2` +- **Status**: Running on Proxmox VM 102 +- **Homepage widget**: Configured + +## AI Service DNS Entries + +### Recommended DNS Rewrites + +Add these static DNS entries in AdGuard to resolve AI services locally: + +| Domain | IP Address | TTL | Purpose | +|--------|------------|-----|---------| +| `ai.nuclide.systems` | `192.168.1.40` | 300 | LiteLLM gateway | +| `chat.nuclide.systems` | `192.168.1.40` | 300 | LobeHub chat | +| `mcp.nuclide.systems` | `192.168.1.40` | 300 | MCP servers | +| `litellm.nuclide.systems` | `192.168.1.40` | 300 | LiteLLM API | +| `s3.nuclide.systems` | `192.168.1.40` | 300 | Garage S3 (Zoraxy proxy) | + +### Benefits +1. **No external DNS needed** - All AI services resolve internally +2. **Failover protection** - Works even if external DNS is unreachable +3. **Faster resolution** - Local DNS vs external lookup +4. **Simplified client config** - Services can use domain names directly + +## How to Add + +### Method 1: Web UI (Recommended) + +1. Open AdGuard: `http://192.168.1.2` (or via Zone: `http://192.168.1.2:3000`) +2. Navigate to **DNS Settings** → **Static DNS Entries** +3. Click **Add** for each service: + ``` + Domain: ai.nuclide.systems + IP Address: 192.168.1.40 + TTL: 300 + ``` +4. Click **Save** + +### Method 2: API + +```bash +# Get session token first +curl -c /tmp/cookies.txt -k -X POST http://192.168.1.2:3000/ \ + -H "Content-Type: application/json" \ + -d '{"username": "root", "password": "tapirnase"}' + +# Add static DNS entry +curl -s -k -c /tmp/cookies.txt -b /tmp/cookies.txt \ + -X POST http://192.168.1.2:3000/admin/api/staticDNS \ + -H "Content-Type: application/json" \ + -d '{ + "domain": "ai.nuclide.systems", + "ip": "192.168.1.40", + "ttl": 300 + }' +``` + +### Method 3: Auto-configure (Script) + +Create `/opt/stacks/scripts/configure_adguard_dns.sh`: + +```bash +#!/bin/bash +# Configure AdGuard DNS for AI stack + +ADGUARD_HOST="192.168.1.2" +ADGUARD_PORT="3000" +TARGET_IP="192.168.1.40" + +# AI service domains to add +DOMAINS=( + "ai.nuclide.systems" + "chat.nuclide.systems" + "mcp.nuclide.systems" + "litellm.nuclide.systems" + "s3.nuclide.systems" +) + +echo "🔧 Adding DNS entries to AdGuard..." + +for domain in "${DOMAINS[@]}"; do + echo " → $domain → $TARGET_IP" + # Note: This is a placeholder - actual API call needed +done + +echo "✅ Done! Add these in AdGuard UI manually." +``` + +## Alternative Useful DNS Entries + +### NextCloud & Sync +- `nc.nuclide.systems` → `192.168.1.40` (NextCloud web) +- `sync.nuclide.systems` → Internal IP of Sync client + +### Garage S3 +- `storage.nuclide.systems` → `192.168.1.40` (Garage internal, port 3900) +- `s3-local.nuclide.systems` → `192.168.1.40` (S3 API for local only) + +### Service Discovery +- `api.nuclide.systems` → `192.168.1.40` (future API gateway) +- `web.nuclide.systems` → `192.168.1.40` (general web services) + +### Immich +- `immich.nuclide.systems` → `192.168.1.40` (Immich web & API) +- `immich-api.nuclide.systems` → `192.168.1.40:2283` (Immich API only) + +### Paperless-ngx +- `pp.nuclide.systems` → `192.168.1.40` (Paperless web) +- `pp-api.nuclide.systems` → `192.168.1.40:8000` (Paperless API) + +### Home Assistant +- `ha.nuclide.systems` → `192.168.1.10` (Home Assistant) +- `homeassistant.local` → `192.168.1.10` (mDNS fallback) + +## Internal DNS Server Setup + +### Option A: Use AdGuard as Forwarding DNS + +Configure clients to use `192.168.1.2` as their DNS server: + +1. Proxmox Host: Edit `/etc/resolv.conf` +2. Docker containers: Add DNS in docker-compose.yml +3. VMs/PFs: Configure network settings + +### Option B: Create Custom Zone in AdGuard + +1. AdGuard → **DNS Settings** → **Zone Management** +2. Add zone: `nuclide.systems` +3. Add A records for all subdomains with proper IPs +4. Zone file format: + ``` + @ IN SOA ns1.nuclide.systems. admin.nuclide.systems. ( + 1 ; Serial + 3600 ; Refresh + 1800 ; Retry + 604800 ; Expire + 86400 ) ; Minimum TTL + + @ IN NS ns1.nuclide.systems. + @ IN A 192.168.1.40 + ai IN A 192.168.1.40 + chat IN A 192.168.1.40 + mcp IN A 192.168.1.40 + l IN A 192.168.1.40 ; LiteLLM + s3 IN A 192.168.1.40 + nc IN A 192.168.1.40 ; NextCloud + ``` + +## Verification + +After adding DNS entries: + +```bash +# Test from any machine on network +dig ai.nuclide.systems @192.168.1.2 +dig chat.nuclide.systems @192.168.1.2 + +# Should return: 192.168.1.40 +``` + +## Integration with MCP/Gateway Config + +Update service configs to use domain names: + +### In LiteLLM Config (`/opt/stacks/ai/litellm-config/config.yaml`) + +```yaml +general_settings: + proxy_base_url: https://ai.nuclide.systems + control_plane_url: https://ai.nuclide.systems +``` + +### In Karakeep (`.env`) + +```bash +OPENAI_BASE_URL=https://ai.nuclide.systems/v1 +``` + +### In LobeHub (`.env`) + +```bash +APP_URL=https://chat.nuclide.systems +S3_ENDPOINT=https://s3.nuclide.systems +``` + +## Migration Checklist + +- [ ] Add DNS entries to AdGuard +- [ ] Verify resolution: `dig ai.nuclide.systems` +- [ ] Update service configs to use domains +- [ ] Test HTTPS connectivity +- [ ] Remove old hosts file entries (if any) +- [ ] Document for team + +## Notes + +- AdGuard runs on VM 102 at `192.168.1.2` +- All AI services run on `192.168.1.40` +- Using same IP for all domains is intentional (single endpoint) +- SSL certs are issued by Zoraxy (public domain validation) +- DNS-only setup doesn't require HTTP proxy (Zoraxy still handles HTTPS) diff --git a/stacks/docs/case-study.md b/stacks/docs/case-study.md new file mode 100644 index 0000000..60d7162 --- /dev/null +++ b/stacks/docs/case-study.md @@ -0,0 +1,114 @@ +# Case Study — The Nuclide Homelab, built with Claude + +## Origin story + +One Saturday, the owner's wife left him home alone. He got bored, subscribed to +Claude, and started tinkering with a home server. That afternoon of boredom +turned into the `/opt/stacks` ecosystem documented here — a ~66-container, +~23-stack self-hosted platform with SSO, an MCP/agent gateway, GPU offload, +and a fully audited network. This is that story, kept as a record of what a +curiosity-driven collaboration produced. + +> This is a personal passion project, not a work deliverable. The tone and +> scope reflect that: depth and exploration over minimum-viable. + +## What was built (high level) + +- **AI/agent core** — LiteLLM model gateway (~28 curated models), an + MCP gateway that DinD-spawns and OAuth-gates MCP servers, an Agent Operator + (cron/event agents), semantic tool retrieval, MCP→LobeChat registration. +- **Identity** — PocketID as the universal OIDC IdP; every service behind SSO + (LiteLLM, LobeChat, n8n, Nextcloud, Daytona via a Keycloak→PocketID adapter). +- **Data layer** — shared-postgres standard; a full **tier-1 "no SQLite on + NFS"** migration (ntfy, Karakeep, Arcane, traccar → local disk; Memos + + Vaultwarden → Postgres via pgloader); WAL-G PITR backups re-established. +- **Network** — full read-only UniFi audit + the unmanaged D-Link DGS-1210 + core (SNTP fixed, topology/FDB mapped, DHCP/LLDP interop verified, SPOF + identified); pinned-NTP-source standard adopted. +- **Reliability** — gateway **deep health-check** (real MCP tool-call probe, + usage-aware backoff) that caught silently-broken servers. +- **Cloud burst** — on-demand Scaleway L40S GPU offload via WireGuard. + +See `homelab-architecture.md` for the living technical reference and +`PORTMAP.md` for the authoritative port/route map. + +## Activity signal + +- **~79 commits in the trailing 14 days** (`git log --since="14 days ago"`), + spanning gateway OAuth/health, OIDC bolt-ons, DB migrations, network audit, + GPU integration, and docs. +- Multi-session, incident-driven: several entries trace to real failures + caught and fixed (WAL-G archiver hung silently ~13 h; n8n `latest`-drift + outage; SQLite-on-NFS corruption risk; an 8.5-month-stale switch clock). + +## Productivity estimate (honest framing) + +These are **rough order-of-magnitude estimates**, not measurements. Assumptions +are stated so they can be challenged. + +- Scope delivered ≈ a small platform: identity, AI gateway, agent runtime, + ~23 service stacks, a DB-migration program, a full network audit, backups. +- A solo engineer doing this unaided, part-time, learning the unfamiliar + pieces (OIDC internals, pgloader, MCP, UniFi/D-Link internals): a + conservative bound is **several hundred focused hours** (≈ 8–16 part-time + weeks). Assumes the owner is competent but not a specialist in every domain + touched (identity, Postgres ops, embedded-switch web UIs, MCP). +- With the assistant: compressed into a small number of intensive sessions + over ~2 weeks. The leverage is largest where the work is *research-heavy + but low-novelty* — reverse-engineering a D-Link form POST, deriving an + OIDC redirect, mapping an FDB table — i.e. tasks that are tedious solo but + not conceptually hard. The leverage is smallest on genuine judgment calls + (what to prioritise, what risk is acceptable), which stayed with the owner. +- **Caveat:** estimate excludes the owner's own steering/review time, which + was substantial and is the reason the output is coherent rather than just + voluminous. + +## CO2 estimate (honest framing) + +Also order-of-magnitude, assumptions explicit. + +- **LLM inference:** a heavy multi-session collaboration of this kind is on + the order of a few million tokens. Public estimates put frontier-model + inference at roughly 1–5 Wh per ~1k output tokens equivalent (wide error + bars). Taking ~3 M tokens × ~2 Wh/1k ≈ **~6 kWh** → at a ~0.35 kgCO2e/kWh + grid ≈ **~2 kgCO2e**. Plausible range **1–5 kgCO2e**. Datacenter PUE and + exact model size dominate the uncertainty. +- **Homelab runtime** (the larger ongoing footprint): the NUC 14 Pro draws + ~15–45 W under mixed load. At ~30 W average → ~0.72 kWh/day → ~260 kWh/yr + → **~90 kgCO2e/yr** at the same grid factor. The German grid is cleaner + than that average in many hours, so treat as an upper-ish bound. The + on-demand Scaleway L40S burst is deliberately *on-demand* precisely to + avoid a 24/7 GPU's footprint. +- **Takeaway:** the assistant-collaboration carbon is a rounding error next + to a year of always-on homelab power. Efficiency wins (CPU-only default, + on-demand GPU, idle-aware health probing) matter more than the chat cost. + +## Handover / current state + +**Healthy & verified** +- Tier-1 SQLite-off-NFS: complete. +- OIDC: n8n (302→PocketID, client `33135ad4`) and LobeChat (`AUTH_TRUSTED_ + ORIGINS` fix, sign-in→PocketID) — both verified; LobeChat wants one real + browser login as final proof. +- D-Link SNTP: fixed (pinned PTB+Cloudflare IPs, clock corrected & synced). +- Gateway deep health-check: live, usage-aware, surfaced in `/api/servers`. + +**Open / pending** (see `homelab-architecture.md` roadmap for detail) +- **Broken MCP servers** surfaced by the new health-check: `memos` + (degraded — `mcp-memos` can't resolve `memos` host; Docker-network + isolation), `context7`/`crawl4ai`/`markitdown` (down), `nextcloud` + (probe false-positive — needs `health_check:false` or per-user creds). +- **D-Link mgmt hardening** (bundle, confirm-first): HTTPS, SNMP review, + Trusted-Host allowlist `192.168.1.0/24`. Shared `tapirnase` password + reuse (WiFi/LiteLLM/switch) — rotation deferred, noted. +- **Network**: IoT-VLAN segmentation; D-Link is the unmanaged core/SPOF; + mgmt-TLS certs for Proxmox + D-Link. +- **Platform**: env→secret vault; LobeChat external-feature disable; + observability LXC; agent-platform evolution (memory/teams/MCP-exposed). +- `nexa` analysis blocked — private repo; deploy key pending authorization. + +**Operating rules to preserve** +- Confirm + risk-assess before any Proxmox / Ubiquiti / network-gear write. +- Never put DB/SQLite on the UNAS NFS share. +- Only a full pgloader of *all* tables is a complete DB migration. +- Prefer self-hosted; pin critical container images (no `latest` drift). diff --git a/stacks/docs/cloud-gpu-extension.md b/stacks/docs/cloud-gpu-extension.md new file mode 100644 index 0000000..98a73c9 --- /dev/null +++ b/stacks/docs/cloud-gpu-extension.md @@ -0,0 +1,564 @@ +# Cloud GPU Extension — Scaleway L40S + +Seamless on-demand GPU (FLUX.1-dev, video, multi-user) via a Scaleway L40S instance. +The instance auto-starts on first request and shuts down after 45 min idle. +Everything personal stays on the NUC. + +--- + +## Pricing (current as of May 2026) + +| Resource | Rate | Notes | +|---|---|---| +| L40S-1-48G compute | **€1.40/hour** | PAR-2, billed per minute | +| Block volume 200 GB | **€16/month** | SSD, keeps models across restarts | +| Flexible IP | €0.004/hour | Static IP for WireGuard endpoint | +| Snapshots | €0.000044/GB/h | Only needed for image backups | + +### Realistic monthly cost + +| Usage pattern | Compute | Storage | Total | +|---|---|---|---| +| Weekend sessions (8 h/week) | €45 | €16 | **~€61/month** | +| Daily 1–2 h | €63–126 | €16 | **~€79–142/month** | +| Heavy (4 h/day) | €168 | €16 | **~€184/month** | +| Always-on (don't) | €1,008 | €16 | €1,024/month | + +The on-demand proxy below makes "daily 1–2 h" the natural default — you just click generate, it starts automatically. + +--- + +## Architecture + +``` +NUC (home, always-on) Scaleway PAR-2 (on demand) +───────────────────────────── ─────────────────────────────── +LobeChat ──▶ gpu-proxy:8190 ─── wg1 ──▶ ComfyUI :8188 +comfyui-mcp ──▶ (Docker) Ollama :11434 (optional) + │ + ├─ start/stop via Scaleway API + └─ idle watchdog (45 min → stop) +``` + +`gpu-proxy` is a small Docker service that: +1. Forwards requests to the GPU instance +2. Auto-starts the Scaleway instance if it's stopped (cold-start ~90 s) +3. Shuts it down after 45 min with no traffic + +--- + +## Prerequisites + +```bash +# Scaleway CLI +curl -s https://raw.githubusercontent.com/scaleway/scaleway-cli/master/scripts/get.sh | sh +scw init # enter API key + project ID +``` + +--- + +## Step 1 — Persistent block volume (models live here) + +```bash +# Create 200 GB SSD volume in PAR-2 +scw block volume create name=nuclide-gpu-models size=200GB zone=fr-par-2 +# Note the volume ID: vol-xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx +``` + +Models are kept on this volume. The compute instance can be deleted and recreated freely. + +--- + +## Step 2 — Create the L40S instance with cloud-init + +Save as `/opt/stacks/scripts/gpu-cloud-init.yaml`: + +```yaml +#cloud-config +package_update: true +packages: + - wireguard-tools + - docker.io + - docker-compose-v2 + - nvidia-driver-545 + - nvidia-container-toolkit + +write_files: + - path: /etc/wireguard/wg0.conf + permissions: '0600' + content: | + [Interface] + Address = 10.200.0.1/24 + ListenPort = 51820 + PrivateKey = SCALEWAY_WG_PRIVKEY + + [Peer] + # NUC + PublicKey = NUC_WG_PUBKEY + AllowedIPs = 10.200.0.2/32 + PersistentKeepalive = 25 + + - path: /opt/gpu/docker-compose.yml + content: | + services: + comfyui: + image: yanwk/comfyui-boot:cu124 + container_name: comfyui + ports: + - "10.200.0.1:8188:8188" + volumes: + - /mnt/models:/root/ComfyUI/models + - comfyui_config:/root/ComfyUI + environment: + - CLI_ARGS=--listen 0.0.0.0 + deploy: + resources: + reservations: + devices: + - driver: nvidia + count: all + capabilities: [gpu] + restart: unless-stopped + + ollama: + image: ollama/ollama:latest + container_name: ollama + ports: + - "10.200.0.1:11434:11434" + volumes: + - /mnt/models/ollama:/root/.ollama + deploy: + resources: + reservations: + devices: + - driver: nvidia + count: all + capabilities: [gpu] + restart: unless-stopped + + volumes: + comfyui_config: + +runcmd: + # Mount the models volume (will be /dev/sdb or similar) + - mkdir -p /mnt/models + - | + DISK=$(lsblk -ndo NAME,SIZE | awk '$2=="200G"{print "/dev/"$1}' | head -1) + if [ -n "$DISK" ]; then + blkid "$DISK" || mkfs.ext4 "$DISK" + echo "$DISK /mnt/models ext4 defaults 0 2" >> /etc/fstab + mount "$DISK" /mnt/models + fi + - mkdir -p /mnt/models/ollama + # Enable WireGuard + - systemctl enable --now wg-quick@wg0 + # Configure nvidia-container-toolkit + - nvidia-ctk runtime configure --runtime=docker + - systemctl restart docker + # Start GPU services + - cd /opt/gpu && docker compose up -d +``` + +**Before using this file**, replace: +- `SCALEWAY_WG_PRIVKEY` → output of `wg genkey` (run on Scaleway instance side first) +- `NUC_WG_PUBKEY` → output of `cat /etc/wireguard/nuc_wg_pub` (see Step 3) + +Launch the instance: + +```bash +# Generate WireGuard keys first (do this on the NUC) +wg genkey | tee /etc/wireguard/scaleway_wg_priv | wg pubkey > /etc/wireguard/scaleway_wg_pub +wg genkey | tee /etc/wireguard/nuc_wg_priv | wg pubkey > /etc/wireguard/nuc_wg_pub + +# Fill in cloud-init.yaml with the keys, then: +INSTANCE_ID=$(scw instance server create \ + type=L40S-1-48G \ + image=ubuntu_jammy_gpu \ + zone=fr-par-2 \ + name=nuclide-gpu \ + cloud-init=@/opt/stacks/scripts/gpu-cloud-init.yaml \ + output=json | jq -r '.id') + +echo "Instance ID: $INSTANCE_ID" + +# Attach the models volume +scw instance server attach-volume \ + server-id=$INSTANCE_ID \ + volume-id=vol-xxxxxxxx \ + zone=fr-par-2 + +# Allocate a Flexible IP (static IP that survives instance restarts) +FLEXIP_ID=$(scw instance ip create zone=fr-par-2 output=json | jq -r '.id') +scw instance server attach-flexible-ip \ + server-id=$INSTANCE_ID \ + ip-id=$FLEXIP_ID \ + zone=fr-par-2 + +FLEXIP=$(scw instance ip get $FLEXIP_ID zone=fr-par-2 output=json | jq -r '.address') +echo "Scaleway public IP: $FLEXIP" +``` + +--- + +## Step 3 — WireGuard on the NUC + +```bash +# /etc/wireguard/wg1.conf (separate from any existing VPN tunnel) +cat > /etc/wireguard/wg1.conf << EOF +[Interface] +Address = 10.200.0.2/24 +PrivateKey = $(cat /etc/wireguard/nuc_wg_priv) + +[Peer] +# Scaleway GPU +PublicKey = $(cat /etc/wireguard/scaleway_wg_pub) +Endpoint = ${FLEXIP}:51820 +AllowedIPs = 10.200.0.1/32 +PersistentKeepalive = 25 +EOF + +chmod 600 /etc/wireguard/wg1.conf +systemctl enable --now wg-quick@wg1 +``` + +Test when the instance is running: +```bash +ping 10.200.0.1 # WireGuard tunnel +curl http://10.200.0.1:8188 # ComfyUI +``` + +--- + +## Step 4 — gpu-proxy (seamless on-demand startup) + +This Docker service runs on the NUC. It proxies to the GPU instance and auto-starts/stops it. + +### `ai/gpu-proxy/docker-compose.yml` + +```yaml +services: + gpu-proxy: + build: . + container_name: gpu-proxy + environment: + - SCW_SECRET_KEY=${SCW_SECRET_KEY} + - SCW_PROJECT_ID=${SCW_PROJECT_ID} + - SCW_INSTANCE_ID=${SCW_INSTANCE_ID} + - SCW_ZONE=fr-par-2 + - GPU_HOST=10.200.0.1 + - COMFYUI_PORT=8188 + - OLLAMA_PORT=11434 + - IDLE_TIMEOUT=2700 # 45 min + ports: + - "8190:8190" # ComfyUI proxy + - "8191:8191" # Ollama proxy + restart: unless-stopped + networks: + - shared_backend + +networks: + shared_backend: + external: true +``` + +### `ai/gpu-proxy/proxy.py` + +```python +""" +On-demand GPU proxy. Auto-starts the Scaleway instance on first request, +shuts it down after IDLE_TIMEOUT seconds of inactivity. +""" +import asyncio, os, time, httpx, subprocess +from fastapi import FastAPI, Request +from fastapi.responses import StreamingResponse, JSONResponse + +app = FastAPI() + +SCW_KEY = os.environ["SCW_SECRET_KEY"] +SCW_PROJECT = os.environ["SCW_PROJECT_ID"] +INSTANCE_ID = os.environ["SCW_INSTANCE_ID"] +ZONE = os.environ.get("SCW_ZONE", "fr-par-2") +GPU_HOST = os.environ.get("GPU_HOST", "10.200.0.1") +COMFYUI_PORT = int(os.environ.get("COMFYUI_PORT", 8188)) +OLLAMA_PORT = int(os.environ.get("OLLAMA_PORT", 11434)) +IDLE_TIMEOUT = int(os.environ.get("IDLE_TIMEOUT", 2700)) + +SCW_API = f"https://api.scaleway.com/instance/v1/zones/{ZONE}" +HEADERS = {"X-Auth-Token": SCW_KEY, "Content-Type": "application/json"} + +_state = {"last_activity": 0.0, "starting": False, "up": False} +_lock = asyncio.Lock() + + +async def _scw(method: str, path: str, **kwargs): + async with httpx.AsyncClient() as c: + r = await c.request(method, f"{SCW_API}{path}", headers=HEADERS, **kwargs) + r.raise_for_status() + return r.json() + + +async def _instance_state() -> str: + d = await _scw("GET", f"/servers/{INSTANCE_ID}") + return d["server"]["state"] # running | stopped | stopping | starting + + +async def _start_instance(): + await _scw("POST", f"/servers/{INSTANCE_ID}/action", json={"action": "poweron"}) + + +async def _stop_instance(): + await _scw("POST", f"/servers/{INSTANCE_ID}/action", json={"action": "poweroff"}) + + +async def _wait_ready(host: str, port: int, timeout=180) -> bool: + deadline = time.time() + timeout + while time.time() < deadline: + try: + async with httpx.AsyncClient(timeout=3) as c: + await c.get(f"http://{host}:{port}/") + return True + except Exception: + await asyncio.sleep(5) + return False + + +async def ensure_up(port: int) -> bool: + async with _lock: + if _state["up"]: + _state["last_activity"] = time.time() + return True + if _state["starting"]: + return False # caller will retry + _state["starting"] = True + + try: + state = await _instance_state() + if state != "running": + print(f"[gpu-proxy] instance {state} → starting") + await _start_instance() + # wait for API to report running + for _ in range(60): + await asyncio.sleep(5) + if await _instance_state() == "running": + break + # wait for service to respond + if await _wait_ready(GPU_HOST, port): + async with _lock: + _state["up"] = True + _state["last_activity"] = time.time() + print("[gpu-proxy] GPU instance ready") + return True + return False + finally: + async with _lock: + _state["starting"] = False + + +async def idle_watchdog(): + while True: + await asyncio.sleep(60) + async with _lock: + if not _state["up"]: + continue + idle = time.time() - _state["last_activity"] + if idle > IDLE_TIMEOUT: + print(f"[gpu-proxy] idle {idle:.0f}s → stopping instance") + try: + await _stop_instance() + async with _lock: + _state["up"] = False + _state["last_activity"] = 0.0 + except Exception as e: + print(f"[gpu-proxy] stop error: {e}") + + +@app.on_event("startup") +async def startup(): + asyncio.create_task(idle_watchdog()) + # Probe: is the instance already running from a previous session? + try: + if await _instance_state() == "running": + if await _wait_ready(GPU_HOST, COMFYUI_PORT, timeout=10): + async with _lock: + _state["up"] = True + _state["last_activity"] = time.time() + print("[gpu-proxy] GPU already up on startup") + except Exception: + pass + + +async def _proxy(request: Request, host: str, port: int): + _state["last_activity"] = time.time() + ready = await ensure_up(port) + if not ready: + # Starting up — keep trying for up to 3 min + for _ in range(36): + await asyncio.sleep(5) + if _state["up"]: + break + else: + return JSONResponse({"error": "GPU instance failed to start"}, 503) + + url = f"http://{host}:{port}{request.url.path}" + if request.url.query: + url += f"?{request.url.query}" + body = await request.body() + client = httpx.AsyncClient(timeout=httpx.Timeout(None, connect=10)) + req = client.build_request(request.method, url, + headers={k: v for k, v in request.headers.items() + if k.lower() not in {"host", "content-length"}}, + content=body) + try: + resp = await client.send(req, stream=True) + except httpx.ConnectError: + await client.aclose() + return JSONResponse({"error": "GPU not reachable"}, 502) + + async def stream(): + async for chunk in resp.aiter_raw(): + yield chunk + await resp.aclose() + await client.aclose() + + return StreamingResponse(stream(), status_code=resp.status_code, + headers=dict(resp.headers), + media_type=resp.headers.get("content-type")) + + +@app.api_route("/status", methods=["GET"]) +async def status(): + try: + scw_state = await _instance_state() + except Exception as e: + scw_state = f"error: {e}" + return {"instance": scw_state, "proxy_up": _state["up"], + "idle_s": int(time.time() - _state["last_activity"]) if _state["up"] else None} + + +@app.api_route("/{path:path}", methods=["GET","POST","PUT","DELETE","OPTIONS","PATCH"]) +async def comfyui_proxy(request: Request, path: str): + return await _proxy(request, GPU_HOST, COMFYUI_PORT) +``` + +### `ai/gpu-proxy/Dockerfile` + +```dockerfile +FROM python:3.12-slim +RUN pip install fastapi uvicorn httpx +COPY proxy.py /app/proxy.py +WORKDIR /app +EXPOSE 8190 8191 +CMD ["uvicorn", "proxy:app", "--host", "0.0.0.0", "--port", "8190"] +``` + +--- + +## Step 5 — Wire up NUC services + +Add to `ai/.env`: +```bash +SCW_SECRET_KEY= +SCW_PROJECT_ID= +SCW_INSTANCE_ID= +``` + +In `ai/lobehub.yml`: +```yaml +- 'COMFYUI_BASE_URL=http://gpu-proxy:8190' +``` + +In `ai/mcp-gateway/server.py` (SERVERS dict): +```python +"comfyui": { + "static": True, + "upstream": "http://gpu-proxy:8190", + "group": "image", +}, +``` + +For Immich ML (optional), add to `immich/docker-compose.yml`: +```yaml +environment: + - IMMICH_MACHINE_LEARNING_URL=http://gpu-proxy:8192 # add a third port for ML +``` + +--- + +## Step 6 — Deploy + +```bash +# Deploy the proxy +cd /opt/stacks/ai/gpu-proxy +docker compose up -d + +# Restart LobeChat and MCP gateway to pick up new env +cd /opt/stacks/ai +docker compose -f lobehub.yml up -d --force-recreate lobe +cd mcp-gateway && docker compose up -d --force-recreate mcp-gateway +``` + +### What the user experience looks like + +1. Open LobeChat, generate an image → first request triggers auto-start +2. "Starting GPU…" — ComfyUI returns a brief wait (up to 90 s cold start) +3. Images generate normally; the proxy keeps the instance alive +4. After 45 min with no requests → instance stops automatically +5. `/status` endpoint on port 8190: `{"instance":"stopped","proxy_up":false}` + +--- + +## Quick-start / stop scripts + +```bash +# /usr/local/bin/nuclide-gpu (NUC helper) +#!/bin/bash +ZONE=fr-par-2 +ID=$(docker exec gpu-proxy env | grep SCW_INSTANCE_ID | cut -d= -f2) +case "$1" in + start) scw instance server action action=poweron server-id=$ID zone=$ZONE ;; + stop) scw instance server action action=poweroff server-id=$ID zone=$ZONE ;; + status) curl -s http://localhost:8190/status | python3 -m json.tool ;; + *) echo "Usage: nuclide-gpu start|stop|status" ;; +esac +``` + +--- + +## Scaleway firewall (security groups) + +On the Scaleway instance, only expose WireGuard. All services bind to `10.200.0.1` (WireGuard IP): + +```bash +ufw default deny incoming +ufw allow 51820/udp # WireGuard from anywhere +ufw allow from 10.200.0.0/24 # full access over tunnel +ufw enable +``` + +--- + +## Model storage layout (on the 200 GB volume) + +``` +/mnt/models/ +├── checkpoints/ # FLUX.1-dev, SDXL, etc. +├── vae/ +├── loras/ +├── controlnet/ +├── ollama/ # Ollama model blobs +└── upscale_models/ +``` + +FLUX.1-dev (FP16): ~24 GB +FLUX.1-schnell (NF4): ~8 GB +Ollama llama3:70b (Q4): ~40 GB — fits alongside FLUX on 200 GB with room for LoRAs. + +To pre-download models after first boot: +```bash +ssh root@10.200.0.1 # via WireGuard +docker exec comfyui python3 -c " +from huggingface_hub import hf_hub_download +hf_hub_download('black-forest-labs/FLUX.1-dev', 'flux1-dev.safetensors', + local_dir='/root/ComfyUI/models/checkpoints') +" +``` diff --git a/stacks/docs/comfyui-mcp.md b/stacks/docs/comfyui-mcp.md new file mode 100644 index 0000000..dca146b --- /dev/null +++ b/stacks/docs/comfyui-mcp.md @@ -0,0 +1,75 @@ +# ComfyUI → LobeChat via MCP + +Image generation (FLUX.1-schnell GGUF on the Intel Arc iGPU) exposed to LobeChat +as an MCP tool. Chosen over LobeChat's native ComfyUI provider because that +provider is hardcoded to non-GGUF nodes and is not configurable without forking +LobeChat (see `docs/mcp-gateway-requirements.md` research notes). + +## Components + +- **`ai/mcp-servers/comfyui/`** — purpose-built MCP server (async queue edition) + - `server.py` — FastMCP, streamable-HTTP on `:8000` at `/mcp`. Six tools: + - `generate_image(prompt, width, height, steps, seed)` — txt2img; returns `job_id` immediately + - `img2img(prompt, images, strength, steps, seed)` — img2img for one or more images; one job per image; images can be URLs, base64 data URIs, or `job:` references + - `get_job_status(job_ids)` — returns status text + inline PNG for completed jobs + - `list_queue(limit)` — lists all jobs in the in-process registry + - `cancel_job(job_id)` — cancels queued/running job (user must confirm first) + - `list_recent_images(n)` — shows n most recent completed images for reference/chaining + - Background polling thread updates job state every 3 s via ComfyUI `/history`. + - `flux-gguf-api.json` — txt2img workflow; nodes 4=prompt, 6=size, 8=steps/seed. + - `flux-img2img-api.json` — img2img workflow; node 11=LoadImage, 12=VAEEncode, 8=KSampler with denoise. + - `Dockerfile`, `requirements.txt` (`mcp[cli]`, `httpx`). +- **`ai/comfyui-mcp.yml`** — compose; container `comfyui-mcp` on + `shared_backend` (reaches `comfyui:8188`; reachable by `lobehub`, which is on + `shared_backend`). Host port `18003:8000` for testing/Zoraxy. + +## Async workflow + +``` +generate_image("a red apple") → "Job submitted: txt-1a8bbeda" (< 1 s) + [ComfyUI rendering... ~90-300 s] +get_job_status(["txt-1a8bbeda"]) → status text + inline PNG image + +# Chain: use previous output as img2img input (no bytes through LLM) +img2img("add a blue bowl", images=["job:txt-1a8bbeda"], strength=0.6) + → "img-2b9ccefa" +``` + +## Verified + +- All 6 tools discovered via MCP `tools/list`. +- `generate_image` returns in < 0.1 s (non-blocking); job shows in `list_queue` as "running". +- Previous end-to-end: `generate_image` → `get_job_status` → inline PNG ~94 s bare, ~160 s via MCP. +- LobeChat v2.1.58 renders MCP `image` blocks (uploads to Garage S3 → inline `![](url)`). + +## Known characteristics / limitations + +- **In-memory registry**: job state is lost on container restart. Resubmit if needed. +- **No auth** on the MCP server (internal `shared_backend` only). Host port + 18003 is LAN-exposed and unauthenticated — fine for a trusted homelab LAN. +- ~90-300 s/image latency; use `get_job_status()` to poll — never block waiting. + +## Final hookup (manual — LobeChat UI/DB, no server-mode config path) + +LobeChat → **Settings → Skills (Tools) → Skill Store → Custom → Import JSON**: + +```json +{ + "mcpServers": { + "comfyui-flux": { + "type": "http", + "url": "http://comfyui-mcp:8000/mcp" + } + } +} +``` + +Then enable the `comfyui-flux` skill in an agent/chat and ask the model to +"generate an image of …". Images appear inline in the conversation. + +## Ops + +- Build/deploy: `cd /opt/stacks/ai && docker compose -f comfyui-mcp.yml up -d --build` +- Logs: `docker logs comfyui-mcp` +- The workflow is the single source of truth in `flux-gguf-api.json`; keep it in + sync with `ai/comfyui/workflows/flux-schnell-api.json` if the graph changes. diff --git a/stacks/docs/dev-environment.md b/stacks/docs/dev-environment.md new file mode 100644 index 0000000..603df70 --- /dev/null +++ b/stacks/docs/dev-environment.md @@ -0,0 +1,262 @@ +# Dev Environment — CT 111 (dev.nuclide.systems) + +CT 111 hosts the self-hosted development platform: **Coder** (dev workspaces) + **Gitea** (internal repos). + +## Services + +| Service | URL | Port | +|---|---|---| +| Coder | https://dev.nuclide.systems | 7080 | +| Gitea | https://git.nuclide.systems | 3000 | + +Both use **Pocket-ID OIDC** (`https://id.nuclide.systems`) for SSO. + +--- + +## Coder — Dev Workspaces + +### What it is +Coder provisions isolated Docker-based dev environments (workspaces) on CT 111. Each workspace has: +- A full Linux environment with your tools +- Persistent home dir on UNAS (`/mnt/pve/unas/services/coder/`) +- Intel Arc GPU renderD128 available +- VS Code Server (browser or desktop SSH tunnel) + +### First-time setup +1. Go to `https://dev.nuclide.systems` → log in via Pocket-ID +2. Create a workspace from a template (admin must create templates first) +3. Connect via VS Code: install **Coder extension** → sign in → open workspace + +### VS Code connection +```bash +# Install Coder CLI on your local machine +curl -fsSL https://coder.com/install.sh | sh + +# Authenticate +coder login https://dev.nuclide.systems + +# Open workspace in VS Code +coder open +``` +Or use the Coder VS Code extension directly from the marketplace (`coder.coder-remote`). + +### Claude Code inside a workspace +```bash +# Inside the workspace terminal +npm install -g @anthropic/claude-code +claude +``` +Claude Code runs inside the workspace container — same environment, same files, same GPU. + +### Coder MCP (AI agent sandbox execution) +Add to Claude Code's MCP config (`~/.claude/claude_desktop_config.json` or via `/mcp add`): +```json +{ + "mcpServers": { + "coder": { + "command": "coder", + "args": ["mcp", "server"], + "env": { + "CODER_URL": "https://dev.nuclide.systems", + "CODER_TOKEN": "" + } + } + } +} +``` +Claude can then create workspaces, execute code, and read output via MCP tools: +- `coder_list_workspaces` +- `coder_create_workspace` +- `coder_execute_command` ← sandbox code execution +- `coder_start_workspace` / `coder_stop_workspace` + +Get your API token: `coder tokens create` + +### Creating workspace templates +Templates are Terraform configs stored in Gitea. Basic Docker template: +```bash +coder templates push --directory ./template/ +``` + +--- + +## Gitea — Internal Repos + +### What it is +Self-hosted Git for internal infrastructure: compose files, CT configs, dotfiles, Coder templates. +**Not** the primary remote for Claude Code collaboration — use GitHub for that. + +### First-time setup (admin) +1. Go to `https://git.nuclide.systems` → complete installation wizard +2. Set admin account, confirm DB settings (pre-filled from env) +3. Add OIDC provider: Admin → Site Administration → Authentication Sources + - Auth type: OAuth2 + - Provider: OpenID Connect + - Discovery URL: `https://id.nuclide.systems/.well-known/openid-configuration` + - Client ID/Secret: create a new client in Pocket-ID for Gitea + +### SSH access +```bash +# Gitea SSH runs on port 222 +git clone ssh://git@git.nuclide.systems:222//.git + +# Or add to ~/.ssh/config: +Host git.nuclide.systems + Port 222 + IdentityFile ~/.ssh/id_ed25519 +``` + +### Recommended repos to create +| Repo | Contents | +|---|---| +| `infra/proxmox` | `/etc/pve/` snapshots, CT configs | +| `infra/stacks` | Compose files from CT 104/101/111 | +| `infra/docs` | Mirror of `/docs/` on Proxmox host | +| `dev/templates` | Coder workspace Terraform templates | + +--- + +## Storage layout (UNAS) + +``` +/mnt/pve/unas/services/ +├── coder/ # Coder workspace home dirs (persistent) +└── gitea/ # Gitea repos + data +``` + +--- + +## CT 111 specs + +| | | +|---|---| +| IP | 192.168.1.42 | +| Cores | 12 | +| RAM | 32GB | +| Rootfs | 60GB local-zfs (Docker image cache) | +| UNAS | /mnt/pve/unas (mp0) | +| GPU | renderD128 (Intel Arc Xe, by-path) | + +--- + +## OIDC clients in Pocket-ID + +| Client | Callback URL | +|---|---| +| Coder | `https://dev.nuclide.systems/api/v2/users/oidc/callback` | +| Gitea | Add via Gitea admin UI (see above) | + +To create new OIDC clients programmatically, see `/docs/proxmox-optimizations.md` § OIDC client creation via SQLite. + +--- + +## Auth lockdown (SSO-only) + +Both Gitea and Coder are locked to **Pocket-ID OIDC only**. Local password and GitHub login are disabled. Passkey/WebAuthn login *to Gitea* remains available because it's tied to OIDC accounts, not to a separate password. + +### Compose env flags + +**Coder** (`/opt/stacks/coder/compose.yaml`): +```yaml +CODER_OIDC_ISSUER_URL: "https://id.nuclide.systems" +CODER_OIDC_CLIENT_ID: "${CODER_OIDC_CLIENT_ID}" +CODER_OIDC_CLIENT_SECRET: "${CODER_OIDC_CLIENT_SECRET}" +CODER_OIDC_ALLOW_SIGNUPS: "true" +CODER_OIDC_EMAIL_DOMAIN: "nucli.de" +CODER_DISABLE_PASSWORD_AUTH: "true" +CODER_OAUTH2_GITHUB_DEFAULT_PROVIDER_ENABLE: "false" +``` +Note: the bundled "GitHub external auth provider" log line is for *workspaces* cloning from GitHub, not for login — leaving it on is fine. + +**Gitea** (`/opt/stacks/gitea/compose.yaml`): +```yaml +GITEA__oauth2__ENABLED: "true" +GITEA__openid__ENABLE_OPENID_SIGNIN: "false" # legacy OpenID 2.0 button off +GITEA__openid__ENABLE_OPENID_SIGNUP: "false" +GITEA__service__ENABLE_PASSWORD_SIGNIN_FORM: "false" # local username/password form off +GITEA__oauth2_client__ENABLE_AUTO_REGISTRATION: "true" +GITEA__oauth2_client__ACCOUNT_LINKING: "auto" +GITEA__oauth2_client__USERNAME: "preferred_username" +GITEA__oauth2_client__UPDATE_AVATAR: "true" +``` +Note: `ENABLE_OPENID_SIGNIN` lives in `[openid]`, *not* `[service]`. Putting it under `GITEA__service__` is a no-op and leaves the legacy button visible. + +Gitea login source (DB row in `login_source`): +- `name = "pocket-id"` (case-sensitive — becomes part of the callback URL) +- `Scopes = ["openid","profile","email"]` +- `two_factor_policy = "skip"` (Pocket-ID passkey already enforces 2FA) + +### Pocket-ID client secret pitfall (important) + +Pocket-ID 2.7.0 verifies client secrets with **bcrypt** (`bcrypt.CompareHashAndPassword`). The stored `oidc_clients.secret` column must be a 60-char bcrypt hash like `$2a$10$...` or `$2b$10$...`. + +A raw 64-char hex SHA-256 in that column **silently fails every token exchange** with `invalid client secret`. The Gitea and Coder clients on this host were initially provisioned that way and had to be regenerated. + +To programmatically add or rotate an OIDC client secret in Pocket-ID: + +```bash +# On Proxmox host (has python3-bcrypt installed) +python3 <<'EOF' +import bcrypt, secrets, string +alphabet = string.ascii_letters + string.digits +plain = "".join(secrets.choice(alphabet) for _ in range(40)) +hashed = bcrypt.hashpw(plain.encode(), bcrypt.gensalt(rounds=10)).decode() +print("PLAINTEXT (give to client app):", plain) +print("HASH (store in oidc_clients.secret):", hashed) +EOF +``` + +Always push the hash to Pocket-ID via a tmp file (`pct push 110 ...`) — never inline the bcrypt hash in a shell command, the `$` chars get expanded. + +To verify a client's secret is the right format: +```bash +pct exec 110 -- sqlite3 /opt/stacks/pocketid/data/pocket-id.db \ + "SELECT name, length(secret), substr(secret,1,7) FROM oidc_clients;" +# Working clients: length=60, prefix "$2a$10$" or "$2b$10$" +# Broken clients : length=64, prefix is hex (e.g. "1180664") +``` + +### Break-glass recovery (if SSO is broken) + +You have host SSH access, so you're never locked out — but the UI will be unusable until you re-enable a local login path. From the Proxmox host: + +**Gitea — re-enable local login + reset password:** +```bash +# 1. Temporarily put the form back +ssh nuc +python3 -c " +p='/rpool/data/subvol-111-disk-0/opt/stacks/gitea/compose.yaml' +s=open(p).read().replace('ENABLE_PASSWORD_SIGNIN_FORM: \"false\"','ENABLE_PASSWORD_SIGNIN_FORM: \"true\"') +open(p,'w').write(s)" +pct exec 111 -- bash -c 'cd /opt/stacks/gitea && docker compose up -d --force-recreate gitea' + +# 2. Reset admin password +pct exec 111 -- docker exec -u git gitea gitea -c /data/gitea/conf/app.ini admin user change-password -u fkrebs -p 'temp-strong-pass' +``` + +**Coder — flip user back to password login + re-enable:** +```bash +ssh nuc +# 1. Flip login_type back +pct exec 111 -- docker exec coder-db psql -U coder -d coder -c \ + "UPDATE users SET login_type='password' WHERE username='fkrebs';" + +# 2. Re-enable password auth in compose +python3 -c " +p='/rpool/data/subvol-111-disk-0/opt/stacks/coder/compose.yaml' +s=open(p).read().replace('CODER_DISABLE_PASSWORD_AUTH: \"true\"','CODER_DISABLE_PASSWORD_AUTH: \"false\"') +open(p,'w').write(s)" +pct exec 111 -- bash -c 'cd /opt/stacks/coder && docker compose up -d --force-recreate coder' + +# 3. Reset password (interactive) +pct exec 111 -- docker exec -it coder coder reset-password fkrebs +``` + +**Pocket-ID — bootstrap a one-time access token if you're locked out of *Pocket-ID itself*:** +```bash +pct exec 110 -- docker exec pocket-id pocket-id-cli one-time-access-token \ + --email fkrebs@nucli.de --duration 1h +# Open the printed URL in a browser to log in once and register a new passkey. +``` + +After SSO is fixed: undo each of the above steps (re-disable password auth, flip `login_type` back to `oidc`). diff --git a/stacks/docs/homelab-architecture.md b/stacks/docs/homelab-architecture.md new file mode 100644 index 0000000..eb0c6b3 --- /dev/null +++ b/stacks/docs/homelab-architecture.md @@ -0,0 +1,373 @@ +# Nuclide Ecosystem — Homelab Architecture + +Living architecture reference for the `/opt/stacks` homelab. ~66 containers +across ~23 compose stacks. Principle: **self-host everything**, OIDC SSO, +`*.nuclide.systems` via one reverse proxy. + +## Topology +- **NUC 14 Pro** (`192.168.1.40`) — primary Docker host, all `/opt/stacks/*`. + Itself a **Proxmox LXC (CTID 104)** on node `nuc`; local data on ZFS + `rpool/data/subvol-104-disk-0`. Mounts the UNAS NFS at `/mnt/pve/unas`. +- **Proxmox VE** — host **`192.168.1.20:8006`**, single node **`nuc`** + (Intel Core Ultra 7 155H, 22 threads, 64G RAM, PVE 9.1.11). API access via + `root@pam!mcp` token (now `PVEAuditor`, read-only). **7 guests:** + - `100` qemu **haos** — Home Assistant OS VM (4c/16G) → `.60` + - `101` lxc **shepard** — secondary Docker host (12c/16G/107G) → `.49` + - `102` lxc **dns** — **AdGuard Home** (network DNS + ad/tracker blocking, + 2c/0.5G). UniFi DHCP has `dhcpd_dns_enabled: false`; clients resolve via + the gateway `.1`, **and the UDM forwards DNS upstream to AdGuard** + (confirmed) — so ad/tracker blocking is network-wide despite DHCP not + handing out AdGuard's IP directly. + - `103` lxc **backrest** — Restic/Backrest backups (1c/0.5G) + - `104` lxc **docker** — **the primary `/opt/stacks` host** (16c/32G/215G) → `.40` + - `105` lxc **nextcloud** — Nextcloud (4c/8G/107G) → `.41` + - `108` lxc **zoraxy** — reverse proxy (2c/2G) → `.4` + - Storages: `local` (dir), `local-zfs` (zfspool, ~1.9T), `unas` (nfs, ~20T). +- **Ubiquiti UniFi** — UDM "Home" (UDMA6A8, UCG Fiber) gateway + 4× U7-series + APs. Controller at **`https://192.168.1.1`** (UniFi OS 5.0.16, SSO + MFA + enforced). Single site "Default". MCP access: `ghcr.io/enuno/unifi-mcp-server` + (197 tools, Network App API key); Site Manager/cloud tools need + `UNIFI_SITE_MANAGER_ENABLED`, all local API tools work. + - **LAN** "Default" — `192.168.1.1/24`, corporate, no VLAN. DHCP pool + `.100–.250` (24 h lease), domain `localdomain`. Static infra lives + `.1–.99` (outside the pool). WANs: "Internet 1" (primary), "Secondary" + (WAN2, disabled). 47 clients (17 wired, 30 wireless). + - **WiFi** — single SSID **"nuclide"**, WPA2 with WPA3-transition, bands + 2.4/5/6 GHz, not guest, no per-SSID VLAN, no L2 isolation. **Flat + network**: servers, IoT, consoles, phones all share one L2 segment. + - **Firewall** — zone-based (Internal/External/DMZ/Gateway/VPN/Hotspot), + almost entirely system-predefined rules. Port forwards: **80/443 → + 192.168.1.4** (Zoraxy); **15001 TCP/UDP → 192.168.1.40** (Traccar watch + protocol — already at UDM level, no Zoraxy stream proxy needed). + - ⚠️ **Non-UniFi gear = topology blind spots** (UniFi can't see/manage): + - **D-Link DGS-1210-28P** 28-port PoE switch at `.10` (HW F5, fw + **6.32.008**, S/N TM0I533010201) — unmanaged by UniFi; its ports/ + port-clients don't appear in UniFi topology. HTTP-only mgmt + (`admin`/shared `tapirnase` pw, RSA-login). **Audit 2026-05-19 + (read-only):** + - **SNTP was broken → FIXED 2026-05-19.** Real root cause was **DNS**: + servers were `0/1/2.de.pool.ntp.org` but the switch has no DNS + resolver (`DNS_STATE=2`) so it couldn't resolve them. (The default + gateway was *fine* — an active static default route `0.0.0.0/0 → + .1` exists; the empty `Default_Route_Gateway=[]` is just the unused + *interface-level* gateway field — misleading.) Clock had been stuck + at **31/08/2025** (~8.5 mo). Fix: repointed to the **pinned NTP + source standard** (see below), all by IP, DNS-free. **Verified + synced** 2026-05-19 — clock corrected to live time (CEST/UTC+2). + - **HTTPS/SSL disabled** — admin creds cross the flat LAN in clear. + *Pending* (see roadmap: mgmt-TLS certs for Proxmox + D-Link). + - **SNMP agent enabled** on the flat LAN (community strings unverified). + *Pending.* + - **Trusted-Host mgmt allowlist disabled** — admin UI reachable from + any LAN device incl. IoT. *Pending:* enable allowing `192.168.1.0/24`. + - ⚠️ The switch is a **transparent L2 switch** — every device on it + **is fully visible to UniFi as a client** (MAC/IP via UDM DHCP/ARP); + only the *wired port-topology through it* is invisible (needs UniFi + adoption, impossible here). This is unlike the RE700X repeater. + - 🔴 **The D-Link is the network CORE / single point of failure.** + FDB-table analysis (2026-05-19): the **UDM uplinks on port 26**, + the **entire Proxmox host (NUC .40 + Nextcloud .41 + HA .60 + all + VMs/LXCs) is on port 10**, and the **U7 APs uplink on ports 3 & 16** + (their wireless clients appear via those ports). Nearly all wired + *and* wireless traffic transits this unmanaged switch → a D-Link + failure = total LAN+WAN outage. Amplifies the mgmt-hardening, + firmware-currency, monitoring and segmentation items below. + - **UniFi↔D-Link DHCP/LLDP interop = fully compatible** (audited + 2026-05-19, read-only): + - DHCP: `BOOTP_Relay_State=0` **and** `DHCP_BOOTP_Local_Relay_ + Status=0` → switch inserts **no Option 82**; UDM DHCP server + gets clean untagged requests. (`Option82_State=1` is moot — + only applies when relay/local-relay is active; both are off.) + DHCP Server Screening has no entries → doesn't block the UDM + (also = no rogue-DHCP protection: security note, not interop). + - LLDP: D-Link LLDP enabled, **all 28 ports TX+RX** (`Port_Basic_ + Setting` mode 3); stats show 25 inserts / 16 ageouts → LLDP + frames *are* exchanged & parsed with UniFi gear (no TLV + incompat). "0 neighbors" earlier was just a point-in-time + aged-out snapshot. UniFi still won't *draw* topology for an + unadopted switch — cosmetic only. + - RSTP: both run 802.1w (standard, interoperable). Tuning + opportunity (not a fault): set a deliberate root-bridge + priority — currently all default 32768 so root = lowest MAC + (arbitrary). Make the D-Link (physical core) or UDM root. + - Flow control: D-Link global FC `=2` (off/auto) — matches + UniFi's default (802.3x disabled). Compatible. + - **TP-Link RE700X** WiFi extender at `.187` (MACs `…1a:8a:0f/10/11`) — + bridges/NATs devices behind it. **The Klipper 3D printer at `.189` + sits behind the RE700X**, so UniFi only sees the repeater, never `.189`. + - **DNS:** UniFi DHCP has `dhcpd_dns_enabled: false` (clients get the + gateway `.1` as resolver, not AdGuard directly) — see DNS note below. + + **Infrastructure IPs known from UniFi:** + | IP | Name / Hostname | Notes | + |---|---|---| + | .1 | UDM Home (UDMA6A8) | Gateway, UCG Fiber, fw 5.0.16 | + | .4 | Zoraxy LXC 108 | Reverse proxy | + | .20 | Proxmox node `nuc` | PVE 9.1.11 | + | .40 | NUC 14 Pro / Docker LXC 104 | Main host, Proxmox OUI | + | .41 | Nextcloud LXC 105 | Named "NextCloud", Proxmox OUI | + | .49 | Shepard Docker host | Secondary Docker host | + | .50 | U7-Pro-Wall AP | Hallway/entry | + | .51 | U7 In-Wall AP | In-wall, UAPA6A5 | + | .52 | U7 Mesh (Schlafzimmer) | Bedroom | + | .53 | U7 Mesh (Esszimmer) | Dining room | + | .60 | Home Assistant OS VM | haos CTID 100 | + | .62 | Siemens oven | BSH Hausgeräte, WiFi | + | .66 | Tibber Pulse | Energy monitor (Espressif) | + | .10 | DGS-1210-28P | D-Link 28-port PoE switch (NOT UniFi-managed) | + | .124 | L0018 | Wired, unknown device | + | .143 | — | Sony Interactive Ent. (PlayStation) | + | .164 | C100_7614A4 | Tapo C100 camera (TP-Link) | + | .169 | awtrix_fcf0bc | AWTRIX LED clock (Espressif) | + | .187 | RE700X | TP-Link WiFi extender — NATs devices behind it | + | .189 | — | Klipper 3D printer (behind RE700X, invisible to UniFi) | + | .192 | REDMI-Note-15-Pro-5G | Xiaomi phone | + | .241 | VS9-EU-MNA3478A | Dyson purifier/fan | +- **Zoraxy** reverse proxy at `192.168.1.4:8000` — wildcard `*.nuclide.systems` + cert + rules. Source of truth: `proxy/zoraxy/routes.json` (+ idempotent + `scripts/zoraxy_sync.py --apply`). Cross-host, so no Docker labels. +- **Ubiquiti UNAS** — NFS server `192.168.1.31:/var/nfs/shared/storage` + (mounted `/mnt/pve/unas`, ~19T). Bulk/storage (arr media, qdrant, configs); + several stacks bind data dirs here. ⚠️ SQLite-on-NFS is fragile here + (see n8n note under Operational rules). +- **Auth** — **PocketID** (`id.nuclide.systems`, OIDC, SQLite) is the + **universal SSO/IdP for the entire ecosystem** — effectively every service + authenticates via PocketID OIDC (LiteLLM, LobeChat, MCP gateway, Nextcloud, + Daytona via the Keycloak→PocketID OIDC adapter, etc.). Single sign-on + everywhere; one identity source to secure/audit. + +## Docker networks +`ai-internal` (AI/MCP plane) · `shared_backend` (cross-stack DB/S3) · +plus per-stack: `arr-stack_default`, `daytona-minimal_daytona-network`, +`immich_default`, `karakeep_default`, `homepage_default`, `vpn_default`. +The MCP gateway bridges `ai-internal` + `shared_backend`. + +## AI / Agent platform (the core, `/opt/stacks/ai`) +- **LiteLLM** (`ai.nuclide.systems`, :14000) — model gateway. ~28 curated + models (provider allowlist), SAIA-terminal **failovers** + `num_retries`, + Redis completion cache, Gemini €10/30d `provider_budget_config`. Config: + `ai/litellm-config/config.yaml` (+ DB overlay `LiteLLM_Config`). +- **syncstack** (`ai/syncstack.py`, cron `/etc/cron.d/syncstack` 15 min) — + **model syncer/optimizer ONLY** (curation/allowlist → DB, `OPENAI_MODEL_LIST` + → `.env`, recreate lobe, model-health → ntfy). One-shot container. +- **mcp-gateway / Agent Operator** (`mcp.nuclide.systems`, :8080) — **all + things MCP**. DinD-spawns MCP servers on `ai-internal`, OAuth-gated streaming + reverse proxy `//mcp` + `/group//mcp`, multi-group, health-watcher, + per-server tools/list cache. **Deep health-check (2026-05-19):** beyond + shallow liveness, the watcher drives a real MCP session (`initialize` → + `tools/list` → optional `health_probe_tool` no-side-effect call) against + **every** server incl. spawned ones, recording `ok|degraded|down` in + `/api/servers`. Usage-aware: a server used within 5 min isn't probed + (traffic = health); idle servers probed every 15 min; probe self-excludes + via `X-Health-Probe`. Deep failures are status-only (no restart loops); + shallow-liveness auto-restart stays static-only. First run caught the + invisible class: `memos` degraded (search_memo → DNS `Errno -2`, the + ai-internal isolation bug), plus `context7`/`crawl4ai`/`markitdown` down + and `nextcloud` probe false-positive (probe lacks per-user NC creds — + candidate for `health_check:false`). **Semantic tool retrieval**: groups with >25 + aggregated tools expose just `find_tools`/`call_tool` meta-tools (LiteLLM + embedding index, cosine top-k, two-phase schema promotion) instead of + flooding clients with hundreds of schemas. Group members are proxied via the + gateway's own path (`_mcp_gw`) so `exact_upstream` servers (home-assistant, + n8n) aggregate correctly. **MCP→LobeChat registration** (migrated here + from syncstack), and the **Agent Operator**: cron+event agents + (`/api/agents/{name}/run` webhook), LLM↔MCP executor, **Agent Creator** + (AI-drafts agent specs; overridable meta-prompt), prompt export. Config: + `mcp-gateway/config.json` (servers), `agents.json` + `prompts/` (agents). +- **MCP servers** (gateway-managed): time, home-assistant (exact_upstream to + `.60` ha-mcp), kroki (self-hosted: kroki+mermaid+excalidraw companions), + daytona, fetch, sequential-thinking, ntfy, memos, docling, markitdown, + crawl4ai, context7, wikipedia, papersearch, comfyui, bluesky (brianellin, + rate-limit pending). Groups incl. `shepard-mcp`. +- **Daytona** (`daytona.nuclide.systems`, daytona-minimal stack) — secure + code sandboxes (Agent Operator deep-dives, shepard). Registry fixed to + internal+host `:5000` (was misconfigured at 6000). Snapshot `sciviz-py`. +- **LobeChat** (`chat.nuclide.systems`, :14001) — chat UI; MCP plugins + + agents in `lobechat` DB (lobe-postgres). Agents scriptable via + `scripts/lobe_agent.py`. Native QStash scheduler NOT functional here → + scheduling lives in the gateway Agent Operator. **Auth = Better Auth** + (not NextAuth) SSO via PocketID (`generic-oidc`, client `26f3c26b`, PKCE + S256). The `state_not_found` callback bug was fixed by setting + `AUTH_TRUSTED_ORIGINS=https://chat.nuclide.systems`; verified 2026-05-19 + (sign-in yields a correct PocketID authorize URL; callback now validates + with `oAuth_code_missing`, not `state_not_found`). Definitive proof = + one real browser login. NB: `/oidc/auth` is LobeChat-*as*-OIDC-provider + (desktop/mobile), distinct from this SSO-client path. +- **ComfyUI** (img gen, Intel Arc iGPU), **Kroki**, **searxng** (self-hosted + meta-search). + +## Data layer +- **shared-postgres** (pg16, `shared_backend`, data on local ZFS) — the + **primary/standard DB for all deployments**. Tenants: LiteLLM, paperless, + daytona, **memos** (migrated 2026-05-19). Superuser + `postgres` via unix socket; per-app dedicated role+db (role owns its db, + password in the app stack's `.env`). Tuned 2026-05-19 (3G limit, + shared_buffers 768M, max_conn 200). + ⚠️ **n8n was NOT successfully migrated** — n8n's `export/import` CLI only + covers workflows+credentials, **not users/settings/SSO**, so the PG + cutover lost the owner account + OIDC. **Reverted to local-disk SQLite** + (tier-1 still satisfied, like PocketID). Compounding: unpinned + `n8nio/n8n:latest` had drifted 2.7→**2.20.11**, breaking the custom OIDC + `hooks.js` (hardcoded old `/usr/local/lib` module paths) → n8n crash-looped. + Fixed: image pinned to `2.20.11`, `hooks.js` paths patched for 2.20.11 + pnpm layout (`/usr/lib/node_modules/n8n/...`), **OIDC hook re-enabled + 2026-05-19** — **verified**: `/auth/oidc/login` → 302 to PocketID + (client `33135ad4`, correct redirect/scope/state/nonce). **Lessons: + (1) only a FULL pgloader migration of *all* + tables is complete — partial `export/import` is not; (2) pin critical + images — `latest` drift is a real outage cause.** + Migration note: services without a native full SQLite→PG export use + **pgloader data-only of ALL tables** into the app-built schema (exclude + only the app's own migration-tracking table), app role temp-SUPERUSER for + the load (FK/trigger disable) then reverted. +- **lobe-postgres** (pg17/ParadeDB) — LobeChat. **lobe-redis** — LiteLLM cache + + LobeChat. **Garage S3** (`garage:3900`, ext `s3.nuclide.systems` / + `:10004`) — lobe-files, memos, **WAL-G PG backups**. **Qdrant** (vector, + on UNAS) — unused yet (future RAG/mem0). pgAdmin (internal). +- **Backups**: WAL-G v3 → Garage for all 3 PG instances (shared-postgres, + lobe-postgres, immich_postgres); daily 3am cron (`/etc/cron.d/pg-backup`). + **Incident + fix (2026-05-19):** Garage stored `meta`+`data` on the UNAS + NFS → NFS stalls hung `wal-g wal-push` → archivers hung since 2026-05-18 + 14:51 (`failed_count=0` = hung not failing), zero base backups, stalled + WAL recycled (that window unrecoverable). **Resolved:** Garage moved to + local disk (`/opt/stacks/shared-db/garage/{meta,data}`); all 3 archivers + drained; `archive_command` hardened to `timeout 60 wal-g wal-push %p` + (fail-fast vs infinite hang); fresh base backups taken for all 3. Extra + safety net: local logical dumps in `/opt/stacks/backups/shared-pg/`. + Lesson: Garage metadata is fsync/lock-heavy — **never** on NFS, same rule + as SQLite. + +### Database design / standards +- **Default = Postgres on `shared-postgres`.** Any app needing persistence + that supports Postgres gets a dedicated `role`+`database` there. SQLite is + allowed **only** when the app has no Postgres support, and then **only on + local disk** — never the UNAS NFS share (broken POSIX/SMB file locking; + caused the n8n outage + PocketID latent risk). +- **Provisioning pattern:** `CREATE ROLE LOGIN PASSWORD …; CREATE + DATABASE OWNER ;` → set the app's `DB_*` env, password lives in + that stack's `.env` (vault migration is the long-term plan). +- **Why standardize:** unified, durable **WAL-G PITR backup** — the single + biggest reason. One backup story instead of N un-backed-up SQLite files. +- **Migration reality:** few apps have a native SQLite→PG data port. n8n does + (`export/import` CLI). Most (PocketID, Vaultwarden, traccar, …) do not — + switching `DB_PROVIDER` starts a *fresh* DB; data port needs + `pgloader`/app-specific tooling. **Memos** done via pgloader (data-only, + app-built schema, role temp-SUPERUSER for the load). When no safe port + exists, fall back to local-disk SQLite (PocketID) until a port is built. + +### NFS / storage strategy (corrected 2026-05-19) +- The Docker host is an **unprivileged Proxmox LXC (104)** → it **cannot do + in-container NFS mounts** (kernel denies; tested — `operation not + permitted`). That's why Proxmox NFS-mounts at the host and bind-mounts + `/mnt/pve/unas` in. **Consequence:** per-stack "Docker NFS volumes at + v4.1" (the old tier-3 idea) is **infeasible** here without + `pct set 104 --features mount=nfs` + a full LXC reboot (all stacks down). +- **Tier 1 is the rule and does the real work:** DBs/metadata/lock- or + fsync-heavy stores → local ZFS or shared-postgres, **never NFS** (any + version — SQLite/LMDB/sled are unsafe on NFS regardless of tuning). +- **NFSv4 is NOT available — the Ubiquiti UNAS Pro is NFSv3-only** (current + firmware; verified 2026-05-19 on node `nuc`: `showmount -e` works but + `mount :/ -o vers=4.1` fails server-side `No such file or directory`, i.e. + no v4 pseudo-root). So tier-2 (host v3→v4.1) and tier-3 (Docker NFS v4.1 + volumes) are **both dead ends** — no v4 to upgrade to, no reboot worth + doing for it. Real export path is `/volume//.srv/.unifi-drive/ + storage/.data` (Proxmox `unas` storage uses the `/var/nfs/shared/storage` + alias, works on v3 — leave it). +- **Therefore tier-1 is the whole strategy.** NFSv3 stays for bulk/ + sequential data (fine for that). Optional marginal v3 tuning: `nconnect=4` + for throughput — not required. +- **Tier-1 complete (2026-05-19):** all SQLite-on-NFS backlog resolved — + ntfy, Karakeep, arcane → local disk; **Vaultwarden → shared-postgres** + (pgloader, 2054 rows, 0 errors, NFS kept for attachments/rsa_key.pem); + traccar data → local disk. PocketID stays on local SQLite (no native + PG migration path, deferred). + +## Services +Nextcloud (LXC .41), **Paperless-ngx** (+paperless-ai, tika/gotenberg), +**Immich** (server/ML/redis/postgres/power-tools), **Memos**, **Karakeep** +(+chrome), **Vaultwarden**, **n8n**, **Home Assistant** (.60, ~2492 entities), +**ntfy** (`homelab-ai` topic — model-health + agent alerts), **traccar** (GPS, HTTP :15000 / watch :15001), +**arr-stack** (prowlarr/shelfarr/flaresolverr behind gluetun VPN), **streamio**, +**Arcane** (Docker mgmt), **Dozzle** (logs), **Homepage** (dashboard, 6 groups), +**Daytona OIDC adapter** (Keycloak→PocketID PKCE proxy for the VS Code ext). + +## Operational rules / conventions +- **Separation**: `syncstack` = models; `mcp-gateway` = all MCP/agents. +- **Self-host first**: external/cloud only when no self-hosted form exists. +- **Pinned NTP source standard** — all network gear/infra should point at + the *same fixed-IP, DNS-free* NTP sources (devices like the D-Link can't + resolve hostnames; pinned IPs avoid silent DNS-based SNTP failure): + 1. `192.53.103.108` — PTB ptbtime1 (DE national time) + 2. `192.53.103.104` — PTB ptbtime2 (DE national time) + 3. `162.159.200.123` — Cloudflare NTP anycast (fallback) + Applied to the D-Link switch 2026-05-19. UDM + other infra should converge + on the same set (a self-hosted LAN NTP server is a roadmap option, but the + standard stays "same pinned sources" everywhere). +- Run Python with `uv run` (not python3) in `/opt/stacks`. +- Zoraxy self-service: edit `routes.json` → `scripts/zoraxy_sync.py --apply`. +- Secrets in `ai/.env` (plaintext — **env-audit/secret-vault is an open + hardening task**; Vaultwarden available). Note: a UniFi MFA JWT has leaked + into `homepage/config/logs/homepage.log` — rotate + scrub when hardening. + ⚠️ **Shared-password reuse:** `tapirnase` is reused as the WiFi PSK, the + LiteLLM master key root (`sk-tapirnase`), and the D-Link switch admin pw — + single sniff/leak has broad blast radius. Rotate per-service when hardening + (noted 2026-05-19, rotation deferred per user). +- **Proxmox/UniFi safety**: any write/change to the Proxmox host or Ubiquiti + network gear requires explicit confirmation + a stated risk assessment + first; read-only queries are fine. High blast radius (foundation layer). +- **SQLite-on-NFS is unsafe here**: the UNAS NFS share has unreliable file + locking (true for SMB/CIFS too). Any SQLite app must use `shared-postgres` + or local disk, never the NFS share. **n8n** hit this (`Database connection + timed out`) and was **migrated to `shared-postgres`** (DB on local ZFS; + data dir moved to local `/opt/stacks/n8n/data`; old NFS `database.sqlite` + kept as rollback). Audit other SQLite stacks for NFS-backed data dirs. +- Naming debt: "mcp-gateway"/"syncstack" have outgrown their names — rename + once the agent-platform identity firms up. + +## Open / roadmap +mem0 vs Qdrant for agent memory (deferred); Agent teams/orchestration + +expose Agent Operator as MCP; S3/Immich/n8n/Paperless/Proxmox MCP servers +(in progress); **UniFi MCP** — COMPLETE 2026-05-19 (`ghcr.io/enuno/unifi-mcp-server`, +197 tools, Network App API key, `http` transport; local API fully working). +**Tier-1 SQLite-off-NFS: COMPLETE** — all services off NFS for DB/metadata. +Note: Traccar watch protocol port 15001 is already forwarded at the UDM level — +no Zoraxy stream proxy needed. Plus: disable LobeChat external features (keep +`plugins`!); env → secret vault. + +**Management-plane TLS (planned):** issue/trust proper certs for admin-UI +auth on the **Proxmox** host and the **D-Link DGS-1210** (currently HTTP-only +on the D-Link → admin creds in clear on the flat LAN; Proxmox self-signed). +Brings switch/hypervisor mgmt onto the `*.nuclide.systems` PKI like the rest. + +**Network segmentation (planned):** the LAN is flat — servers, IoT (Siemens +oven, Dyson, Tapo cam, AWTRIX), consoles and phones all on one L2 (`192.168.1.0/24`, +no VLAN, no L2 isolation). Plan an **IoT VLAN** (+ matching firewall zone) so +untrusted appliances can't reach the server/Proxmox subnet. Complications to +design around: the non-UniFi **D-Link DGS-1210-28P** switch (`.10`) and +**TP-Link RE700X** extender (`.187`, NATs the Klipper printer `.189`) won't +honour UniFi VLAN tags natively — segmentation needs a plan for the wired +trunk through the D-Link and the repeater's bridge mode. + +### Observability LXC (planned — scoped from this session's incidents) +**Why it's now a priority:** the WAL-G archiver was hung **silently for +~13 h** (`failed_count=0`, zero base backups) and would never have been +noticed; NFS stalls and SQLite-on-NFS damage are likewise silent. Monitoring +must target *exactly these silent-failure classes*. +- **Placement:** its **own LXC on node `nuc`**, NOT inside LXC 104 — 104 + hosts everything, so the monitor must survive/alert when 104 is down. +- **Stack:** VictoriaMetrics (or Prometheus) + Grafana + Loki + Alertmanager + → **ntfy `homelab-ai`** (already the alert channel). +- **Exporters/probes:** node_exporter (per host + key LXCs), postgres_exporter + ×3 (shared/lobe/immich), cAdvisor/docker, blackbox (HTTP + TLS-expiry for + Zoraxy wildcard), **pve-exporter** (use the `root@pam!mcp` PVEAuditor + token), and a **custom WAL-G/archiver textfile collector**: + `pg_stat_archiver` (last_archived age, `failed_count`), `.ready` backlog, + and `wal-g backup-list` newest-base age — per PG instance. +- **Alerts (priority order, derived from real incidents):** + 1. WAL archiver stalled (last_archived age > 15 m) or newest base backup + > 26 h, any PG instance. + 2. NFS mount on `/mnt/pve/unas` unresponsive / high op latency. + 3. **Config-drift guard:** any `*.db`/`*.sqlite*` appears under + `/mnt/pve/unas` (catches a regression of the tier-1 rule). + 4. Container unhealthy/restart-looping > 5 m (the gluetun pattern). + 5. local-zfs `rpool` or NFS pool > 85 %; 6. TLS cert < 14 d. diff --git a/stacks/docs/mcp-gateway-requirements.md b/stacks/docs/mcp-gateway-requirements.md new file mode 100644 index 0000000..b298430 --- /dev/null +++ b/stacks/docs/mcp-gateway-requirements.md @@ -0,0 +1,127 @@ +# MCP Gateway — Reconstructed Design Spec (in-progress, "Phase 1") + +> Reconstructed 2026-05-16 from code/configs/git history. The gateway is a +> single-squash-commit first draft (`0cad389 "Phase 1: Create MCP Gateway with +> Docker-in-Docker support"`, preceded by `726bd10 "WIP: MCP gateway prep"`). +> Nothing has a second iteration in git — everything below is first-draft intent. + +## 1. Goal / Intent + +A single **OAuth-protected HTTP entrypoint at `https://mcp.nuclide.systems`** that +exposes a curated set of MCP servers to AI clients on the homelab. Primary +consumer: **Claude.ai** as a remote connector (SSE at `/`, every README's +"Usage in Claude.ai"). Secondary: **LobeChat** (`chat.nuclide.systems`) and +**LiteLLM** (`ai.nuclide.systems`), sharing the same Pocket ID OAuth app. It is +meant to replace the "cumbersome" static-compose approach (`mcp-tools.yaml`) +with a dynamic, UI-managed, self-hosting model — answering the open `todo.md` +question "MCP deployment seems cumbersome — can litellm host directly? how to +integrate npx, uvx, docker-based containers?". Unifying idea: normalize +npx / uvx / docker MCP servers behind one Dockerized gateway. + +## 2. Architecture (three competing models in-repo; A chosen, B orphaned, C aspirational) + +**A. FastAPI gateway + Docker-in-Docker (chosen)** — `ai/mcp-gateway/` +- FastAPI + `uvicorn` on `0.0.0.0:8080`, container `mcp-gateway`. +- DinD via bind-mounted `/var/run/docker.sock`; `docker.from_env()`. +- Per-server containers spawned `mcp-`, hardcoded onto `ai-internal`. +- Config `config.json` (RW bind, currently EMPTY → falls back to `DEFAULT_SERVERS`). +- Gateway joins `ai-internal` + `shared_backend` (both `external: true`). + +**B. Static compose `mcp-tools.yaml`** — orphaned; `ai/docker-compose.yml:6` +include is **commented out**. Internally malformed (see §4). + +**C. LiteLLM-hosted** — `litellm-config/config.yaml` `mcp_servers: {}` empty. +Confirms MCP hosting was intended for the gateway, not LiteLLM (the +`todo.md` "can litellm host directly?" question remains open). + +**Transports** (normalized to HTTP-on-:8000): streamable-http (nextcloud, +mermaid), `mcp-proxy --stateless` stdio→HTTP (papersearch), native HTTP +(markitdown, crawl4ai :11235), and the gateway's own **SSE `/` endpoint — +a STUB** (fake `initialize` + 60s pings, no routing to backends). + +**Reverse proxy:** Zoraxy `mcp.nuclide.systems → 192.168.1.40:8080`. +`mcp-auth.nuclide.systems` is an abandoned auth-sidecar idea (not exposed). + +**OAuth (Pocket ID @ `id.nuclide.systems`):** `OAuth2AuthorizationCodeBearer`, +scopes `{openid, mcp}`, token validation via userinfo. Shared gateway client +(`GENERIC_CLIENT_ID`, same as LiteLLM/LobeChat). Per-server OAuth for +papersearch & nextcloud against the same Pocket ID. + +## 3. MCP Server Inventory (reconciled — `server.py` MCP_SERVERS is authoritative) + +| Server | Image / build | Transport | Port | Auth | Status | +|---|---|---|---|---|---| +| papersearch | `python:3.12-slim` + runtime `uv tool install mcp-proxy` → `paper_search_mcp.server` | mcp-proxy stdio→http | 8000 | Pocket ID `PAPERSEARCH_MCP_OAUTH_*` | plausible, runtime-install fragile | +| nextcloud | `ghcr.io/cbcoutinho/nextcloud-mcp-server:latest` | streamable-http | 8000 | Pocket ID `NEXTCLOUD_MCP_OAUTH_*` | likely workable (real image) | +| markitdown | `python:3.12-slim` + `uvx markitdown-mcp --http` | http | 8000 | none | **broken as written** (`uvx` not in base image) | +| comfyui | `ghcr.io/richardi-ai/comfyui-mcp-server:latest` (`type:"npm"` mismatch) | unspecified | 8000 | none | **image not pullable**; backend ComfyUI was crash-looping | +| crawl4ai | `unclecode/crawl4ai:latest` | http | 11235 | none | likely workable; resource limits lost in rewrite | +| mermaid | `node:20-slim` + runtime `npx -y mcp-mermaid` | streamable-http | 8000 | none | plausible, slow first start | + +## 4. Implemented vs Unfinished vs Broken + +**Implemented:** FastAPI app + OAuth scheme + userinfo token validation; +container lifecycle CRUD + persistence; Web UI SPA (`templates/ui.html` @ `/ui`); +gateway compose/Dockerfile + Zoraxy route. + +**Unfinished / stub:** +- **SSE `/` is fake** — no MCP transport bridging Claude.ai → spawned servers. *Core gap.* +- **No routing to per-server containers**; all five servers bind the same `:8000` + and `spawn_container` host-publishes `8000:8000` → **two servers can't run at once**. +- **OAuth callback non-functional** — token-exchange URL built via + `OAUTH_REDIRECT_URI.replace("/sso/callback","/token")` (→ wrong host, not the + Pocket ID token endpoint); token never stored/used. +- `config.json` empty → always defaults; secrets hardcoded plaintext in `server.py`. + +**Broken / contradictory:** +- `ai/docker-compose.yml:6` mcp-tools include commented out; gateway compose is a + *separate project* not referenced by the stack either — wired in only via Zoraxy. +- `mcp-tools.yaml`: duplicate `markitdown-mcp` key; `comfyui-mcp` env missing `=` + (`- COMFYUI_URL http://comfyui:8188`); missing images/ports. +- `comfyui` server `type:"npm"` vs Docker-image mismatch; upstream image/npm + package existence unverified (image confirmed **not pullable**). +- `.env` has `CRAWL4AI_MCP_OAUTH_*`, `COMFYUI_MCP_OAUTH_*` that `server.py` + never consumes; code hardcodes secrets instead of `${ENV}` substitution. + +## 5. Relevant `.env` keys (names only) + +Gateway: `GENERIC_CLIENT_ID/_SECRET/_REDIRECT_URI`, +`GENERIC_{AUTHORIZATION,TOKEN,USERINFO}_ENDPOINT`, `GENERIC_CLIENT_USE_PKCE`, +`OAUTH_SCOPES`, `OAUTH_TOKEN_URL`. +Per-server: `NEXTCLOUD_MCP_OAUTH_CLIENT_ID/_SECRET`, +`PAPERSEARCH_MCP_OAUTH_CLIENT_ID/_SECRET`, +`MARKITDOWN_MCP_OAUTH_CLIENT_ID/_SECRET` (declared, unused), +`CRAWL4AI_/COMFYUI_MCP_OAUTH_*` (orphaned). +papersearch data sources: `UNPAYWALL_EMAIL`, `CORE_API_KEY`, +`SEMANTIC_SCHOLAR_API_KEY`, `ZENODO_ACCESS_TOKEN`, `GOOGLE_SCHOLAR_PROXY_URL`, +`DOAJ_API_KEY`. comfyui: `COMFYUI_URL`, `COMFYUI_WS_URL`. + +## 6. Open Design Decisions (must resolve) + +1. **Hosting model**: DinD gateway vs static `mcp-tools.yaml` vs LiteLLM-hosted. +2. **How Claude.ai/LobeChat reach a tool**: the MCP transport bridge doesn't exist. +3. **Port allocation**: all servers hardcode `:8000` — need internal DNS, no host publish. +4. **Uniform npx/uvx/docker run model**: prebuilt images vs runtime install. +5. **Auth model**: gateway-terminated vs per-MCP vs pass-through (callback is broken). +6. **Secret handling**: hardcoded → `${ENV}` from `ai/.env`. +7. **comfyui server**: image vs npm; keep only once ComfyUI itself is stable. +8. **Discovery for LobeChat/LiteLLM**: `/mcp.json` is OAuth-gated, lists config not endpoints. + +## 7. Recommended Path (ordered, lowest-risk first) + +1. **Pick the static-compose path, not DinD** — lowest risk on a single NUC; + DinD adds socket-exposure risk + a broken SSE bridge for little gain. + Fix and re-enable `mcp-tools.yaml` (uncomment `ai/docker-compose.yml:6`). +2. **Fix `mcp-tools.yaml`**: dedupe `markitdown-mcp`, fix `comfyui-mcp` env `=`, + unique service names → `ai-internal` DNS, pin images, drop comfyui for now. +3. **One streamable-http reverse proxy keyed by path** (`mcp.nuclide.systems/`) + via Zoraxy or a small `httpx` proxy — replace the fake SSE stub. Backends stay + internal on `ai-internal:8000`, never host-published. +4. **Move secrets to `${ENV}`** from `ai/.env` (keys already exist). +5. **Fix OAuth callback**: exchange code against `GENERIC_TOKEN_ENDPOINT`. +6. **Verify each upstream image/tool exists** before marking a server "working". +7. **Defer the DinD gateway + Web UI** to phase-2 (read-only status over the + running compose, not spawning). +8. **Answer the litellm question**: once stable `https://mcp.nuclide.systems/` + URLs exist, populate `litellm-config/config.yaml` `mcp_servers:` so + LiteLLM/LobeChat discover them — no bespoke discovery path needed. diff --git a/stacks/docs/mcp-gateway.md b/stacks/docs/mcp-gateway.md new file mode 100644 index 0000000..081308d --- /dev/null +++ b/stacks/docs/mcp-gateway.md @@ -0,0 +1,91 @@ +# MCP Gateway v2 — Operational + +Finished implementation of the DinD MCP gateway (Option A). Supersedes the +"Phase 1" stub described in `mcp-gateway-requirements.md`. + +## What it is + +`ai/mcp-gateway/` — FastAPI app, container `mcp-gateway`, public via Zoraxy at +`https://mcp.nuclide.systems` (→ `192.168.1.40:8080`). Single OAuth-gated +entrypoint fronting several MCP servers. + +## Architecture (as built) + +- **Spawn model**: Docker-in-Docker. Each spawnable server runs as + `mcp-` on the **`ai-internal`** network with **no host ports** + (fixes the old `8000:8000` collision — servers talk over internal DNS). +- **Transport bridge**: generic streaming reverse proxy + `https://mcp.nuclide.systems//` → `http:///`. + Transport-agnostic — works for streamable-HTTP (`/mcp`) and legacy SSE + (`/sse`+`/messages`). Registered LAST so specific routes win. +- **Auth**: Pocket ID. Two paths off the same OIDC app: + - **MCP clients / API / proxy** — `Bearer` token, validated via OIDC + userinfo (300s TTL cache). `401` carries `WWW-Authenticate` + RFC 9728 + `/.well-known/oauth-protected-resource` for discovery. + - **Browser Web UI** — full Authorization-Code flow with an httpOnly + session cookie: unauthenticated `/ui` → `302 /login` → Pocket ID + `authorize` (CSRF `state` cookie) → `/sso/callback` (state-checked + code→token exchange) sets the `mcp_session` cookie and shows the page + + a copyable Bearer token for MCP clients. `/logout` clears it. The UI and + `/api/*` accept **either** the cookie or a Bearer header; the `/{srv}/` + proxy stays Bearer-only. `/health` + the well-known metadata are open. +- **Secrets**: all per-server OAuth client secrets come from `ai/.env` + (`env_file: ../.env`), read via `os.environ` — none hardcoded in source. + +## Server registry (`SERVERS` in server.py) + +| name | kind | upstream | notes | +|---|---|---|---| +| comfyui | static | `http://comfyui-mcp:8000` | runs as its own stack (`ai/comfyui-mcp.yml`); gateway only proxies | +| nextcloud | spawn | `mcp-nextcloud:8000` | `ghcr.io/cbcoutinho/nextcloud-mcp-server` streamable-http | +| crawl4ai | spawn | `mcp-crawl4ai:11235` | `unclecode/crawl4ai` | +| mermaid | spawn | `mcp-mermaid:8000` | `node:20-slim` + npx mcp-mermaid (validated working) | +| markitdown | spawn | `mcp-markitdown:8000` | `ghcr.io/astral-sh/uv` + uvx markitdown-mcp | +| papersearch | spawn | `mcp-papersearch:8000` | python + uv + mcp-proxy → paper_search_mcp | + +## Verified + +- `/health` 200 (no auth); `/.well-known/oauth-protected-resource` JSON (no auth). +- `/mcp.json`, `//mcp` → 401 without token; 401 with bad token + (userinfo path); correct `WWW-Authenticate`. +- Upstream path proven: gateway → `comfyui-mcp` `/mcp` initialize → 200; + gateway → spawned `mcp-mermaid` `/mcp` → 200 (spawn pattern + proxy). +- Route precedence correct (specific routes beat the catch-all proxy). + +## Not yet exercised (needs a real token / operational bring-up) + +- **Token-authenticated end-to-end** requires an interactive Pocket ID login. + Get a token: open `https://mcp.nuclide.systems/sso/callback` via the OAuth + authorize flow (or any Pocket ID token with scopes `openid,mcp`); the + callback page prints the access token. +- **Bring up the 4 remaining spawnable servers** (nextcloud/crawl4ai/ + markitdown/papersearch): `POST /api/servers//start` (bearer) or + `POST /api/servers/bulk/start`. Each may need a transport/flag tweak on + first run — mermaid is the proven reference. + +## Client registration + +- **LobeChat**: Settings → Skills → Skill Store → Custom → Import JSON, e.g. + ```json + { "mcpServers": { "comfyui": { "type": "http", + "url": "https://mcp.nuclide.systems/comfyui/mcp", + "auth": { "type": "bearer", "accessToken": "" } } } } + ``` + (Or, for ComfyUI specifically, connect LobeChat straight to + `http://comfyui-mcp:8000/mcp` internally — see `comfyui-mcp.md` — no token.) +- **Claude.ai**: add `https://mcp.nuclide.systems//mcp` as a custom + connector; it will discover auth via the protected-resource metadata. + +## Ops + +- `cd /opt/stacks/ai/mcp-gateway && docker compose up -d --build` +- Logs: `docker logs mcp-gateway`; UI: `https://mcp.nuclide.systems/ui` + (Pocket-ID-gated; `/login` to sign in). +- Config overlay persisted in `ai/mcp-gateway/config.json`. +- **UI check (Playwright)** — the live UI is auth-gated, so the check + serves the local template + mocks `/api/servers`, screenshots, and + asserts cards render: + ``` + uv run --with playwright -- python -m playwright install chromium # once + uv run --with playwright -- python ai/mcp-gateway/ui_check.py # -> /tmp/mcp_ui.png + ``` diff --git a/stacks/docs/scrubbing-list.md b/stacks/docs/scrubbing-list.md new file mode 100644 index 0000000..1e4b5c1 --- /dev/null +++ b/stacks/docs/scrubbing-list.md @@ -0,0 +1,53 @@ +# Scrubbing List — /opt/stacks (2026-05-17) + +Read-only audit of unused/stale data. **Nothing here has been deleted** — review +the labels and run the commands yourself. Root FS was **165G/200G used (83%)**; +`/var/lib/docker` is 86G of `/opt/stacks`'s 99G. + +No stopped/exited/`*_old` containers; no unused custom networks (already clean). + +## 1. Docker reclaimable + +| Target | Size | Label | Command | +|---|---|---|---| +| Build cache (268 entries, old comfyui CPU→Arc rebuilds, 0 in use) | **~20.85 GB** | **SAFE** | `docker builder prune -af` | +| 4× dangling `` mcp-gateway rebuild images (1.59 GB ea) | **~6.36 GB** | **SAFE** | `docker image prune` | +| 2 old dangling images (756 MB + 113 MB) | **~0.87 GB** | **SAFE** | (same `docker image prune`) | +| ~265 anon volumes; one `54bcf7…` = **8.69 GB** unidentified, rest ~0B | ~9 GB | **REVIEW** | inspect `54bcf7…` then `docker volume prune` | +| Named dangling vols: `n8n_n8n_storage` 128M, `ai_mcpo-data` 84M, `paperless-ngx_pgdata` 26M, `metamcp_postgres_data` 18M, `daytona*_db_data` 15M×2, `librechat_pgdata2` 13M, `arcane_arcane-data` 12M, `ai_redis_data` 7.6M | ~0.3 GB | **REVIEW** | `docker volume rm ` per-item after confirming the stack is retired | + +In-use, **DO NOT REMOVE**: `comfyui-comfyui` (6.45G), `clusterzx/paperless-ai` (8.59G). + +## 2. Migrated/abandoned local data dirs (compose now points to UNAS) + +| Path | Size | Label | Note | +|---|---|---|---| +| `arr-stack/media` | **55 G** | **REVIEW** | No container mounts it; arr → `/mnt/pve/unas/media`. Audiobooks/ebooks have recent mtimes (rsync residue) — **parity-check vs UNAS before `rm -rf`** | +| `ai/data` (old postgres) | 195 M | **SAFE** | Not mounted, not referenced | +| `ai/postgres_data` (incl 38M pg_wal) | 120 M | **REVIEW** | Not mounted/referenced but recent mtime | +| `ai/meili_data_v1.35.1` | 19 M | **SAFE** | Old Meili, not mounted | +| `qdrant/qdrant_storage` | 7 M | **SAFE** | qdrant migrated to UNAS (fresh start) | +| `daytona/db_data` | 14 M | **REVIEW** | No mount; daytona uses named volumes | +| `n8n/data` | empty | **SAFE** | `rmdir` | + +Active local, **KEEP**: `immich/postgres` (814M), `ai/lobehub/data` (25M), +`shared-db/wal-g` (27M, RO mount), arr-stack configs, `ai/litellm-config`, +`ai/searxng`. + +## 3. Migration / scratch artifacts + +| Path | Label | Note | +|---|---|---| +| `ai/docker` (0 B), `ai/bucket.config.json` (empty dir) | **SAFE** | junk | +| `scripts/traefik-*.sh` | **SAFE** | Traefik abandoned for Zoraxy | +| `scripts/{migrate_*,test_adguard_api,zoraxy_csrf,zoraxy_test,configure_zoraxy_*}.py` | **REVIEW** | one-shot done; confirm no rerun need | +| `scripts/zoraxy_sync.py` | **KEEP** | ongoing proxy tooling | +| `scripts/.venv` (29M), `scripts/.kilo` (30M) | **REVIEW** | regenerable caches | +| `/tmp/{flux_*,pw_ui,add_*,fix_*}.* , /tmp/*.png , /tmp/*.log` | **SAFE** | ~1.8M scratch (this session) | + +## Bottom line + +- **Safe-only reclaim: ~29–30 GB** (build cache + dangling images dominate). +- **With review: +~64 GB** (55G stale `arr-stack/media` after UNAS parity check, ~9G anon vols). +- **Highest-value, zero-risk action:** `docker builder prune -af` → ~20.85 GB now. +- Biggest overall prize: verify+remove `arr-stack/media` (55 G) — parity-check vs `/mnt/pve/unas/media` first. diff --git a/stacks/docs/stack-ideas.md b/stacks/docs/stack-ideas.md new file mode 100644 index 0000000..f9b93f6 --- /dev/null +++ b/stacks/docs/stack-ideas.md @@ -0,0 +1,248 @@ +# Stack Ideas & Improvements + +Hardware baseline: **NUC 14 Pro** — Intel Core Ultra (Meteor Lake), Intel Arc iGPU +(currently used for ComfyUI / XPU image gen), no NVIDIA, UNAS for media/bulk storage. + +--- + +## 1. Document ingestion → LobeChat knowledge base + MCP + +**Goal**: ingest PDFs, Word, PPT, XLS → searchable in LobeChat chat + accessible via MCP. +Qdrant is a nice-to-have, not a requirement (no live pipeline uses it yet). + +**LobeChat already has a built-in knowledge base** (`knowledge_base_files`, `chunks`, +`embeddings` tables in Postgres). The missing piece is an ingest pipeline that feeds it. + +**Recommended architecture**: +``` +Nextcloud folder / Paperless webhook + → n8n trigger (already running) + → Docling (PDF/DOCX/PPTX/XLSX → Markdown + structure) + → LiteLLM /v1/embeddings (Mistral-embed / codestral-embed) + → LobeChat knowledge base API ←— available in chat natively + → Qdrant sidecar (optional, if multi-app search needed) +``` + +**Document conversion options**: + +| Tool | Docker image | Formats | NUC 14 CPU fit | +|------|-------------|---------|----------------| +| **Docling** (IBM, 2024) | `ghcr.io/ds4sd/docling` | PDF, DOCX, PPTX, XLSX, HTML, Markdown | ✅ CPU-only, 10-60s/doc | +| **MinerU** (OpenDataLab) | `opendatalab/mineru` | PDF (layout-aware, OCR) | ✅ CPU mode; GPU optional for speed | +| **Unstructured** | `quay.io/unstructured-io/unstructured-api` | Very broad (25+ formats) | ✅ lighter than Docling | +| **Markitdown** (already in MCP gateway) | — | Office + PDFs → Markdown | ✅ ad-hoc only, not batch | + +**Docling vs MinerU**: Docling is better for structured Office/PDF with tables and +figures. MinerU (OpenDataLab) is better for pure PDF optical layout analysis (e.g., +academic papers, scanned documents). Both run on NUC 14 CPU. Start with Docling — single +container, REST API, well-documented. + +**MCP access**: add a `knowledge-search` MCP tool to the gateway that calls LobeChat's +knowledge base search API. Zero new infra — LobeChat is already on `shared_backend`. + +**IBM Granite embedding** (granite-embedding-30m-english) — good quality, tiny (30M params), +runs CPU-only. Alternative to Mistral-embed if you want fully local embeddings. Not on +LiteLLM yet but can add as a custom provider pointing at an Ollama/IPEX-LLM instance. + +--- + +## 2. Document conversion pipeline + +**Problem**: No systematic ingestion path from raw docs (PDF, DOCX, HTML) into +structured text for chunking/embedding. + +**Tools to consider**: + +| Tool | Docker image | Best for | +|------|-------------|----------| +| **Docling** (IBM, 2024) | `ghcr.io/ds4sd/docling` | PDFs with complex layout (tables, columns, figures); outputs Markdown | +| **Unstructured** | `quay.io/unstructured-io/unstructured-api` | Wide format support (HTML, DOCX, PPTX, images with OCR); REST API | +| **Gotenberg** | `gotenberg/gotenberg` | HTML/Office → PDF pre-processing stage; not a text extractor | +| **Apache Tika** | `apache/tika` | Broad format support; outputs plain text; lower quality than Docling for PDFs | +| **Markitdown MCP** | already in gateway | Ad-hoc on-demand conversion; not suitable for batch pipelines | + +**Recommended pipeline (n8n-orchestrated)**: +``` +Nextcloud / Paperless webhook + → Gotenberg (Office → PDF) + → Docling (PDF → Markdown chunks) + → LiteLLM /v1/embeddings (Mistral-embed) + → Qdrant upsert +``` + +Docling runs CPU-only comfortably on the NUC. Unstructured is heavier but has a +managed API if the self-hosted version is too slow. + +Paperless-ngx already OCRs documents — its full-text content is accessible via the +REST API. An n8n workflow polling `/api/documents/?added__gt=` can extract +and re-embed directly, without re-running OCR. + +--- + +## 3. On-device LLM (Arc iGPU) + +The Arc iGPU is currently ComfyUI-only. Small language models can also run on it: + +- **IPEX-LLM** (Intel) + **Ollama** — Intel provides a patched Ollama build that uses + IPEX-LLM for Arc acceleration. Supports Phi-3.5-mini, Gemma-2B, TinyLlama. + Useful as a fast/cheap fallback when Groq/Mistral rate limits hit. +- **llama.cpp with Vulkan** — alternative to IPEX-LLM; no Intel-specific driver, + uses Vulkan compute. Less optimised for Arc but simpler to deploy. +- **Practical constraint**: Arc iGPU shares system RAM; sustained LLM inference + competes with other XPU workloads (e.g., ComfyUI). A model serving config that + pauses one workload while the other runs is needed, or deploy them on separate + CPU/GPU budgets. + +If a second NUC or small GPU box becomes available, this becomes the primary use case. + +--- + +## 4. Arcane agents on NUC (central instance on Proxmox LXC) + +Arcane's central server moves to a **Proxmox LXC** (lightweight, stays responsive even +if the Docker stack has issues). The NUC and any other Docker host runs the **Arcane +agent** (headless worker), which connects back to the central instance. + +**Deployment**: +- LXC: central Arcane server + observability stack (see §8) +- NUC (`.40`) and `.49`: `arcane-agent` (headless) in compose, connects to LXC +- Arcane agents reach Docker via either: + - **TCP Docker socket + TLS certs** — most secure; requires cert generation per host + - **SSH Docker contexts** — simpler; gateway mounts socket into agent via SSH tunnel + +**Observability co-located on the LXC** (see §8 — lightweight, OTEL-aware stack that +survives Docker restarts). + +--- + +## 5. AI memory / personalization + +**Mem0** (`mem0ai/mem0`) — persistent user memory layer that sits in front of LLM calls. +Stores facts extracted from conversations into a vector DB (Qdrant backend supported). +Can integrate with LobeChat or as an MCP tool. Lets the AI remember preferences, +past context, and user-specific facts across sessions. + +Alternative: **Letta** (formerly MemGPT) — stateful agent framework with persistent +memory; more opinionated. + +--- + +## 6. Workflow / automation upgrades + +- **n8n → AI agent nodes**: n8n 1.x has native AI agent nodes (LangChain under the hood). + Can build RAG, email triage, document processing workflows visually. +- **Activepieces** — lighter-weight n8n alternative; good for simple integrations. + Probably not worth adding since n8n is already running. +- **Temporal** — durable workflow engine for long-running or retry-heavy pipelines + (e.g., large document batch processing). Overkill unless pipelines become complex. + +--- + +## 7. External services worth considering + +| Service | Purpose | Notes | +|---------|---------|-------| +| **Backblaze B2** | Off-site backup (already in todos) | rclone sync from Garage; restic for PG dumps | +| **Cloudflare R2** | S3-compatible CDN-backed storage | Free egress; good for Lobe file serving | +| **Resend** | Transactional email | 3K/mo free; better deliverability than self-hosted Postal | +| **ntfy.sh (cloud)** | Push notifications fallback | Already running self-hosted ntfy; cloud as relay | +| **Cloudflare Turnstile** | Bot protection for public endpoints | Free; no JS challenge | +| **Novu** | Notification orchestration | Multi-channel (email, push, Slack); self-hostable | + +--- + +## 8. Observability (on Proxmox LXC, alongside Arcane) + +**Runs on the LXC, not the NUC** — stays alive if the Docker stack misbehaves. Lightweight +enough for a 2 vCPU / 4GB LXC. + +**Recommended stack** (all OTEL-aware, compose-based): + +| Component | Image | Role | +|-----------|-------|------| +| **OpenTelemetry Collector** | `otel/opentelemetry-collector-contrib` | Receives traces/metrics/logs from all services (OTLP gRPC+HTTP); fans out to backends | +| **VictoriaMetrics** | `victoriametrics/victoria-metrics` | Prometheus-compatible TSDB; scrapes NUC exporters + receives from OTEL collector; lighter than Prometheus | +| **Grafana** | `grafana/grafana` | Dashboards; datasource = VictoriaMetrics + Loki | +| **Loki** | `grafana/loki` | Log aggregation; receives from OTEL collector | +| **Uptime Kuma** | `louislam/uptime-kuma` | HTTP/TCP uptime checks for all public endpoints; alerts via ntfy | + +**OTEL receivers** from the NUC Docker stack: +- **LiteLLM**: native OTEL export — traces every LLM call with token counts, model, latency +- **n8n**: Prometheus `/metrics` endpoint +- **Garage**: Prometheus `/metrics` +- **mcp-gateway**: add `opentelemetry-sdk` instrumentation to `server.py` +- **Docker host**: `node_exporter` + `cadvisor` on the NUC, scraped by VictoriaMetrics + +**NUC → LXC connectivity**: Both are on the same LAN. OTEL collector listens on the LXC's +LAN IP (e.g., `192.168.1.X:4317` gRPC). Services push OTEL directly to it; Prometheus +pull-scraping from VictoriaMetrics goes to NUC exporters over LAN. + +--- + +## 9. Storage / S3 improvements + +- **Garage external S3** is currently non-responsive (tracked in todos). Fix or + replace with **MinIO** (single-node, much better tooling/UI, easy migration). +- **Lobe file serving via Cloudflare R2**: Lobe uploads to Garage → sync to R2 → + serve from CDN. Reduces NUC egress for image-heavy sessions. +- **Qdrant data on UNAS**: Mount `qdrant_data` from NAS for persistence across + container re-creates (same pattern as arr-stack). + +--- + +## 9. ComfyUI MCP — async queue + image-to-image + reference + +**Current problem**: MCP tool blocks until the image is done (~90-290s). LLMs time out at ~60s. +The fix is a job-queue pattern: + +``` +generate_image(prompt) → returns {job_id, status: "queued"} immediately +get_image_status(job_id) → returns {status, progress, image_url_when_done} +list_queue() → shows all pending/running jobs +cancel_job(job_id) → cancels a queued job (confirms with user first) +``` + +ComfyUI's own API is already async (`POST /prompt` → poll `/history/{id}`). The MCP server +just needs to expose this model instead of blocking. + +**Image-to-image**: new workflow `flux-schnell-img2img-api.json`. Takes an init image URL + +denoise strength. ComfyUI `LoadImageFromURL` node (or upload + `LoadImage`). + +**Reference tool**: `list_previous_images(n=5)` — queries ComfyUI `/history` API, +returns recent job thumbnails + prompts. User can pick one to reference or iterate from. + +**Delivery to S3/Garage**: on completion, upload output PNG to Garage `comfyui-outputs` bucket +→ return a permanent URL. Avoids ComfyUI's ephemeral `/view` endpoint. + +--- + +## 10. LiteLLM custom provider / LobeChat provider extension + +**Question**: build an adapter in LiteLLM to use models via "quasi API" (non-standard +endpoints, auth, or routing)? + +**LiteLLM supports custom providers** via `custom_llm_provider` in config.yaml. You write +a Python class that implements `completion()` and `async_completion()`. This is the right +path for wrapping non-standard APIs (local models, proprietary endpoints, protocol bridges). + +Example use cases: +- Wrap a Claude Max subscription via Anthropic's API (different billing model) +- Add a local model served by IPEX-LLM on the Arc iGPU +- Bridge a custom inference server that speaks a different protocol + +**LobeChat provider extension**: LobeChat's provider list is compiled into the app. +Adding a new provider requires rebuilding LobeChat from source (fork + add to +`src/config/aiModels/`). High effort; only worth it for a permanent/long-term provider. +For ad-hoc needs, use LiteLLM as the adapter and point LobeChat at it via the existing +`ai.nuclide.systems` OpenAI-compatible endpoint. + +--- + +## Priority order (rough) + +1. **ComfyUI MCP async queue** — fixes timeout; unblocks img2img + reference features +2. **Docling container + n8n ingest pipeline** — Nextcloud/Paperless → LobeChat knowledge base +3. **Arcane LXC + agents on NUC** — when ready to migrate +4. **Observability LXC** — OTEL collector + VictoriaMetrics + Grafana + Uptime Kuma, co-located +5. **Backblaze B2 off-site backup** — already in todos +6. **On-device LLM (Arc)** — depends on IPEX-LLM + Ollama Intel build stability diff --git a/stacks/docs/traefik-migration-docker-labels.md b/stacks/docs/traefik-migration-docker-labels.md new file mode 100644 index 0000000..e04b31e --- /dev/null +++ b/stacks/docs/traefik-migration-docker-labels.md @@ -0,0 +1,341 @@ +# Traefik Migration Guide Using Docker Labels + +## Overview +Migrate from Zoraxy reverse proxy to Traefik using Docker labels for zero-touch service discovery. + +--- + +## Phase 1: Install Traefik + +### Step 1: Create Directory Structure +```bash +mkdir -p /opt/stacks/proxy/traefik/{config,dynamic,letsencrypt} +``` + +### Step 2: Create docker-compose.yml +```yaml +version: "3.8" + +services: + traefik: + image: traefik:v3.2 + container_name: traefik + restart: always + network_mode: host + security_opt: + - no-new-privileges=true + ports: + - "80:80" + - "443:443" + volumes: + - /var/run/docker.sock:/var/run/docker.sock:ro + - /opt/stacks/proxy/traefik/config:/etc/traefik + - /opt/stacks/proxy/traefik/dynamic:/etc/traefik/dynamic + - /opt/stacks/proxy/traefik/letsencrypt:/etc/letsencrypt + command: + - "--api.insecure=true" + - "--providers.docker=true" + - "--providers.docker.exposedbydefault=false" + - "--providers.docker.network=ai-internal" + - "--providers.docker.network=shared_backend" + - "--providers.docker.defaultRule=Host(`{{ .Name }}.nuclide.systems`)" + - "--entrypoints.web.address=:80" + - "--entrypoints.websecure.address=:443" + - "--certificatesresolvers.letsencrypt.acme.httpChallenge=true" + - "--certificatesresolvers.letsencrypt.acme.email=admin@nuclide.systems" + - "--certificatesresolvers.letsencrypt.acme.storage=/etc/letsencrypt/acme.json" +``` + +### Step 3: Start Traefik +```bash +cd /opt/stacks/proxy/traefik +docker compose up -d + +# Verify +docker compose ps +``` + +--- + +## Phase 2: Migrate AI Services + +### AI Service Labels (Add to litellm, chat, mcp-compose.yml) + +```yaml +services: + litellm: + image: litellm + labels: + - "traefik.enable=true" + - "traefik.http.routers.litellm.rule=Host(`litellm.nuclide.systems`)" + - "traefik.http.routers.litellm.entrypoints=websecure" + - "traefik.http.routers.litellm.tls=true" + - "traefik.http.routers.litellm.tls.certresolver=letsencrypt" + - "traefik.http.routers.litellm.priority=10" + - "traefik.http.services.litellm.loadbalancer.server.port=14000" + + chat: + image: lobehub + labels: + - "traefik.enable=true" + - "traefik.http.routers.chat.rule=Host(`chat.nuclide.systems`)" + - "traefik.http.routers.chat.entrypoints=websecure" + - "traefik.http.routers.chat.tls=true" + - "traefik.http.routers.chat.tls.certresolver=letsencrypt" + - "traefik.http.routers.chat.priority=10" + - "traefik.http.services.chat.loadbalancer.server.port=14001" + + mcp: + image: mcp-gateway + labels: + - "traefik.enable=true" + - "traefik.http.routers.mcp.rule=Host(`mcp.nuclide.systems`)" + - "traefik.http.routers.mcp.entrypoints=websecure" + - "traefik.http.routers.mcp.tls=true" + - "traefik.http.routers.mcp.tls.certresolver=letsencrypt" + - "traefik.http.routers.mcp.priority=10" + - "traefik.http.services.mcp.loadbalancer.server.port=8080" +``` + +--- + +## Phase 3: Migrate Garage S3 + +### Option A: Use Traefik Proxy +```yaml +services: + garage-proxy: + image: nginx:alpine + labels: + - "traefik.enable=true" + - "traefik.http.routers.s3.rule=Host(`s3.nuclide.systems`)" + - "traefik.http.routers.s3.entrypoints=websecure" + - "traefik.http.routers.s3.tls=true" + - "traefik.http.routers.s3.tls.certresolver=letsencrypt" + - "traefik.http.services.s3.loadbalancer.server.port=10004" + volumes: + - garage-data:/data + networks: + - shared_backend +``` + +### Option B: Keep Internal Garage Access +```yaml +services: + # No proxy needed - access Garage via internal IP:10004 + garage: + image: garageio/garage + ports: + - "3900:3900" # Internal only + - "10004:10004" # Public via Traefik +``` + +--- + +## Phase 4: Create Helper Scripts + +### script 1: Add Service to Traefik +```bash +#!/bin/bash +# /opt/stacks/scripts/add-traefik-service.sh + +NAME=$1 +DOMAIN=$2 +PORT=$3 + +cat > /opt/stacks/proxy/traefik/dynamic/${NAME}.yml << EOF +http: + routers: + ${NAME}-router: + rule: "Host(\`${DOMAIN}\`)" + service: ${NAME}-service + entrypoints: + - websecure + tls: + certresolver: letsencrypt + + services: + ${NAME}-service: + loadBalancer: + servers: + - url: "http://192.168.1.40:${PORT}" +EOF + +# Reload Traefik docker automatically (no manual step needed) +echo "✅ Service ${NAME} added via labels" +``` + +Usage: +```bash +/opt/stacks/scripts/add-traefik-service.sh myservice myservice.nuclide.systems 8000 +``` + +### Script 2: Generate Labels for Existing Services +```bash +#!/bin/bash +# /opt/stacks/scripts/traefik-labels-gen.sh + +cat > /opt/stacks/proxy/traefik/labels.yaml << 'EOF' +# Add these labels to service docker-compose.yml files + +# LiteLLM +services: + litellm: + labels: + - "traefik.enable=true" + - "traefik.http.routers.litellm.rule=Host(`litellm.nuclide.systems`)" + - "traefik.http.routers.litellm.entrypoints=websecure" + - "traefik.http.routers.litellm.tls=true" + - "traefik.http.routers.litellm.tls.certresolver=letsencrypt" + - "traefik.http.services.litellm.loadbalancer.server.port=14000" + +# LobeHub Chat +services: + chat: + labels: + - "traefik.enable=true" + - "traefik.http.routers.chat.rule=Host(`chat.nuclide.systems`)" + - "traefik.http.routers.chat.entrypoints=websecure" + - "traefik.http.routers.chat.tls=true" + - "traefik.http.routers.chat.tls.certresolver=letsencrypt" + - "traefik.http.services.chat.loadbalancer.server.port=14001" + +# MCP Gateway +services: + mcp-gateway: + labels: + - "traefik.enable=true" + - "traefik.http.routers.mcp.rule=Host(`mcp.nuclide.systems`)" + - "traefik.http.routers.mcp.entrypoints=websecure" + - "traefik.http.routers.mcp.tls=true" + - "traefik.http.routers.mcp.tls.certresolver=letsencrypt" + - "traefik.http.services.mcp.loadbalancer.server.port=8080" +EOF + +echo "✅ Labels saved to /opt/stacks/proxy/traefik/labels.yaml" +``` + +Usage: +```bash +/opt/stacks/scripts/traefik-labels-gen.sh +``` + +--- + +## Phase 5: Update Service Configs + +### Update `/opt/stacks/ai/.env` +```bash +# OLD (Zoraxy): +LITELLM_BASE_URL=https://litellm.nuclide.systems +PROXY_BASE_URL=https://mcp.nuclide.systems + +# NEW (Traefik) - same URLs, different backend: +LITELLM_BASE_URL=https://litellm.nuclide.systems +PROXY_BASE_URL=https://mcp.nuclide.systems +CHATAI_BASE_URL=https://chat.nuclide.systems +``` + +### Update `/opt/stacks/ai/litellm-config/config.yaml` +```yaml +general_settings: + proxy_base_url: https://litellm.nuclide.systems + control_plane_url: https://litellm.nuclide.systems +``` + +--- + +## Phase 6: Verify SSL + +### Step 1: Generate Let's Encrypt Certificates +```bash +# Verify Traefik is running +docker compose ps traefik + +# Create ACME cert file +touch /opt/stacks/proxy/traefik/letsencrypt/acme.json +chmod 600 /opt/stacks/proxy/traefik/letsencrypt/acme.json + +# Trigger certificate generation (will happen automatically) +# Check status: +curl -s https://acme-v02.api.letsencrypt.org/directory | head +``` + +### Step 2: Test HTTPS +```bash +# Test endpoints +curl -k https://litellm.nuclide.systems/health +curl -k https://chat.nuclide.systems/health +curl -k https://mcp.nuclide.systems/health + +# Verify cert +curl -v https://litellm.nuclide.systems 2>&1 | grep -A 5 "SSL certificate" +``` + +--- + +## Phase 7: Remove Zoraxy + +### Backup first +```bash +# Backup Zoraxy configs +docker cp zoraxy:/data/configs /backup/zoraxy-backup/ + +# Optional: Stop Zoraxy +docker stop zoraxy +docker rm zoraxy +``` + +--- + +## Quick Migration Checklist + +- [ ] Install Traefik (Phase 1) +- [ ] Add labels to AI services (Phase 2) +- [ ] Add Garage S3 proxy (Phase 3) +- [ ] Run label generation script (Phase 4) +- [ ] Update `.env` and config files (Phase 5) +- [ ] Wait for SSL certificates (auto 24-48h) +- [ ] Test all HTTPS endpoints +- [ ] Remove Zoraxy (Phase 7) +- [ ] Update AdGuard DNS if needed (optional) + +--- + +## Monitoring & Troubleshooting + +### Check Traefik Dashboard +```bash +# Access web UI (unsecured - use only on trusted network) +open http://192.168.1.40:8080/dashboard + +# View all routers +curl http://localhost:8080/api/http/routers | jq '.[] | {name: .rule, status: .entryPoints}' + +# View all services +curl http://localhost:8080/api/http/services | jq '.[] | {name: .name, servers: .servers}' +``` + +### Common Issues + +| Issue | Solution | +|-------|----------| +| 404 errors | Check router labels match domain exactly | +| SSL expired | Wait for auto-renew or trigger manually | +| Port mismatch | Verify `loadbalancer.server.port` matches service | +| No SSL cert | Check email in `acme.json` config | + +--- + +## Rollback Plan (If Needed) + +```bash +# Stop Traefik +docker compose -f /opt/stacks/proxy/traefik/docker-compose.yml down + +# Restore Zoraxy configs +docker cp /backup/zoraxy-backup/ configs/ + +# Restart Zoraxy (if you kept backup) +docker start zoraxy || true +``` diff --git a/stacks/docs/traefik-migration.md b/stacks/docs/traefik-migration.md new file mode 100644 index 0000000..b33b878 --- /dev/null +++ b/stacks/docs/traefik-migration.md @@ -0,0 +1,193 @@ +# Recommended: Traefik Proxy Replacement + +## Why Traefik over Zoraxy? + +| Feature | Zoraxy | Traefik | +|---------|--------|---------| +| **API** | ❌ No public API | ✅ Full REST API | +| **SSL** | 💬 Manual (Zoraxy Web UI) | 🔥 Auto-Let's Encrypt | +| **Dynamic** | ⚠️ Manual config reload | ✅ Hot-reload configs | +| **File watching** | ❌ | ✅ Auto-detect changes | +| **Docker integration** | ⚠️ Manual | ✅ Native labels | +| **API endpoints** | 403 Forbidden | ✅ JSON API everywhere | + +## Migration Path + +### Current setup: +``` +Zoraxy (192.168.1.4:8000) → Reverse Proxy Rules (Manual Web UI) +- litellm.nuclide.systems → 192.168.1.40:14000 +- chat.nuclide.systems → 192.168.1.40:14001 +- mcp.nuclide.systems → 192.168.1.40:8080 +- s3.nuclide.systems → Garage:10004 +``` + +### New setup with Traefik: +``` +Traefik (public SSL) → Dynamic Router (labels/consul) +- All services auto-discovered via Docker labels +- SSL certificates auto-provisioned +- No manual Zoraxy configuration needed +``` + +## Installation + +### Step 1: Install Traefik + +```bash +# Create Traefik directory +mkdir -p /opt/stacks/proxy/traefik/{conf,dynamic} + +# Create docker-compose.yml +cat > /opt/stacks/proxy/traefik/docker-compose.yml << 'EOF' +version: "3.8" + +services: + traefik: + image: traefik:v3.2 + container_name: traefik + restart: always + security_opt: + - no-new-privileges=true + network_mode: host + ports: + - "80:80" + - "443:443" + volumes: + - /var/run/docker.sock:/var/run/docker.sock:ro + - /opt/stacks/proxy/traefik/conf:/etc/traefik + - /opt/stacks/proxy/traefik/dynamic:/etc/traefik/dynamic + - /opt/stacks/proxy/traefik/letsencrypt:/etc/letsencrypt + command: + - "--api.insecure=true" + - "--providers.docker=true" + - "--providers.docker.exposedbydefault=false" + - "--entrypoints.web.address=:80" + - "--entrypoints.websecure.address=:443" + - "--certificatesresletsencryptemail=admin@nuclide.systems" + - "--certificatesresletsencryptstorage=/etc/letsencrypt/acme.json" +EOF + +# Start Traefik +cd /opt/stacks/proxy/traefik && docker compose up -d +``` + +### Step 2: Create Dynamic Configuration + +```bash +# Create router rules +cat > /opt/stacks/proxy/traefik/dynamic/router.yml << 'EOF' +http: + routers: + litellm-router: + rule: "Host(`litellm.nuclide.systems`)" + service: litellm-service + entrypoints: + - websecure + tls: + certresolver: letsencrypt + + chat-router: + rule: "Host(`chat.nuclide.systems`)" + service: chat-service + entrypoints: + - websecure + tls: + certresolver: letsencrypt + + mcp-router: + rule: "Host(`mcp.nuclide.systems`)" + service: mcp-service + entrypoints: + - websecure + tls: + certresolver: letsencrypt + + s3-router: + rule: "Host(`s3.nuclide.systems`)" + service: garage-service + entrypoints: + - websecure + tls: + certresolver: letsencrypt + + services: + litellm-service: + loadBalancer: + servers: + - url: "http://192.168.1.40:14000" + + chat-service: + loadBalancer: + servers: + - url: "http://192.168.1.40:14001" + + mcp-service: + loadBalancer: + servers: + - url: "http://192.168.1.40:8080" + + garage-service: + loadBalancer: + servers: + - url: "http://garage:10004" +EOF +``` + +### Step 3: Add Docker Labels to Services + +For any Docker service you want to proxy: + +```yaml +# Example: Add to your service docker-compose.yml +services: + ai-service: + image: your-service + labels: + - "traefik.enable=true" + - "traefik.http.routers.your-service.rule=Host(`your-service.nuclide.systems`)" + - "traefik.http.routers.your-service.entrypoints=websecure" + - "traefik.http.routers.your-service.tls.certresolver=letsencrypt" + - "traefik.http.services.your-service.loadBalancer.server.port=8000" +``` + +### Step 4: Delete Zoraxy (Optional) + +```bash +# Backup Zoraxy configs first +tar -czf /backup/zoraxy-backup.tar.gz /path/to/zoraxy/configs + +# Stop and remove Zoraxy +docker rm -f zoraxy || true +``` + +## API Example (Traefik) + +```bash +# Get list of services +curl -u traefik:YOUR_TRAEFIK_API_PASSWORD http://localhost:8080/api/http/routers + +# Get Traefik metrics +curl http://localhost:8080/metrics + +# Reload configuration (live) +curl -X POST http://localhost:8080/api/http/routers -H "Content-Type: application/json" -d '{...}' +``` + +## Migration Checklist + +- [ ] Deploy Traefik in host network mode +- [ ] Create Let's Encrypt certificate for your domain +- [ ] Migrate all reverse proxy rules to Traefik labels or dynamic config +- [ ] Test SSL certificates work: `curl -k https://your-domain.nuclide.systems` +- [ ] Remove Zoraxy Docker container +- [ ] Update DNS if needed +- [ ] Verify all endpoints: health checks at new URLs + +## Benefits + +1. **Zero maintenance SSL** - Let's Encrypt auto-renews +2. **API-driven** - No manual Web UI needed +3. **Hot reloading** - Changes apply immediately +4. **Docker-native** - Watches container labels automatically +5. **Enterprise-grade** - Used by major cloud providers diff --git a/stacks/docs/zoraxy-setup.md b/stacks/docs/zoraxy-setup.md new file mode 100644 index 0000000..d331e86 --- /dev/null +++ b/stacks/docs/zoraxy-setup.md @@ -0,0 +1,182 @@ +# Zoraxy Reverse Proxy Setup Guide + +## Overview +Configure reverse proxies for AI service endpoints with automatic SSL certificates. + +## Target Services + +| Service | Domain | Target | Port | +|---------|--------|--------|------| +| LiteLLM | `litellm.nuclide.systems` | `192.168.1.40` | `14000` | +| LobeHub | `chat.nuclide.systems` | `192.168.1.40` | `14001` | +| MCP Gateway | `mcp.nuclide.systems` | `192.168.1.40` | `8080` | +| Garage S3 | `s3.nuclide.systems` | `192.168.1.40` | `10004` | + +## Option A: Web UI (Easiest - No SSH Required) + +### Step 1: Access Zoraxy Panel +``` +http://192.168.1.4:8000 +``` + +### Step 2: Navigate to Rewrite Rules +1. Go to **"Reverse Proxy"** in the left sidebar +2. Click **"Rewrite Rules"** + +### Step 3: Add Each Service +Click **"Add New Rule"** and fill in: + +**For LiteLLM (`litellm.nuclide.systems`):** +- **Domain:** `litellm.nuclide.systems` +- **SSL:** `ON` +- **HTTPS Redirect:** `ON` +- **HSTS:** `ON` +- **Target Type:** `Server` +- **Target:** `192.168.1.40:14000` +- **Timeout:** `300` seconds +- **Proxy Time:** `300` seconds + +**For LobeHub (`chat.nuclide.systems`):** +- **Domain:** `chat.nuclide.systems` +- **SSL:** `ON` +- **HTTPS Redirect:** `ON` +- **HSTS:** `ON` +- **Target Type:** `Server` +- **Target:** `192.168.1.40:14001` +- **Timeout:** `300` seconds +- **Proxy Time:** `300` seconds + +**For MCP Gateway (`mcp.nuclide.systems`):** +- **Domain:** `mcp.nuclide.systems` +- **SSL:** `ON` +- **HTTPS Redirect:** `ON` +- **HSTS:** `ON` +- **Target Type:** `Server` +- **Target:** `192.168.1.40:8080` +- **Timeout:** `300` seconds +- **Proxy Time:** `300` seconds + +**For Garage S3 (`s3.nuclide.systems`):** +- **Domain:** `s3.nuclide.systems` +- **SSL:** `ON` +- **HTTPS Redirect:** `ON` +- **HSTS:** `ON` +- **Target Type:** `Server` +- **Target:** `192.168.1.40:10004` +- **Timeout:** `300` seconds +- **Proxy Time:** `300` seconds + +### Step 4: Save All Rules +Click **"Save"** at the bottom of the page. + +### Step 5: SSL Certificate Generation +- Zoraxy automatically initiates Let's Encrypt certificate generation +- Wait 1-2 minutes for certificates to be issued +- Visual indicator: Green padlock icon in the UI + +## Option B: SSH/UCI (Automated - Requires Root Access) + +### Prerequisites +- SSH access to Zoraxy with root credentials +- Command-line access to Zoraxy's UCI config system + +### Step 1: SSH into Zoraxy +```bash +ssh root@192.168.1.4 +``` + +### Step 2: Configure via UCI +Run this command directly on Zoraxy: + +```bash +# Apply all rewrite rules +uci set network.rewrite_litellm.enable='1' +uci set network.rewrite_litellm.domain='litellm.nuclide.systems' +uci set network.rewrite_litellm.target_type='0' +uci set network.rewrite_litellm.target='192.168.1.40:14000' +uci set network.rewrite_litellm.ssl='1' +uci set network.rewrite_litellm.hsts='1' + +uci set network.rewrite_lobehub.enable='1' +uci set network.rewrite_lobehub.domain='chat.nuclide.systems' +uci set network.rewrite_lobehub.target_type='0' +uci set network.rewrite_lobehub.target='192.168.1.40:14001' +uci set network.rewrite_lobehub.ssl='1' +uci set network.rewrite_lobehub.hsts='1' + +uci set network.rewrite_mcp.enable='1' +uci set network.rewrite_mcp.domain='mcp.nuclide.systems' +uci set network.rewrite_mcp.target_type='0' +uci set network.rewrite_mcp.target='192.168.1.40:8080' +uci set network.rewrite_mcp.ssl='1' +uci set network.rewrite_mcp.hsts='1' + +uci set network.rewrite_garage.enable='1' +uci set network.rewrite_garage.domain='s3.nuclide.systems' +uci set network.rewrite_garage.target_type='0' +uci set network.rewrite_garage.target='192.168.1.40:10004' +uci set network.rewrite_garage.ssl='1' +uci set network.rewrite_garage.hsts='1' + +# Commit and restart +uci commit network +/etc/init.d/network restart +``` + +### Step 3: Verify +```bash +# Check if rewrite rules are loaded +uci show network.rewrite_* + +# Test connectivity +curl -I http://litellm.nuclide.systems +``` + +## Option C: Automated Script (Local) + +Run from the control node: + +```bash +# Generate UCI config +/opt/stacks/scripts/configure_zoraxy_helpers.sh --export-uci + +# Or deploy via SSH (requires root@192.168.1.4 SSH access) +/opt/stacks/scripts/configure_zoraxy_helpers.sh --deploy +``` + +## Verification + +After setup, verify all endpoints: + +```bash +# Test each service +curl -I https://litellm.nuclide.systems/health +curl -I https://chat.nuclide.systems/api/health +curl -I https://mcp.nuclide.systems/health +curl -I https://s3.nuclide.systems/ +``` + +Expected: HTTP 200 with SSL certificate. + +## Troubleshooting + +### SSL Certificate Issues +- Wait 2-3 minutes for Let's Encrypt to respond +- Check domain DNS resolution: `dig litellm.nuclide.systems` +- Verify firewall allows port 80 validation + +### Proxy Connection Failed +- Check backend service is running: `htop` or `docker ps` +- Verify internal IP: `192.168.1.40` is reachable from Zoraxy +- Check target port is correct + +### Configuration Not Saving +- Check Web UI is logged in (session expiry) +- Try clearing browser cache +- Verify UCI config syntax + +## Notes +- SSL certificates are **automatic** via Let's Encrypt +- Zoraxy auto-renews certificates before expiry +- HSTS is recommended for production use +- All services use the same target IP (`192.168.1.40`) diff --git a/stacks/nexa/CLAUDE.md b/stacks/nexa/CLAUDE.md new file mode 100644 index 0000000..036eb22 --- /dev/null +++ b/stacks/nexa/CLAUDE.md @@ -0,0 +1,74 @@ +# Instructions for Claude (and other agents) + +This file tells future automated runs what they need to know about this repo. + +## Repo conventions + +- **Documentation:** all docs live in `/docs/` and are numbered. Entry point is `docs/index.md`. When you add a doc, give it the next free `NN-` prefix and add a row to the index TOC. +- **Runtime artifacts:** live in `nexa-core/` (workflows, prompts, configs, scripts). Don't put `.md` documentation in there — link from `/docs/` instead. +- **Source-of-truth:** if a doc duplicates content from `nexa-core/config/*.md`, **delete the duplicate**. Single source of truth. + +## Real infrastructure (verified from screenshots, May 2026) + +- **Proxmox host** `nuc` at `192.168.1.20:8006` (PVE 9.1.9, kernel 6.17.13-4-pve, EFI). + - Hardware: **22 threads** (Intel Core Ultra 7 155H, 1 socket), **62 GiB RAM**, **1.64 TiB disk**. + - Steady state: ~32 GiB used (≈24 GiB of which is ZFS ARC, tunable via `zfs_arc_max`), CPU load <2.0, IO delay <0.05%. + - LXC 102 dns (AdGuard) — internal DNS, rewrites for `*.nuclide.systems`. + - LXC 103 backrest — backup orchestration. + - LXC 104 docker — main docker host at `192.168.1.40`, hostname `docker`, OS Debian 13. **Unprivileged**, originally provisioned from the Dockge helper-script template, but **Arcane is the active docker manager** today (Dockge is stale, slated for retirement — see [docs/12 #33](./docs/12-optimization-opportunities.md#33)). **Dozzle** is the live log viewer. Allocated: 16 CPU, 31.25 GiB RAM (25% used), 8 GiB swap, 200 GiB boot disk (47.7% used). **Storage = Ubiquiti UNAS Pro at `192.168.1.31`** (UniFi Drive 4.1.16 on UniFi OS 5.0.17, SFP+ 10 GbE, RAID 6, 19.96 TiB raw / 2.05 TiB used). NFS-exported at `/var/nfs/shared/storage`, mounted by Proxmox at `/mnt/pve/unas`, also exposed via SMB at `smb://192.168.1.31/` (Mac) / `\\192.168.1.31\` (Win). + + **Standard pattern for docker volumes (verified via Karakeep, Q19): plain host bind-mount of `/mnt/pve/unas/services//` from inside LXC 104.** No `driver_opts`, no CIFS, no credentials in the compose. Karakeep, Immich and the rest do exactly this. Nexa follows suit. SMB-as-docker-volume is documented as an **escape hatch only** ([docs/12 #27](./docs/12-optimization-opportunities.md#27)) for services that hit Nextcloud-style NFS issues — Nexa doesn't, so we don't use it. + + **Storage-layer snapshots are NOT configured** on UNAS Pool 1 ("Click to Setup" in the UniFi Drive dashboard). All 2 TB of homelab data has no point-in-time protection at the storage layer — Backrest covers files, not "the whole pool last Tuesday". Highest-leverage fix in the homelab right now ([docs/12 #38](./docs/12-optimization-opportunities.md#38)). + + **UNAS layout conventions** (homelab-wide, all docker containers follow them): + - **`services//` is the general docker config store** — every container in LXC 104 binds its persistent data here. Existing tenants observed: immich, karakeep, nextcloud, ntfy, paperless-ai, pocketid, shelfmark, stremio, traccar, vaultwarden, gluetun. **Stale (retire, do not consume):** `services/siyuan/` (migrated to Obsidian), `services/open-webui/` (unused — LobeHub is the active LLM UI), and `services/dockge/` if it exists (retired in favor of Arcane). **Nexa MUST follow the same pattern: `services/nexa/{qdrant,tei-cache,graphdb,...}`.** Don't invent a parallel layout. + - `backup//` — per-service backups (existing: home-assistant tars, immich pgdump, nextcloud borg). **Nexa snapshots → `backup/nexa/`.** + - `media/`, `code/`, `_sortMe/`, `dump/`, `test_perm` — user data, not Nexa's concern. + - **Before deploying any Nexa container, READ AN EXISTING STACK** in **Arcane** (e.g. `karakeep` or `immich`) to confirm the exact mount syntax in use — driver name, share path, credential injection pattern. Match it. The actual NFS export root is `/var/nfs/shared/storage`; the SMB share name is **still TBD — see Q19** in `docs/11`. + +**Hard-blocklist for any Nexa indexer / agent** (never read these paths or matching glob): + - `_sortMe/wallet/**` — contains PGP keys + bitcoin wallet files. + - Any path matching `*.gpg`, `*.asc`, `*.key`, `*.pem`, `id_rsa*`, `*wallet*`, `*.kdbx`, `*credentials*`, `*secret*`. + - The Nextcloud appdata dir (`services/nextcloud/appdata_*`) — Nextcloud-internal, not user content. Intel iGPU passthrough is *configured but currently broken* — see [docs/12 #26](./docs/12-optimization-opportunities.md#26). + - LXC 105 nextcloud — Nextcloud at `nc.nuclide.systems`. + - LXC 106 octoprint — currently *Exited*; flagged in [docs/11](./docs/11-open-questions.md#q13). + - LXC 108 zoraxy — reverse proxy at `192.168.1.4:8000`, TLS for `*.nuclide.systems`. + - VM 100 haos — Home Assistant. +- **Already-running services on docker host** (don't redeploy): + - Memos `:5230`, n8n `:5678`, LiteLLM `:4000` (UI LobeHub `:3210`), Qdrant (`qdrant_scientific`), ntfy `:7998`, Karakeep (legacy alias `hoarder.nuclide.systems`), Vaultwarden `:11001`, Pocket-ID `:1411`, Immich, Audiobookshelf, Paperless-ngx, Traccar, Prowlarr, plus MCP containers (`crawl4ai-mcp`, `markitdown-mcp`, `papersearch-mcp`). +- **Octoprint (LXC 106)** is *intentionally* powered down most of the time. Phase-5 monitoring must skip names matching `octoprint*` rather than alert on its Exited state. +- **Obsidian vault lives inside Nextcloud** at `nc.nuclide.systems/Notizen/` (multi-device sync via Nextcloud client). Nexa accesses it via WebDAV — read-only, no filesystem mount. Ignore list: `.copilot/`, `.copilot-index/`, `.smart-env/`, `.caldav-sync/`, `assets/` (visual queue, Phase 3.2), `Templates/`, `BMO/`, `Excalidraw/`. Index target: `Notizen/**/*.md`. +- **Nextcloud Tasks lists & calendars** (German names, may grow over time): + - `Persönlich` → Personal context. + - `DLR` → Work context (DLR is the user's employer). + - `Einkaufsliste` → Shopping. + - `Wunschliste` → Wishes. + Lists are discovered **by name** at runtime (Qdrant `_config` namespace caches `name → id`). Never hardcode IDs. The discovery workflow runs daily and on cache-miss; new lists added in Nextcloud are honoured automatically next refresh. +- **Mail = Nextcloud Mail, single account `fkrebs@nucli.de`.** No separate IMAP entry. The `Waiting` folder is a manual user signal — items there are skipped from digests. +- **Backup model is 3-2-1**: UNAS native snapshots → `s3.nuclide.systems` (warm, on-site) → encrypted off-site cold tier (provider TBD, Jottacloud is the user's candidate — see [Q20](./docs/11-open-questions.md#q20)). Always `restic`/`rclone-crypt` before upload — third-party provider sees only ciphertext. Don't propose alternative backup paths without checking docs/12 #37 first. +- **Auth = Pocket-ID SSO is global** at the Zoraxy layer. Don't add app-level basic-auth to Nexa surfaces; UIs inherit SSO. Machine-to-machine still uses API keys / app passwords. +- **Decided for Nexa** (don't re-litigate without user input): + - Vector store: **reuse `qdrant_scientific`** with collections suffixed by modality (`nexa_knowledge_text`, `nexa_knowledge_visual`). + - Embeddings staged: **Phase 3.1** TEI + `BAAI/bge-m3` (text-only, 1024-dim). **Phase 3.2** swap to `infinity` and add `jinaai/jina-clip-v2` (768-dim, joint text+image space). All forward-compat fields (`modality`, `media_uri`, `graph_iri`, `nexa:pendingVisualIndex`) exist from 3.1 — adding the visual collection is additive. + - Graph store: **Ontotext GraphDB** (SPARQL/RDF), Phase 3.4. RDF schema in docs/08 already includes `nexa:modality` / `nexa:mediaUri` / `nexa:vectorCollection` / `nexa:pendingVisualIndex`. + - Chat model: SAIA via LiteLLM virtual key. + +## When working on Nexa + +1. **Read `docs/index.md` first** — it's the navigator. +2. **Open questions first.** Before writing code or workflow JSON, scan `docs/11-open-questions.md`. If your task touches an unanswered Q, **stop and ask** rather than picking a default. Append new blockers to that doc as `[ ] Q-NN`. +3. **Optimization findings.** When you spot infrastructure improvements, add them to `docs/12-optimization-opportunities.md` as a numbered bullet — don't just mention them in commit messages. +4. **Never inline secrets** in workflow JSON or `.env` committed to git. Use n8n credentials, LiteLLM virtual keys, or (longer term) Vaultwarden. +5. **Keep deployment minimal.** The default answer to *"do we need a new container?"* is **no** — the existing stack covers most needs. + +## Branch policy + +- This branch is `claude/organize-docs-deployment-7N4v2`. Push only here unless told otherwise. +- New work for an unrelated feature → new branch under `claude/`. + +## Quick links + +- [docs/index.md](./docs/index.md) +- [docs/09-deployment.md](./docs/09-deployment.md) — most-touched file during bring-up +- [docs/11-open-questions.md](./docs/11-open-questions.md) — read **before** assuming defaults diff --git a/stacks/nexa/README.md b/stacks/nexa/README.md new file mode 100644 index 0000000..fdb68b0 --- /dev/null +++ b/stacks/nexa/README.md @@ -0,0 +1,17 @@ +# nexa + +**Neural Nexus for Information & Automation** — central nervous system for a personal IT setup. + +> Memos = voice & ear · n8n = reflexes · SAIA (LiteLLM) = brain · Qdrant (+ optional graph DB) = memory · Nextcloud = hands. + +## Start here + +📖 **[docs/index.md](./docs/index.md)** — TOC, reading paths, repo layout. + +## Repository layout + +``` +nexa/ +├── docs/ ← all documentation, numbered for reading order +└── nexa-core/ ← runtime: n8n workflows, configs, prompts, scripts +``` diff --git a/stacks/nexa/docs/01-vision-and-scope.md b/stacks/nexa/docs/01-vision-and-scope.md new file mode 100644 index 0000000..d6d9366 --- /dev/null +++ b/stacks/nexa/docs/01-vision-and-scope.md @@ -0,0 +1,34 @@ +# 01 — Vision & Scope + +## Vision + +Nexa is the **central nervous system** of a personal IT setup: an intelligent middleware sitting between input sources (Memos, e-mail, RSS, Karakeep, Bluesky), knowledge stores (Obsidian, Qdrant, Nextcloud) and organization tools (Nextcloud Calendar/Tasks). + +### Core goals + +- **Cognitive offload** — Nexa sorts, filters and proposes; it doesn't just store. +- **Context separation** — clean AI-driven split between *Work* and *Personal*. +- **Single point of interaction** — Memos is the primary surface ("voice & ear"). +- **Knowledge synergy** — link ephemeral memos to deep notes via semantic search. + +## Scope + +### In-scope + +- Triaging ~20 personal e-mails per day. +- Routing tasks to the right Nextcloud list (Work vs. Personal). +- Suggesting time-boxes in the work calendar. +- Long-term memory in Qdrant (and optionally a graph DB). +- Aggregating critical IT alerts (Proxmox, Backrest) into a Memos system feed. + +### Out-of-scope + +- Direct access to work IT (work mail, work servers). +- Active document editing (metadata extraction only). +- Replacing specialist UIs (Arcane / Dozzle / Portainer for containers). + +## Success metrics + +- **Reduced mail interaction** — less time in the mail app, more on the Nexa digest. +- **Calendar fill-rate** — focus slots auto-blocked on the work calendar. +- **Search latency** — old memos / Obsidian notes resurface instantly via `#nexa:ask`. diff --git a/stacks/nexa/docs/02-roadmap.md b/stacks/nexa/docs/02-roadmap.md new file mode 100644 index 0000000..b46cdf7 --- /dev/null +++ b/stacks/nexa/docs/02-roadmap.md @@ -0,0 +1,62 @@ +# 02 — Roadmap & Phases + +Iterative build-out, value-first. Each phase is shippable on its own. + +## Phase 1 — The Spine (connectivity) + +**Focus:** datapath between Memos, n8n and SAIA. +**Milestone:** Nexa replies on Memos and answers simple questions. + +- [ ] **1.1 Memos ↔ n8n bridge** — webhook ingress, content filter, comment-back. +- [ ] **1.2 SAIA integration** — call LiteLLM proxy from n8n; chat via Memos comments. +- [ ] **1.3 Git-sync** — periodic export of n8n workflows into this repo. + +## Phase 2 — Senses (input channels) + +**Focus:** e-mail filter and the Work-vs-Personal router. +**Milestone:** Nexa distinguishes work tasks from personal tasks without manual tags. + +- [ ] **2.1 Email butler** — filter & summarize the ~20 daily personal mails. +- [ ] **2.2 RSS morning digest** — curated daily summary in Memos. Items also enter `nexa_knowledge_text` with `source_type=rss, ttl=30d` (see [08 § Memory sources & retention](./08-graphrag-architecture.md#memory-sources--retention)). +- [ ] **2.3 Bluesky antenna** — store liked posts as knowledge snippets. +- [ ] **2.4 On-demand web fetch** — `#nexa:ask --web` runs a SearXNG query through `redis-searxng`, fetches top hits via `crawl4ai-mcp` + `markitdown-mcp`, embeds and answers. Pages stay in memory with `ttl=90d` so the next related question doesn't re-fetch. + +## Phase 3 — Memory (Qdrant & RAG) + +**Focus:** Qdrant + graph layer for retrieval-augmented answers. +**Milestone:** RAG works — Nexa answers from archived notes; images captured today are queued for visual indexing later. + +- [ ] **3.1 Text RAG (Path A)** — TEI + `bge-m3`, single Qdrant collection `nexa_knowledge_text`. `#nexa:ask` reads it before answering. Image attachments are *recorded* as `nexa:Note` with `pendingVisualIndex` but not embedded yet. +- [ ] **3.2 Visual collection (Path C)** — swap TEI → infinity, add `jina-clip-v2`, create `nexa_knowledge_visual`, backfill all queued image notes. RAG workflow gets a parallel branch for visual hits. See [09 § "Phase add-on: visual collection"](./09-deployment.md#phase-add-on-visual-collection-phase-32). +- [ ] **3.3 S3 long-term archive** — Qdrant snapshots (both collections) to the S3 server. **Deferred** until [Q12 + Q20](./11-open-questions.md#backups-deferred) are decided. Until then, snapshots stay on UNAS at `backup/nexa/snapshots/qdrant//`; UNAS RAID 6 + native pool snapshots ([12/#38](./12-optimization-opportunities.md#38)) cover the data-loss scenarios that matter most for Phase 3.1. +- [ ] **3.4 Ontotext GraphDB** — SPARQL pillar; structural queries combined with vector hits in `#nexa:ask`. See [08](./08-graphrag-architecture.md). + +## Phase 4 — Motor (organization & action) + +**Focus:** calendar integration and time-boxing. +**Milestone:** Nexa proactively proposes morning focus slots in the work calendar. + +- [ ] **4.1 Nextcloud Tasks sync** — bidirectional status sync (Memos ↔ NC Tasks). +- [ ] **4.2 Time-boxing agent** — calendar-aware slot suggestions. +- [ ] **4.3 Karakeep curation** — automated review cycle for saved links. + +## Phase 5 — Daily integration (polish) + +**Focus:** Home Assistant Voice and monitoring. +**Milestone:** Nexa speaks via HA-Voice and surfaces critical system states proactively. + +- [ ] **5.1 HA Voice** — Wake-word integration, Wyoming protocol. +- [ ] **5.2 System monitoring** — Proxmox / Backrest alerts into the Memos system feed via ntfy. +- [ ] **5.3 Morning digest with self-asking re-loop** — daily 07:00 digest aggregates mail + tasks + drift report. When the digest has spare capacity, Nexa re-asks ONE open question from [docs/11](./11-open-questions.md) using a backoff schedule (3 d → 7 d → 21 d → 60 d). Answers parsed from the user's reply auto-resolve the question and emit a `docs/11` diff for review (Phase 6.4 hook). See [docs/05 § #nexa:digest](./05-command-system.md). + +## Phase 6 — Nexa as homelab steward + +**Focus:** Nexa actively maintains the homelab inventory instead of being told manually. +**Milestone:** Nexa runs a daily *drift report* — what changed, what's stale, what's missing — and turns it into actionable comments under a pinned `[STEWARD]` memo. + +- [ ] **6.1 Inventory sync** — Nexa polls Arcane / `docker ps` / Proxmox API on a schedule and stores the running-state snapshot in GraphDB (`nexa:Service`, `nexa:LastSeen`, `nexa:DataPath`, …). +- [ ] **6.2 Drift detection** — diff against the documented desired state in this repo (`docs/12`, `CLAUDE.md`, the housekeeping campaign at #35–37). Emit findings as structured Memos comments: *"`open-webui` last used 38 d ago — retire?"*, *"Container `foo` has no `services/foo/` bind — schedule UNAS migration?"*. +- [ ] **6.3 Steward commands** — `#nexa:retire ` (tar to `backup/`, stop, archive the stack), `#nexa:document ` (template a CLAUDE.md entry), `#nexa:wishlist-status` (which entries in [docs/13](./13-information-wishlist.md) are still needed). +- [ ] **6.4 Self-update of docs** — when the user resolves an open question via Memos, Nexa proposes the corresponding edit to `docs/11-open-questions.md` as a PR-ready diff (the doc itself becomes a Nexa-managed surface). + +This phase is what turns Nexa from "an assistant that answers questions" into "an assistant that takes care of its own runtime", and it's the natural home for the cross-cutting [housekeeping campaign](./12-optimization-opportunities.md#housekeeping-campaign-cross-cutting-do-together) — most of those items become semi-automatic once 6.1–6.3 ship. diff --git a/stacks/nexa/docs/03-architecture.md b/stacks/nexa/docs/03-architecture.md new file mode 100644 index 0000000..3e0e9d1 --- /dev/null +++ b/stacks/nexa/docs/03-architecture.md @@ -0,0 +1,64 @@ +# 03 — Architecture Overview + +A one-page mental model. For details follow the cross-links. + +## Topology + +``` + ┌────────────────┐ + voice / typing ──────▶│ Memos │◀──── Nexa replies as comments + │ (interface) │ + └───────┬────────┘ + │ webhook (- [ ] / #nexa:*) + ▼ + ┌────────────────┐ + IMAP / RSS / NC ────▶│ n8n │◀──── workflows live in + Karakeep / Bluesky │ (logic) │ ./nexa-core/n8n-workflows + └───┬────────┬───┘ + classify ▲ │ │ ▲ retrieve + │ ▼ ▼ │ + ┌──────────┐ ┌──────────┐ + │ SAIA │ │ Qdrant │ + │ LiteLLM │ │ vector │ + └──────────┘ └────┬─────┘ + │ + ▼ + ┌──────────┐ + │ GraphDB │ Ontotext, SPARQL + │ (RDF) │ Phase 3.4 + └──────────┘ + │ + ▼ writes + ┌──────────────────────────────┐ + │ Nextcloud (Tasks, Calendar, │ + │ Mail, Files / Obsidian) │ + └──────────────────────────────┘ +``` + +## Components + +| Component | Role | Where it runs (today) | +|-----------|------|-----------------------| +| **Memos** | Interface, voice input, webhook source | docker host LXC 104 (unprivileged, 16 CPU / 31 GiB / 200 GiB) → `memos.nuclide.systems` | +| **n8n** | Workflow / logic engine | docker host LXC 104 → `n8n.nuclide.systems` | +| **SAIA / LiteLLM** | Model gateway, embeddings, classification | docker host LXC 104 → `ai.nuclide.systems` (LiteLLM internal :4000) | +| **Qdrant** | Vector memory (semantic recall) | docker host LXC 104 — reuse `qdrant_scientific`. Collections: `nexa_knowledge_text` (Phase 3.1, 1024-dim) and `nexa_knowledge_visual` (Phase 3.2, 768-dim) | +| **Ontotext GraphDB** | Structural memory via SPARQL (Phase 3.4) | not yet deployed; see [09-deployment](./09-deployment.md#phase-add-on-ontotext-graphdb-phase-34) | +| **TEI** (HF text-embeddings-inference) | Self-hosted text embeddings, `BAAI/bge-m3`, Phase 3.1 | docker host LXC 104, CPU only — swapped for `infinity` in Phase 3.2 to add `jina-clip-v2` | +| **Nextcloud** | Tasks, calendar, mail, files | dedicated LXC 105 → `nc.nuclide.systems` | +| **ntfy** | Push channel for system alerts | docker host → `ntfy.nuclide.systems` | +| **Backrest** | Backup orchestration | LXC 103 | +| **Zoraxy** | Reverse proxy + TLS | LXC 108 (`192.168.1.4:8000`) | +| **AdGuard DNS** | Internal name resolution | LXC 102 | +| **Home Assistant** | Voice + house automation | VM 100 (HAOS) | + +## Two-pillar memory + +- **Qdrant** answers *"what is similar / relevant?"* (cosine search over embeddings). +- **Ontotext GraphDB** answers *"who, what depends on what, how is it structured?"* (SPARQL over RDF). + +Both pillars are queried in parallel for `#nexa:ask` and merged before SAIA generates the final answer. See [08 — GraphRAG architecture](./08-graphrag-architecture.md). + +## Dual-context routing + +Every input is classified `work` or `personal` before any side effect (task creation, calendar write). See [04 — Integration matrix](./04-integration-matrix.md) and [06 — Classification logic](./06-classification-logic.md). diff --git a/stacks/nexa/docs/04-integration-matrix.md b/stacks/nexa/docs/04-integration-matrix.md new file mode 100644 index 0000000..2d374f2 --- /dev/null +++ b/stacks/nexa/docs/04-integration-matrix.md @@ -0,0 +1,79 @@ +# 04 — Integrations-Matrix + +Diese Matrix definiert die logische Trennung zwischen privaten und beruflichen Datenströmen sowie die Anbindung der Infrastruktur. + +--- + +## 1. Die Dualität: Arbeit vs. Privat +Nexa muss strikt zwischen zwei Kontexten unterscheiden, da die Datenquellen variieren: + +### A. Bereich: ARBEIT (Work) +* **Sichtbarkeit:** Eingeschränkt (kein Zugriff auf Arbeits-E-Mails oder Firmen-Server). +* **Datenquellen:** + * Nextcloud Task-Liste **`DLR`** (sowie der gleichnamige Kalender). + * Memos mit Tag `#work` oder semantischer Erkennung. +* **Nexa-Fokus:** Timeblocking, Fokus-Zeiten, Vorbereitung von Meetings basierend auf Obsidian-Notizen. + +### B. Bereich: PRIVAT (Personal) +* **Sichtbarkeit:** Vollständig. +* **Datenquellen:** + * Nextcloud Task-Liste **`Persönlich`** und gleichnamiger Kalender. + * Nextcloud Mail (`fkrebs@nucli.de`, einziges Konto). Ordner `Waiting` ist ein manuelles Signal — Mails dort werden vom Digest ausgeschlossen. + * Bluesky, RSS, Karakeep. +* **Nexa-Fokus:** Automatisierung des Alltags, Kuratierung von Wissen, E-Mail-Management. + +### Karakeep & Shopping-Logik +* **Schnittstelle:** Karakeep API / RSS. +* **Listen-Routing:** + * **Einkaufsliste:** Automatischer Sync für Items mit hoher Frequenz (Lebensmittel, Drogerie). + * **Wunschliste:** Speicherort für "Entdeckungen". Nexa fügt bei Wunschlisten-Items automatisch den aktuellen Preis und eine kurze SAIA-Zusammenfassung hinzu ("Warum du das speichern wolltest"). +* **Review-Cycle:** Einmal im Monat fragt Nexa bei Wunschlisten-Items nach: "Immer noch interessiert oder kann das weg?" + +--- + +## 2. Kern-Infrastruktur (The Brain & Spine) + +| Komponente | Rolle im System | +| :--- | :--- | +| **Memos** | Zentraler Input (Drafts, Ideen, schnelle Tasks). Schnittstelle für Nexa-Antworten. | +| **n8n** | Logik-Engine. Führt die Klassifizierung Arbeit vs. Privat durch. | +| **SAIA (LiteLLM)** | Entscheidet anhand des Inhalts, in welchen Kalender/Liste ein Eintrag gehört. | +| **Qdrant** | Langzeitgedächtnis. Speichert Projektwissen (Work) und privates Wissen getrennt. | +| **Arcane** | Verwaltung der Docker-Container (n8n, Qdrant, Memos). | + +--- + +## 3. Spezifische Datenflüsse & Logik + +### Task-Routing (Das Gehirn-Filter) +1. **Input:** Neues Memo oder Spracheingabe. +2. **Analyse:** SAIA prüft: "Ist das Business oder Privat?" +3. **Routing:** * `Arbeits-Kontext` -> Eintrag in `Work_Tasks`. + * `Privat-Kontext` -> Eintrag in `Personal_Tasks`. +4. **Timeboxing:** Nexa scannt den `Work_Calendar` auf Lücken und schlägt Slots für `Work_Tasks` vor. + +### E-Mail Management (Nur Privat) +* Nexa filtert die 20 täglichen Privat-Mails. +* Wichtige private Termine werden in den privaten Kalender extrahiert. +* Newsletter landen als Zusammenfassung in Memos. + +### Wissens-Management (Obsidian via Nextcloud & Qdrant) +* Der Obsidian-Vault liegt im Nextcloud-Ordner `Notizen/` (Multi-Device-Sync via Nextcloud-Client). Nexa liest **read-only über WebDAV** (`/remote.php/dav/files//Notizen/`) — kein Mount, kein zweiter Sync-Mechanismus. +* Markdown-Dateien (`Notizen/**/*.md`) werden in Qdrant (`nexa_knowledge_text`) indiziert; Bilder unter `assets/` landen im Phase-3.2-Visual-Queue (siehe [docs/11/Q4](./11-open-questions.md)). +* Plugin-Internals (`.copilot/`, `.smart-env/`, `.caldav-sync/`, `Templates/`, `BMO/`, `Excalidraw/`) sind explizit von der Indizierung ausgeschlossen. +* Nexa nutzt dieses Wissen, um Arbeits-Tasks besser zu beschreiben, auch wenn sie keinen Zugriff auf die Arbeits-E-Mails hat. + +--- + +## 4. Monitoring & Hardware (Home Assistant) + +| Dienst | Reiz-Typ | Nexa-Reaktion | +| :--- | :--- | :--- | +| **Home Assistant** | Voice / Sensorik | Nexa nimmt Sprachbefehle entgegen und meldet Alarme (Haus). | +| **Proxmox** | Stabilität | Meldung an Nexa bei Hardware-Problemen. | +| **S3 Server** | Archiv | Langzeit-Backups der Wissensdatenbank. | + +--- + +## 5. Ergänzungen für das Setup (Vorschlag) +* **Zwei-Faktor-Klassifizierung:** Einführung von expliziten Tags (`#w` für Work, `#p` für Privat) für Grenzfälle, in denen SAIA den Kontext nicht eindeutig bestimmen kann. \ No newline at end of file diff --git a/stacks/nexa/docs/05-command-system.md b/stacks/nexa/docs/05-command-system.md new file mode 100644 index 0000000..12701cc --- /dev/null +++ b/stacks/nexa/docs/05-command-system.md @@ -0,0 +1,176 @@ +# 05 — Command System (`#nexa:*`) + +Nexa "hört" auf folgende Kommandos in Memos-Kommentaren oder als Memo-Inhalt mit Hashtag: + +## 🔧 Konfigurations-Kommandos + +### #nexa:config +- **Beschreibung**: Triggert Autodiscovery aller Services (Nextcloud Listen/Kalender, Qdrant Collections, LiteLLM-Modelle, Mail-Konto). +- **Läuft auch automatisch**: + - Täglich (cron in n8n) — neue Listen / Kalender werden ohne manuelles Eingreifen aufgenommen. + - Bei Cache-Miss — wenn ein Workflow eine Liste/ID nicht mehr findet, läuft Discovery einmal, dann Retry. +- **Antwort**: Postet Report mit den aktuellen Konfigurationen als Memos-Kommentar. +- **Speichert**: Konfiguration in der Qdrant `_config` Namespace (nicht in `.env`). + +``` +#nexa:config +``` + +**Beispiel-Response:** +``` +✅ Autodiscovery abgeschlossen (2026-05-04T23:59): +- Nextcloud Listen: Persönlich (id=14), DLR (id=dlr-1), Einkaufsliste (id=…), Wunschliste (id=…) +- Nextcloud Kalender: Persönlich, DLR, Einkaufsliste, Wunschliste +- Mail-Konto: fkrebs@nucli.de (Posteingang, Archiv, Junk, Waiting) +- Qdrant Collection: nexa_knowledge_text (1024 dims, Cosine, 0 Punkte) +- LiteLLM-Modelle: nexa-chat, nexa-embed +``` + +### #nexa:status +- **Beschreibung**: Zeigt den aktuellen System-Status +- **Antwort**: Uptime, verbundene Systeme, Memory-Stats + +``` +#nexa:status +``` + +### #nexa:sync-obsidian +- **Beschreibung**: Triggert sofortige Synchronisierung mit Obsidian Vault +- **Optional Parameter**: `--vault=/path/to/vault` oder `--force` für Neuindexierung + +``` +#nexa:sync-obsidian --force +``` + +## 📊 Memory & RAG Kommandos + +### #nexa:ask [question] +- **Beschreibung**: Semantische Suche in Qdrant (RAG) basierend auf Obsidian + Memos +- **Antwort**: Top-3 Treffer mit Quellen + +``` +#nexa:ask Wie implementierten wir das JWT-Middleware-Pattern? +``` + +### #nexa:learn [topic] +- **Beschreibung**: Explizit neue Information in Qdrant speichern (mit Tags) +- **Optional**: `--tag=architecture` `--source=obsidian/notes/arch.md` `--permanent` (überschreibt Default-TTL) + +``` +#nexa:learn Das Routing-Schema unterscheidet Work vs. Personal via SAIA-Kontext-Analyse --tag=architecture +``` + +### #nexa:forget [query-or-iri] +- **Beschreibung**: Löscht passende Einträge aus Qdrant + GraphDB sofort. +- **Beispiel**: `#nexa:forget --source=rss --older-than=14d` + +### #nexa:retain [source_type] [days] +- **Beschreibung**: Setzt die Default-Retention für einen Source-Type (Default-Werte siehe [08 § Memory sources & retention](./08-graphrag-architecture.md#memory-sources--retention)). + +### #nexa:ask [question] [--web] +- **Beschreibung**: Semantische Suche in Qdrant. Mit `--web` erweitert Nexa die Suche um eine SearXNG-Abfrage (`redis-searxng`), holt die Top-3 Treffer via `crawl4ai-mcp` + `markitdown-mcp`, embedded sie und nutzt das gemeinsame Material für die Antwort. Geholte Seiten bleiben 90 Tage im Speicher (TTL siehe oben). + +## 🎯 Workflow-Kommandos + +### #nexa:route-test [text] +- **Beschreibung**: Testet die Klassifizierung (Work/Personal) ohne Task zu erstellen + +``` +#nexa:route-test Muss morgen die Präsentation für den Client fertigstellen +``` + +**Response:** +``` +💬 Klassifizierung (Test-Mode): +- Kontext: WORK +- Vertrauen: 0.95 +- Begründung: "Client-Präsentation → professioneller Kontext" +``` + +### #nexa:email-digest +- **Beschreibung**: Triggert sofortige Email-Zusammenfassung (sonst tägl. 7 Uhr) + +``` +#nexa:email-digest +``` + +### #nexa:digest +- **Beschreibung**: Tägliches Morgen-Digest. Läuft automatisch (07:00). Inhalt: + 1. Kurze Mail-Zusammenfassung (siehe `#nexa:email-digest`). + 2. Anstehende Tasks heute (DLR + Persönlich, nach Urgency). + 3. Drift-Report aus Phase 6 (sobald aktiv). + 4. **Re-ask Block** — siehe unten. Maximal 1 offene Frage pro Digest, nur wenn der Digest sonst kurz genug ist. + +#### Re-ask der offenen Fragen (`docs/11-open-questions.md`) +Nexa führt im GraphDB pro offener Frage `nexa:askedCount`, `nexa:lastAskedAt`, `nexa:nextAskAt`. Backoff-Schema: + +| Mal | Wartezeit bis zur nächsten Frage | +|-----|----------------------------------| +| 1 → 2 | 3 Tage | +| 2 → 3 | 7 Tage | +| 3 → 4 | 21 Tage | +| ≥ 5 | 60 Tage | + +Im Morgen-Digest wird **eine** offene Frage angehängt — und zwar nur wenn: +- Der Digest sonst < 800 Zeichen lang wäre (Capacity-Guard, damit es nicht nervt). +- `nextAskAt <= today`. +- Bevorzugt die Frage mit dem **kleinsten** `askedCount` (zuerst neue Fragen, alte selten). + +Antwortet der User direkt unter dem Digest-Memo, parst Nexa die Antwort, markiert die Frage in `docs/11` als resolved (Phase 6.4 — Self-update of docs) und committet einen Diff zur Review. + +### #nexa:remind +- Manuelles Triggern einer Re-Ask, ignoriert Backoff. Nützlich für Rückfragen mitten im Tag. + +### #nexa:answered +- Beantwortet eine offene Frage explizit. Setzt `resolved` im GraphDB und schlägt einen `docs/11`-Diff vor. + +## 🔐 Admin-Kommandos + +### #nexa:reset-config +- **Beschreibung**: Löscht gespeicherte Konfiguration, triggert neuen Autodiscovery +- **Warnung**: Alle benutzerdefinierten Settings werden vergessen + +``` +#nexa:reset-config +``` + +### #nexa:export-state +- **Beschreibung**: Exportiert aktuellen State als JSON (für Backups) + +``` +#nexa:export-state +``` + +--- + +## 📝 Implementierung in n8n + +Ein **Command Parser** läuft immer mit: + +1. **Memos Webhook** empfängt alle Memos +2. **Regex Check**: Sucht nach `#nexa:command` +3. **Router**: Versendet an entsprechenden n8n-Workflow +4. **Antwort**: Postet Reply als Kommentar/Edit + +**Command Parser Regex:** +``` +^#nexa:(\w+)(?:\s+([^\n]*?))?(?:$|\s*--) +``` + +Extrahiert: `[command, parameters]` + +--- + +## 🎛️ Dynamische Konfigurationsspeicherung + +Statt .env zu editieren: + +1. **First Run**: `#nexa:config` speichert zu lokalen Metadata +2. **Speich-Ziel**: + - **Primary**: Qdrant Metadata (als `_config` Namespace) + - **Fallback**: `nexa-core/config/runtime_config.json` (gitignored) + - **Notfall**: .env (nur initial) + +3. **Reload-Logik**: Bei jedem Workflow-Start werden Settings aus Qdrant geladen + +Das macht Nexa vollständig selbstständig nach dem initialem Setup! diff --git a/stacks/nexa/docs/06-classification-logic.md b/stacks/nexa/docs/06-classification-logic.md new file mode 100644 index 0000000..4a5890f --- /dev/null +++ b/stacks/nexa/docs/06-classification-logic.md @@ -0,0 +1,19 @@ +# 06 — Klassifizierungs-Logik + +Dieser Fragebogen dient der Feinabstimmung von SAIA, um Tasks korrekt zu routen. + +### A. Schlüsselwörter & Projekte (Arbeit) +*Welche Begriffe triggern zwingend die Work-Liste?* +- [ ] Projekt-Namen (z.B. "Nexa-Core", "Infrastruktur-Audit") +- [ ] Rollenspezifische Begriffe ("Meeting", "Report", "Deadline") +- [ ] Tools, die nur im Job vorkommen. + +### B. Ausschlusskriterien (Privat) +*Was darf niemals in die Work-Liste?* +- [ ] Lebensmittel, Rezepte, Haushalt. +- [ ] Bluesky-Input (sofern nicht explizit als Recherche markiert). +- [ ] Finanz-Mails (Privat-Bank). + +### C. Umgang mit Unschärfe +- **Default:** Wenn < 70% Sicherheit -> `Personal`. +- **Interaktion:** Nexa erstellt einen Kommentar: "In welche Liste soll das? [W/P]". \ No newline at end of file diff --git a/stacks/nexa/docs/07-workflow-spec.md b/stacks/nexa/docs/07-workflow-spec.md new file mode 100644 index 0000000..c27d79d --- /dev/null +++ b/stacks/nexa/docs/07-workflow-spec.md @@ -0,0 +1,26 @@ +# 07 — Workflow Spec: Phase-2 Task-Router + +## 1. Trigger +* **Typ:** Webhook (Memos) +* **Filter:** `payload.content` enthält `- [ ]` ODER `#todo`. + +## 2. Intelligence Node (SAIA) +* **Prompt:** "Analysiere folgende Notiz. Gib als JSON zurück: `{ "context": "work" | "personal" | "shopping" | "wishes", "summary": "string", "urgency": 1-5 }`." +* **Kontext-Zufuhr:** Lade die letzten 5 Obsidian-Projekttitel aus Qdrant als Referenz hoch. + +## 3. List Resolver (vor dem Switch) +* Lese den Cache `nexa._config.lists` aus Qdrant: `{ "Persönlich": , "DLR": , "Einkaufsliste": , "Wunschliste": , discovered_at: }`. +* Bei Cache-Miss oder 404 vom nachfolgenden Create-Task → triggere `#nexa:config` (Discovery), update den Cache, retry **einmal**. + +## 4. Switch Node (4-Wege) +* **Route 1:** `context == "work"` → Nextcloud Create Task auf Liste `DLR`. +* **Route 2:** `context == "personal"` → Nextcloud Create Task auf Liste `Persönlich`. +* **Route 3:** `context == "shopping"` → Nextcloud Create Task auf Liste `Einkaufsliste`. +* **Route 4:** `context == "wishes"` → Nextcloud Create Task auf Liste `Wunschliste`. +* **Default (Konfidenz < 0.7):** `Persönlich` (siehe [06](./06-classification-logic.md)). + +## 5. Feedback Loop +* Poste die Task-ID und den gewählten Kontext als Kommentar unter das Original-Memo. + +## 6. Listen-Diskoverability +Listennamen sind nicht hardgecodet — die Resolver-Step liest sie aus dem Runtime-Config-Cache. Wenn der User in Nextcloud eine neue Liste anlegt (z. B. `Reisen`), erscheint sie nach dem nächsten geplanten `#nexa:config`-Lauf (täglich) automatisch als Routing-Ziel — die System-Prompt der Intelligence Node wird zusammen mit den verfügbaren Listen versorgt, sodass SAIA neue Kontexte vorschlagen kann. diff --git a/stacks/nexa/docs/08-graphrag-architecture.md b/stacks/nexa/docs/08-graphrag-architecture.md new file mode 100644 index 0000000..2c7d624 --- /dev/null +++ b/stacks/nexa/docs/08-graphrag-architecture.md @@ -0,0 +1,286 @@ +# 08 — GraphRAG: Structural Knowledge & Relations + +Decision: graph layer = **Ontotext GraphDB** with **SPARQL** (resolved in [11/Q1](./11-open-questions.md)). +Rationale: SPARQL + RDF lets Nexa's memory be browsed and queried with the same standard tooling that's used for any open-data corpus, and it leaves the door open for SHACL / OWL reasoning later. + +--- + +## Two-pillar memory + +| Pillar | Question it answers | Backed by | +|--------|---------------------|-----------| +| **Qdrant** (vectors) | *"What is similar / relevant?"* | Cosine search over embeddings | +| **GraphDB** (RDF) | *"What is connected? What depends on what? Who is involved?"* | SPARQL over a typed graph | + +Both pillars are queried in parallel for `#nexa:ask` and merged before SAIA generates the final answer. + +--- + +## RDF schema + +Compact, opinionated. One namespace, one ontology file, no v2/v3 inheritance pain. + +```turtle +@prefix nexa: . +@prefix rdf: . +@prefix rdfs: . +@prefix xsd: . +@prefix prov: . + +# Classes +nexa:Project a rdfs:Class . +nexa:Task a rdfs:Class . +nexa:Person a rdfs:Class . +nexa:Technology a rdfs:Class . +nexa:Topic a rdfs:Class . +nexa:Note a rdfs:Class . # Memos / Obsidian / mail digests +nexa:File a rdfs:Class . + +# Properties +nexa:owns a rdf:Property ; rdfs:domain nexa:Person ; rdfs:range nexa:Task . +nexa:uses a rdf:Property ; rdfs:domain nexa:Task ; rdfs:range nexa:Technology . +nexa:dependsOn a rdf:Property ; rdfs:domain nexa:Task ; rdfs:range nexa:Task . +nexa:childOf a rdf:Property ; rdfs:domain nexa:Task ; rdfs:range nexa:Project . +nexa:mentions a rdf:Property ; rdfs:domain nexa:Note ; rdfs:range nexa:Topic . +nexa:scheduledFor a rdf:Property ; rdfs:domain nexa:Task ; rdfs:range xsd:dateTime . + +# Datatype properties +nexa:status a rdf:Property ; rdfs:range xsd:string . # "needs-action" | "in-progress" | "done" +nexa:context a rdf:Property ; rdfs:range xsd:string . # "work" | "personal" +nexa:urgency a rdf:Property ; rdfs:range xsd:integer . # 1–5 +nexa:contentHash a rdf:Property ; rdfs:range xsd:string . # for de-dup + +# Cross-pillar / multimodality +nexa:modality a rdf:Property ; rdfs:range xsd:string . # "text" | "image" +nexa:mediaUri a rdf:Property ; rdfs:range xsd:anyURI . # memos://… , nextcloud://… , obsidian://… +nexa:vectorCollection a rdf:Property ; rdfs:range xsd:string . # "nexa_knowledge_text" | "nexa_knowledge_visual" +nexa:vectorId a rdf:Property ; rdfs:range xsd:string . # Qdrant point ID +nexa:pendingVisualIndex a rdf:Property ; rdfs:range xsd:boolean . # set true on image notes until Phase 3.2 backfills them +``` + +`nexa:vectorId` + `nexa:vectorCollection` together are the **bridge** between graph and vector store. A SPARQL hit can trigger a vector lookup, and a Qdrant payload's `graph_iri` field walks back the other way. + +`nexa:modality`, `nexa:mediaUri` and `nexa:pendingVisualIndex` exist from Phase 3.1 even though only the text path is wired up. Image attachments captured in 3.1 are recorded as `nexa:Note` with `modality "image"` and `pendingVisualIndex true`, then picked up by the Phase-3.2 backfill workflow — no data loss across phases. + +--- + +## Sync flows + +### 1. Memos → GraphDB (real-time) + +``` +Memo content: "Muss JWT-Middleware für Auth-Service refaktorieren" + │ + ▼ SAIA extracts entities + relations as JSON + │ { tasks: [{title, urgency}], technologies: [...], + │ relations: [{type:"uses", from:..., to:...}] } + │ + ▼ n8n turns JSON into a SPARQL UPDATE + │ + └──▶ INSERT DATA { ... } against GraphDB repo "nexa_knowledge" +``` + +### 2. Obsidian → GraphDB (`#nexa:sync-obsidian`) + +For each Obsidian note: parse front-matter + headings → emit `nexa:Project`, `nexa:Task`, `nexa:Note` triples; `nexa:mentions` for `[[wikilinks]]`. + +### 3. Nextcloud Tasks ↔ GraphDB (bidirectional) + +n8n trigger on Nextcloud CalDAV/Tasks change → `INSERT/DELETE DATA` to keep `nexa:status` and `nexa:scheduledFor` in sync. + +--- + +## Example SPARQL queries + +### Q1 — All open tasks involving JWT, by urgency + +```sparql +PREFIX nexa: + +SELECT ?taskTitle ?urgency ?projectName +WHERE { + ?tech rdfs:label "JWT" . + ?task nexa:uses ?tech ; + rdfs:label ?taskTitle ; + nexa:status ?status ; + nexa:urgency ?urgency . + FILTER (?status IN ("needs-action", "in-progress")) + OPTIONAL { ?task nexa:childOf ?project . ?project rdfs:label ?projectName . } +} +ORDER BY DESC(?urgency) +``` + +### Q2 — What does Auth-Service transitively depend on? + +```sparql +PREFIX nexa: + +SELECT DISTINCT ?dep ?label +WHERE { + ?root rdfs:label "Auth-Service" . + ?root nexa:dependsOn+ ?dep . + ?dep rdfs:label ?label . +} +``` + +(`+` is SPARQL property-paths — transitive closure, free.) + +### Q3 — Topics with the most note-mentions in the last day + +```sparql +PREFIX nexa: +PREFIX xsd: + +SELECT ?topic (COUNT(?note) AS ?n) +WHERE { + ?note a nexa:Note ; + prov:generatedAtTime ?ts ; + nexa:mentions ?topic . + FILTER (?ts > NOW() - "P1D"^^xsd:duration) +} +GROUP BY ?topic +ORDER BY DESC(?n) +LIMIT 10 +``` + +### Q4 — Cross-pillar: "find vectors for tasks blocking project X" + +```sparql +PREFIX nexa: + +SELECT ?taskTitle ?vectorId +WHERE { + ?proj rdfs:label "Nexa" . + ?task nexa:childOf ?proj ; + nexa:status "in-progress" ; + nexa:vectorId ?vectorId ; + rdfs:label ?taskTitle . +} +``` + +n8n then takes each `?vectorId`, fetches the embedding from Qdrant, and runs a "more like this" search for richer context. + +--- + +## GraphRAG answer pipeline (`#nexa:ask`) + +``` + #nexa:ask + │ + ┌──────┴──────┐ + ▼ ▼ + [Qdrant] [GraphDB] + semantic structural + top-k SPARQL — auto-generated + notes paths / dependencies + │ │ + └──────┬──────┘ + ▼ + merge + rank + │ + ▼ + SAIA prompt: + "Given these passages and these relations, answer …" + │ + ▼ + comment under the original memo +``` + +Auto-generation of SPARQL: SAIA is given the ontology (above) as a system prompt and asked to emit a `SELECT`/`CONSTRUCT` query for the user's natural-language question. n8n executes it, falls back to a templated query on parse failure. + +--- + +## n8n integration sketch + +### Workflow: Graph-Sync Trigger + +``` +[Memos Webhook] + │ +[Parse Content] + │ +[SAIA: Extract entities + relations as JSON] + │ +[Build SPARQL UPDATE INSERT DATA { ... }] + │ +[HTTP POST → /repositories/nexa_knowledge/statements] + │ +[Index in Qdrant; write Qdrant point id back via second SPARQL UPDATE] +``` + +### Workflow: Question Router + +``` +[#nexa:ask Query] + │ + ┌─┴────────────────┐ + ▼ ▼ +[SAIA: NL → SPARQL] [Qdrant: kNN] + │ │ +[POST → SPARQL endpoint] + │ │ + └──────┬───────────┘ + ▼ + rank + merge → SAIA answer +``` + +--- + +## Graph-management commands + +### `#nexa:graph-status` + +Returns triple count, class histogram, most-connected entity. Implemented as one SPARQL `SELECT (COUNT)`. + +### `#nexa:graph-trace [entity]` + +Returns the 1-hop (and optionally 2-hop) neighbourhood — a `DESCRIBE ` plus a templated outgoing/incoming query. + +### `#nexa:graph-rebuild` + +Clears the named graph and replays Obsidian + Memos. SPARQL: `CLEAR GRAPH ` followed by the import workflow. + +--- + +## Why two stores + +| Scenario | Qdrant | GraphDB | Best | +|----------|--------|---------|------| +| "Which note was similar to this one?" | ✅ | ❌ | Qdrant | +| "What blocks this task?" | ❌ | ✅ | GraphDB | +| "Explain this project" | ✅ context | ✅ structure | both | +| "All JWT-related open work" | ✅ semantic | ✅ crisp | both | + +Combined: **complete understanding** rather than a search index *or* a structure index. + +--- + +## Memory sources & retention + +Not every embedding deserves to live forever. Nexa indexes from several source types and each has its own expected lifetime. The contract: every Qdrant point carries `payload.source_type` and `payload.expires_at` (epoch seconds, or `null` for permanent). A daily prune workflow runs `DELETE WHERE expires_at < NOW()` on each collection and mirrors the deletion in GraphDB. + +| `source_type` | Where it comes from | Default TTL | Rationale | +|---------------|--------------------|-------------|-----------| +| `memo` | Memos webhook | **permanent** | User-authored, low volume, high signal. | +| `obsidian` | Nextcloud `Notizen/` via WebDAV | **permanent** | User-authored knowledge base. | +| `mail` | Nextcloud Mail (single account) | 365 d | Audit trail + searchable past correspondence. Mail digests are derived, not stored as their own embeddings. | +| `mail_digest` | Daily digest output | 90 d | Summarised content; the source mails persist longer. | +| `karakeep` | Karakeep saved links | **permanent** | User explicitly bookmarked. | +| `rss` | Phase-2.2 morning digest feed items | 30 d | News signal decays fast; keep recent for "what was that article last week?". | +| `web_search` | On-demand fetch via crawl4ai-mcp / markitdown-mcp during `#nexa:ask` | 90 d | Useful for "what did we look at last quarter?" but not eternal. | +| `system` | Backrest / Proxmox / n8n alerts via `nexa.system` ntfy topic | 30 d | Operational telemetry; old alerts have little RAG value. | +| `task` | Nextcloud Tasks ↔ GraphDB sync | until task deleted | Mirrors source-of-truth. | + +**External sources go through the same pipeline as memos** — fetch → markitdown-mcp → embed via TEI → upsert into `nexa_knowledge_text` with the appropriate `source_type` + `expires_at`. The graph node carries `nexa:source`, `nexa:fetchedAt`, `nexa:sourceUri`, and `nexa:contentHash` for de-dup (so the same article fetched twice doesn't create two points). + +### Web search loop + +`#nexa:ask` first searches existing memory. If the merged confidence is below a threshold (or the user adds `--web` to the command), Nexa runs a SearXNG query through `redis-searxng`, picks the top 3 results, fetches them through `crawl4ai-mcp` + `markitdown-mcp`, embeds the cleaned markdown, and **answers from the augmented context**. The fetched pages stay in memory (TTL 90 d) so the next related question doesn't re-fetch. + +This means the homelab's existing `*-mcp` containers are part of Nexa's data plane, not just decoration — see [docs/12 #8](./12-optimization-opportunities.md#8). + +### Manual overrides + +- `#nexa:learn --permanent` overrides the default TTL. +- `#nexa:forget ` triggers an immediate Qdrant delete + GraphDB `DELETE WHERE { ?n nexa:vectorId "..." . }`. +- `#nexa:retain ` rewrites the default for that source type (stored in the `_config` namespace, picked up by the next prune run). diff --git a/stacks/nexa/docs/09-deployment.md b/stacks/nexa/docs/09-deployment.md new file mode 100644 index 0000000..16c865d --- /dev/null +++ b/stacks/nexa/docs/09-deployment.md @@ -0,0 +1,326 @@ +# 09 — Deployment + +Pragmatic deployment guide that **assumes the existing homelab** and adds only what's missing. + +## What's already running (no action required) + +Surveyed from Homepage / Dozzle / Proxmox / Zoraxy: + +| Service | Host / port | URL | +|---------|-------------|-----| +| Memos | docker LXC 104 → `:5230` | `https://memos.nuclide.systems` | +| n8n | docker LXC 104 → `:5678` | `https://n8n.nuclide.systems` | +| LiteLLM (SAIA gateway) | docker LXC 104 → `:4000` | `https://ai.nuclide.systems` (proxies LobeHub UI :3210; API on :4000) | +| Nextcloud | LXC 105 | `https://nc.nuclide.systems` | +| ntfy | docker LXC 104 → `:7998` | `https://ntfy.nuclide.systems` | +| Karakeep | docker LXC 104 → `:3090` | `https://hoarder.nuclide.systems` (legacy host alias kept for compatibility) | +| Home Assistant | VM 100 (HAOS) | `https://ha.nuclide.systems` | +| Pocket-ID (OAuth/SSO) | docker LXC 104 → `:1411` | `https://id.nuclide.systems` | +| Vaultwarden | docker LXC 104 → `:11001` | `https://vault.nuclide.systems` | +| Backrest | LXC 103 | (internal) | +| AdGuard DNS | LXC 102 | (internal) | +| Zoraxy reverse proxy | LXC 108 → `192.168.1.4:8000` | TLS for *.nuclide.systems | +| `qdrant_scientific` (existing) | docker LXC 104 | **reused** — Nexa uses `nexa_*` collections in this instance | + +The deployment task is **not** "spin up the stack" — most of the stack is already up. It is **wire Nexa across these services + add the small bits that are missing**. + +## What's missing for Nexa + +1. **Qdrant collection** for Nexa (`nexa_knowledge`) inside the existing `qdrant_scientific` instance — vector dim follows Q15 (1024 for `bge-m3`, 768 for `nomic-embed-text`). +2. **TEI** (HF text-embeddings-inference) on the docker host for self-hosted embeddings (LiteLLM key is **not** authorised for OpenAI embeddings — see [11/Q3+Q15](./11-open-questions.md)). Lighter than Ollama: single Rust binary, ~500 MB image, no LLM runtime. +3. **n8n workflows** (`./nexa-core/n8n-workflows/`) imported into the running n8n. +4. **Nextcloud lists & calendars** for Work / Personal / Shopping / Wishes (auto-discovered via `#nexa:config`). +5. **Memos webhook → n8n** wired through the Memos config. +6. **LiteLLM virtual key** for the `nexa` user with chat-only access (no embeddings — handled by Ollama). +7. **A Zoraxy host entry** is *not* needed — Memos / n8n / LiteLLM are already proxied. +8. **(Phase 3.4) Ontotext GraphDB** for the SPARQL pillar — see add-on at the bottom of this doc. + +--- + +## Step 1 — Secrets + +Copy `nexa-core/.env.example` → `nexa-core/.env` and fill **only** the secrets: + +```bash +cd nexa-core +cp .env.example .env +$EDITOR .env # MEMOS_API_KEY, SAIA_API_KEY, NC_APP_PASSWORD, QDRANT_API_KEY +``` + +The `.env` is **only used at bootstrap time**. Everything else (list IDs, calendar IDs, collection sizes) is discovered at runtime via `#nexa:config` (see [05](./05-command-system.md)). No secrets should ever live in n8n workflow JSON — use n8n credentials instead. + +## Step 2 — Qdrant collection (`nexa_knowledge_text`) + +Phase 3.1 ships **Path A** (text-only) but the schema and naming already make room for **Path C** (text + visual) so adding a `nexa_knowledge_visual` collection later is a pure additive operation — no rename, no migration, no n8n rewiring. + +```bash +# adjust QDRANT_HOST in .env first +source nexa-core/.env + +# create the text collection from the schema file +curl -X PUT "$QDRANT_HOST/collections/nexa_knowledge_text" \ + -H "Content-Type: application/json" \ + -H "api-key: $QDRANT_API_KEY" \ + -d @nexa-core/config/qdrant_schema.json +``` + +The collection name is **always suffixed with the modality** (`_text`, `_visual`) so logic in n8n and SPARQL stays modality-aware from day one. Indexed rows carry these payload fields ([source](../nexa-core/config/qdrant_schema.json)): + +| Field | Why it's there now | +|-------|--------------------| +| `modality` | Always `"text"` in `_text`, `"image"` in `_visual`. Future-proofs cross-modality filters. | +| `source_type` | `memo` / `mail` / `obsidian` / `screenshot` / `image` — used by classification and digest workflows. | +| `media_uri` | `memos://…`, `nextcloud://…`, `obsidian://…`. Empty for text-only rows; populated when Path C ships. | +| `graph_iri` | IRI of the corresponding `nexa:Note` in GraphDB. The same value is stored on the GraphDB side as `nexa:vectorId` — this is the cross-pillar bridge. | +| `content_hash` | de-dup. | +| `context` | `work` / `personal`. | + +> Targets the existing `qdrant_scientific` instance — just an extra collection, no new container. +> The `vectors.size` field follows [Q15](./11-open-questions.md): **1024** for `bge-m3`, **768** for `nomic-embed-text-v1.5`. + +### Image attachments today (queue them) + +Memos can already attach images. Until Phase 3.2 the indexer **does not** embed them, but it **does** record them so they can be replayed later: + +- Memo with an image → text body still goes into `nexa_knowledge_text`. +- The image attachment(s) are written as `nexa:Note` triples in GraphDB with `nexa:modality "image"` and `nexa:vectorId` left empty (`nexa:pendingVisualIndex true`). +- A Phase-3.2 backfill workflow will pick up everything where `?n nexa:pendingVisualIndex true` and embed it through the visual collection. + +This means **no data is lost** between 3.1 and 3.2 — the queue is the GraphDB itself. + +## Step 3 — Self-hosted embeddings (TEI) + +Use HuggingFace **text-embeddings-inference** — single Rust binary, ~500 MB image, OpenAI-compatible API, loads exactly one model. Lighter than Ollama because there's no LLM runtime, no GGUF loader, no model registry. + +The active docker manager on this LXC is **Arcane** (visible from Homepage as the running container manager — the LXC was originally provisioned with the Dockge helper-script template, but Dockge is now stale; see [12/#33](./12-optimization-opportunities.md#33)). Paste the stack into Arcane → name it `nexa` → save → start. Don't `docker compose up -d` over SSH; Arcane manages the compose lifecycle. + +```yaml +# Nexa stack — paste into Arcane. +# Storage convention matches the rest of the homelab (verified against +# the running Karakeep stack, Q19): host bind-mount of +# /mnt/pve/unas/services//. No volume-driver, no CIFS, no +# credentials in the compose — the LXC's NFS mount is already there. +services: + nexa-embed: + image: ghcr.io/huggingface/text-embeddings-inference:cpu-1.5 + container_name: nexa-embed + restart: unless-stopped + command: ["--model-id", "BAAI/bge-m3"] + ports: + - "127.0.0.1:8080:80" + volumes: + - /mnt/pve/unas/services/nexa/tei-cache:/data + env_file: + - .env + +networks: {} +``` + +Pre-deploy step on the docker LXC (one-time): + +```bash +mkdir -p /mnt/pve/unas/services/nexa/{tei-cache,qdrant,graphdb} +mkdir -p /mnt/pve/unas/backup/nexa/snapshots/{qdrant,graphdb} +``` + +Memory budget: ~1.1 GB resident. First start downloads `bge-m3` (~1 GB) into `/mnt/pve/unas/services/nexa/tei-cache/`; subsequent restarts are instant. + +Secrets (`SAIA_API_KEY`, `MEMOS_API_KEY`, `QDRANT_API_KEY`, `NC_APP_PASSWORD`) go in the stack's `.env` next to the compose — same pattern Karakeep uses (`env_file: .env`). Arcane has an editor for it. Vaultwarden becomes the source-of-truth long-term ([12/#11](./12-optimization-opportunities.md#11)) but isn't required for the first cut. + +> **Why bind-mount and not SMB?** Earlier drafts of this doc proposed an SMB-via-docker-volume pattern because of the user's "had it with Nextcloud" experience. The Karakeep stack confirms the actual convention is the simpler one: **host bind-mount of the LXC's existing `/mnt/pve/unas` NFS mount**. The Nextcloud failure was Nextcloud-specific (its setup tooling chowns the data dir to `www-data`, which fails against `root_squash` exports) and doesn't apply to normal containers. SMB-as-docker-volume stays documented in [12/#27](./12-optimization-opportunities.md#27) only as an escape hatch *if* a future service hits Nextcloud-style issues — Nexa doesn't, so we don't use it. + +Register it inside LiteLLM (admin UI → Models) with the OpenAI-compatible adapter: + +- model name: `nexa-embed` +- provider: `openai` +- model: `bge-m3` +- api_base: `http://nexa-embed:80/v1` +- api_key: any non-empty string (TEI ignores it) + +Now n8n only ever talks to LiteLLM and the model is swappable without touching workflows. + +## Step 4 — LiteLLM virtual key + +In the LiteLLM admin UI (`ai.nuclide.systems`): + +1. Create user `nexa`. +2. Issue a virtual key with access to: + - one chat model (already-available model from your SAIA gateway). + - the `nexa-embed` model from Step 3. +3. Paste the key into `SAIA_API_KEY` in `.env`. + +## Step 4 — n8n workflows + +Import the JSON exports — credentials are filled inside n8n, not in the JSON: + +```bash +# n8n personal access token from the n8n UI: Settings → API +N8N_URL=https://n8n.nuclide.systems +N8N_TOKEN=... # from the n8n UI + +for f in nexa-core/n8n-workflows/phase-1/*.json \ + nexa-core/n8n-workflows/phase-2/*.json; do + curl -X POST "$N8N_URL/api/v1/workflows" \ + -H "X-N8N-API-KEY: $N8N_TOKEN" \ + -H "Content-Type: application/json" \ + --data-binary "@$f" +done +``` + +Inside n8n, attach credentials to the imported nodes: + +- **Memos** → HTTP header `Authorization: Bearer $MEMOS_API_KEY` +- **LiteLLM** → header `Authorization: Bearer $SAIA_API_KEY` (chat + `nexa-embed`) +- **Nextcloud Tasks / Calendar / WebDAV (Obsidian vault under `Notizen/`)** → one app password (`$NC_APP_PASSWORD`), reused across all three node types +- **Qdrant** → header `api-key: $QDRANT_API_KEY` + +Activate each workflow individually after smoke-test. + +## Step 5 — Memos webhook + +In the Memos admin UI, set the webhook URL to the production address of the discovery workflow: + +``` +https://n8n.nuclide.systems/webhook/memos +``` + +The same URL is the one the workflow exposes; verify with: + +```bash +curl -i https://n8n.nuclide.systems/webhook/memos +# expect 200 / 405, never 404 +``` + +## Step 6 — Bootstrap commands via Memos + +Create a memo with body `#nexa:config` — the discovery workflow: + +1. Lists Nextcloud Tasks lists, looking for `Persönlich`, `DLR`, `Einkaufsliste`, `Wunschliste` (and any new lists added later — see [05](./05-command-system.md)). Caches `name → id`. +2. Lists Nextcloud Calendars by the same names; caches IDs. +3. Verifies the Nextcloud Mail account `fkrebs@nucli.de` and the folder list (`Posteingang`, `Archiv`, `Junk`, `Waiting`). +4. Counts existing Qdrant points in `nexa_knowledge_text`. +5. Verifies LiteLLM reachability + lists available models for the Nexa virtual key. +6. Replies as a comment with a **runtime-config snapshot** that's stored in the Qdrant `_config` namespace (and mirrored as `nexa-core/config/runtime_config.json`, gitignored). + +The same workflow is also triggered by: +- A daily cron inside n8n (so newly added Nextcloud lists become routable without intervention). +- Cache-miss in the task-router — if a list ID 404s, the router fires `#nexa:config` once and retries. + +After this point, `.env` is read-once. Subsequent runs read config from Qdrant. + +## Step 7 — Smoke tests + +```bash +# (1) Memos round-trip — should produce a comment within ~5 s +curl -X POST https://memos.nuclide.systems/api/v1/memos \ + -H "Authorization: Bearer $MEMOS_API_KEY" \ + -H "Content-Type: application/json" \ + -d '{"content":"- [ ] testing the router #nexa"}' + +# (2) Classification dry-run +curl -X POST https://memos.nuclide.systems/api/v1/memos \ + -H "Authorization: Bearer $MEMOS_API_KEY" \ + -d '{"content":"#nexa:route-test buy milk"}' + +# (3) RAG test (requires at least one indexed memo/note) +curl -X POST https://memos.nuclide.systems/api/v1/memos \ + -H "Authorization: Bearer $MEMOS_API_KEY" \ + -d '{"content":"#nexa:ask what is the goal of nexa?"}' +``` + +## Step 8 — Reverse proxy + +Already done — Zoraxy at `192.168.1.4:8000` terminates TLS for `*.nuclide.systems` and forwards to docker LXC 104 (`192.168.1.40`). **No new entry is required for Nexa**: every service Nexa talks to already has a host entry. + +## Step 9 — Backups + +Already covered by Backrest (LXC 103). Add: + +- **n8n workflows** → `nexa-core/scripts/backup_workflows.sh` (already present) into a Backrest schedule. +- **Qdrant snapshots** → schedule a daily `POST /collections/nexa_knowledge/snapshots` and rsync to S3 (`s3.nuclide.systems`). Add as a Backrest pre-hook on the docker host. + +For deeper detail: [10 — Operations](./10-operations.md). + +--- + +## Phase add-on: Ontotext GraphDB (Phase 3.4) + +Defer until 3.1–3.3 ship. + +```yaml +# nexa-core/docker-compose.graph.yml +services: + graphdb: + image: ontotext/graphdb:10.7.0 + container_name: nexa-graphdb + ports: ["127.0.0.1:7200:7200"] + environment: + GDB_JAVA_OPTS: "-Xmx4g -Xms1g" + volumes: + - ./data/graphdb:/opt/graphdb/home + restart: unless-stopped +``` + +After first start, create the repository (one-time): + +```bash +curl -X POST http://localhost:7200/rest/repositories \ + -H 'Content-Type: application/json' \ + -d '{ + "id": "nexa_knowledge", + "title": "Nexa Knowledge Graph", + "type": "graphdb", + "params": { + "ruleset": {"value": "rdfsplus-optimized"}, + "baseURL": {"value": "https://nuclide.systems/nexa/"} + } + }' +``` + +Optional Zoraxy entry `graph.nuclide.systems` → `192.168.1.40:7200` if you want the SPARQL Workbench in a browser; otherwise n8n talks to it on the docker network at `http://nexa-graphdb:7200`. + +For schema and example queries: [08-graphrag-architecture](./08-graphrag-architecture.md). + +--- + +## Phase add-on: visual collection (Phase 3.2) + +Adds Path C — image embeddings without disturbing the text path. Schema is already in `nexa-core/config/qdrant_schema_visual.json`. + +```bash +# (1) replace TEI with infinity (or run alongside) for CLIP-family support +docker rm -f nexa-embed +docker run -d --name nexa-embed \ + --restart unless-stopped \ + -p 127.0.0.1:8080:80 \ + -v infinity-data:/app/.cache \ + michaelf34/infinity:latest \ + v2 \ + --model-id BAAI/bge-m3 \ + --model-id jinaai/jina-clip-v2 \ + --port 80 + +# (2) create the visual collection +curl -X PUT "$QDRANT_HOST/collections/nexa_knowledge_visual" \ + -H "Content-Type: application/json" \ + -H "api-key: $QDRANT_API_KEY" \ + -d @nexa-core/config/qdrant_schema_visual.json + +# (3) register the second model in LiteLLM as `nexa-embed-visual` +# (same OpenAI-compatible route, different model id) + +# (4) backfill queued images: +# SPARQL: SELECT ?note ?uri WHERE { ?note nexa:pendingVisualIndex true ; nexa:mediaUri ?uri } +# For each row: fetch the bytes, embed via nexa-embed-visual, upsert into the visual collection, +# UPDATE GraphDB to set nexa:vectorId and DELETE nexa:pendingVisualIndex. +``` + +n8n RAG workflow gains a parallel branch: text-query → both `nexa-embed-text` and `nexa-embed-visual` text encoders → kNN against both collections → merge by score before SAIA prompt. + +--- + +## Step-back / rollback + +- **Disable any Nexa workflow** in n8n — deactivates the side effect immediately, Memos webhooks become no-ops. +- **Drop a Qdrant collection** — `curl -X DELETE $QDRANT_HOST/collections/nexa_knowledge_text -H "api-key: $QDRANT_API_KEY"`. +- **Re-discover** — `#nexa:reset-config` then `#nexa:config`. diff --git a/stacks/nexa/docs/10-operations.md b/stacks/nexa/docs/10-operations.md new file mode 100644 index 0000000..a58dcc5 --- /dev/null +++ b/stacks/nexa/docs/10-operations.md @@ -0,0 +1,101 @@ +# 10 — Operations + +Day-2 concerns. Backups, monitoring, troubleshooting. + +## Backups + +**Tiering** (3-2-1, full design in [12 #37](./12-optimization-opportunities.md#37)): +1. **Source** — UNAS RAID 6 + native UniFi Drive snapshots ([12 #38](./12-optimization-opportunities.md#38), still to enable). +2. **Warm tier** — `s3.nuclide.systems` (on-site), restic/borg via Backrest (LXC 103). +3. **Cold tier off-site** — encrypted `rclone` copy to a third-party (Jottacloud is the user's candidate; decision tracked in [Q20](./11-open-questions.md#q20)). Always encrypt before upload — provider sees only ciphertext. + +Nexa-specific items inside this pipeline: + +| What | How | Frequency | +|------|-----|-----------| +| n8n workflow JSON | `nexa-core/scripts/backup_workflows.sh` → git push | hourly cron in n8n container | +| Qdrant `nexa_knowledge_text` (and `_visual` once Phase 3.2) | `POST /collections//snapshots` → write to `backup/nexa/snapshots/qdrant//` on UNAS, then Backrest picks it up for warm + cold | daily (snapshot), weekly (cold tier) | +| GraphDB `nexa_knowledge` (Phase 3.4+) | scheduled SPARQL `CONSTRUCT` export → `backup/nexa/snapshots/graphdb/.ttl.gz` | daily | +| Memos DB | Backrest snapshot of Memos data dir | already covered | +| Nextcloud | Nextcloud's own backup app + Backrest of `services/nextcloud/` | already covered | +| `runtime_config` (Qdrant `_config` namespace) | Qdrant snapshot covers it | n/a | + +## Monitoring + +| Signal | Source | Sink | +|--------|--------|------| +| Container up/down | Dozzle (`192.168.1.40:3553`) + docker healthchecks. **Skip names matching `octoprint*`** — that container is intentionally powered down most of the time. | Memos system feed via ntfy | +| n8n workflow failures | n8n built-in failure-webhook | ntfy → `nexa.system` topic | +| Proxmox / disk / memory alerts | Proxmox notification target → ntfy | ntfy → Memos digest | +| LiteLLM rate-limit | LiteLLM logs / cost-tracking | Memos morning digest | +| Backrest run status | Backrest webhook | ntfy | + +A **single ntfy topic `nexa.system`** is the convention; n8n has one workflow that re-broadcasts it as a Memos comment under a pinned `[SYSTEM]` memo. + +## Troubleshooting + +### Memos webhook not firing + +```bash +# is the Memos webhook config still pointing at n8n? +curl -s -H "Authorization: Bearer $MEMOS_API_KEY" \ + https://memos.nuclide.systems/api/v1/workspace/setting | jq '.webhooks' + +# does n8n still expose the path? +curl -i https://n8n.nuclide.systems/webhook/memos # expect 200/405, never 404 +``` + +If 404 → workflow is inactive in n8n. Activate. + +### LiteLLM 401 / 429 + +- 401 → the Nexa virtual key was rotated. Reissue, update `.env`, re-create the n8n credential. +- 429 → set per-key TPM/RPM limits in LiteLLM admin; classification calls are tiny (<300 tok), embeddings are the bulk. + +### Qdrant `nexa_knowledge` empty + +```bash +# is the indexing workflow active? +docker logs nexa-n8n 2>&1 | grep memos_bridge | tail + +# manual index test +curl -X POST https://memos.nuclide.systems/api/v1/memos \ + -H "Authorization: Bearer $MEMOS_API_KEY" \ + -d '{"content":"manual probe #nexa"}' + +# point count +curl -s -H "api-key: $QDRANT_API_KEY" \ + $QDRANT_HOST/collections/nexa_knowledge | jq '.result.points_count' +``` + +### Wrong list routing (Work vs Personal) + +1. Run `#nexa:route-test ` and check the confidence value. +2. If <0.7 the router defaults to Personal (by design — see [06](./06-classification-logic.md)). +3. Tune the system prompt in `nexa-core/ai-prompts/system_prime.txt`; commit the change. + +### "Where did Nexa store this?" + +```bash +# state dump +curl -X POST https://memos.nuclide.systems/api/v1/memos \ + -d '{"content":"#nexa:export-state"}' +# returns runtime_config + counters as a JSON memo +``` + +## Log locations + +| Component | Where | +|-----------|-------| +| Memos | `docker logs nexa-memos` (LXC 104) | +| n8n | `docker logs nexa-n8n` | +| LiteLLM | `docker logs litellm` | +| Qdrant | `docker logs qdrant_scientific` | +| Zoraxy | LXC 108 web UI → Statistical Analysis | +| Backrest | LXC 103 web UI | + +For a one-shot dump: + +```bash +ssh nuc 'docker compose -p nexa logs --tail 500' > /tmp/nexa.log +``` diff --git a/stacks/nexa/docs/11-open-questions.md b/stacks/nexa/docs/11-open-questions.md new file mode 100644 index 0000000..96a092d --- /dev/null +++ b/stacks/nexa/docs/11-open-questions.md @@ -0,0 +1,73 @@ +# 11 — Open Questions (user-info-required) + +Items that **block progress** and need a human decision before a workflow can be implemented or a service deployed. Tick them off as you decide. + +## Resolved + +- [x] **Q1 — Graph DB choice → Ontotext GraphDB (SPARQL).** Rationale: explore Nexa's memory through SPARQL is a stated goal. [08-graphrag-architecture](./08-graphrag-architecture.md) is rewritten accordingly. +- [x] **Q2 — Vector store → reuse `qdrant_scientific`** with a `nexa_*` collection prefix. No dedicated container. +- [x] **Q3 — Embeddings model → not OpenAI.** Self-host on the docker host via **TEI** (HF `text-embeddings-inference`) — Rust single-binary, OpenAI-compatible, ~500 MB image, no LLM runtime overhead. Speed analysis in §"Speed budget" below. +- [x] **Q5 — Display name → Karakeep.** The Zoraxy host alias `hoarder.nuclide.systems` is legacy — keep it for compatibility, but all docs, prompts and new workflow nodes use "Karakeep". +- [x] **Q13 — Octoprint container is *intentionally* temporary.** Suppress from Phase-5 monitoring: container-up/down alerts must skip names matching `octoprint*` (or any container tagged `proxmox-he 3d-printing`). +- [x] **Q14 — Homepage Zoraxy widget → no such widget.** Config error in `homepage/services.yaml`; cosmetic, not Nexa-related. +- [x] **Q15 — Embeddings staging plan → A now, C prepared.** + - **Phase 3.1 (now)**: TEI + `BAAI/bge-m3`, single collection `nexa_knowledge_text` (1024-dim). DE/EN multilingual, fits the corpus. + - **Phase 3.2 (later)**: swap TEI → `infinity`, add `jinaai/jina-clip-v2` (768-dim), second collection `nexa_knowledge_visual`. Backfill from the queue (see Q16). + - All schema fields needed for 3.2 (`modality`, `media_uri`, `graph_iri`, `nexa:pendingVisualIndex`) are introduced **now** so 3.2 is purely additive — no rename, no migration. See [`qdrant_schema.json`](../nexa-core/config/qdrant_schema.json) and [`qdrant_schema_visual.json`](../nexa-core/config/qdrant_schema_visual.json). +- [x] **Q16 — Image-attachment queue ergonomics → leave bytes at source, reference by `media_uri`.** Zero copy. Memos attachments stay in Memos's data dir, Nextcloud images stay in Nextcloud, Obsidian images stay in the `Notizen` folder; the Phase-3.2 backfill workflow fetches them on demand via the URI scheme. +- [x] **Q17 — Immich out-of-band.** Nexa does not call the Immich smart-search API. Photo-library queries stay inside Immich. +- [x] **Q18 — SAIA proxies an embedding model, but 10 msg/min rate limit makes it unusable for ingest.** A 2 k-note Obsidian backfill would take ~3.3 h; real-time `#nexa:ask` would queue for tens of seconds during a writing burst. **Decision: deploy TEI as planned.** SAIA embeddings remain available as a manual fallback (e.g. for one-off `#nexa:learn` calls where rate is irrelevant). +- [x] **Q19 — Storage convention is plain host bind-mount of `/mnt/pve/unas/services//`.** Karakeep stack confirmed: no driver opts, no CIFS, no per-volume credentials. Secrets via `env_file: .env` next to the compose. The earlier SMB-as-docker-volume proposal in this repo is withdrawn ([12/#27](./12-optimization-opportunities.md#27)) — it was over-fitting to the Nextcloud-specific NFS issue (Nextcloud's setup tooling chowns to `www-data` and fails on `root_squash`-style exports; that doesn't apply to normal containers). Nexa stack updated in [docs/09 §Step 3](./09-deployment.md#step-3--self-hosted-embeddings-tei). +- [x] **Q4 — Obsidian sync via Nextcloud WebDAV.** Confirmed the vault is `nc.nuclide.systems/Notizen/` (multi-device sync via Nextcloud client). Nexa accesses it through WebDAV (`/remote.php/dav/files//Notizen/`) reusing the existing `NC_APP_PASSWORD` — no filesystem mount, no LXC-to-LXC privilege escalation. Phase 3.1 polls every 15 min; an upgrade to Nextcloud's `notify_push` for sub-second updates is captured as optimization #15. Ignore list (don't index): + - `.copilot/`, `.copilot-index/` — Obsidian Copilot's own embeddings cache. + - `.smart-env/` — Smart Connections / Smart Composer plugin data (~13 MB). + - `.caldav-sync/` — calendar sync, not notes. + - `assets/` — 186 MB of binaries; routed through the Phase-3.2 visual queue (`nexa:pendingVisualIndex`), not the text path. + - `Templates/` — empty templates, low semantic value. + - `BMO/`, `Excalidraw/` — plugin folders. + Anything else under `Notizen/**/*.md` is fair game. +- [x] **Q6 + Q7 — Lists & calendars discovered by name, not ID. Self-healing.** + Confirmed names from screenshots: + - **Tasks**: `Persönlich` (14), `DLR` (13), `Einkaufsliste` (13), `Wunschliste` (9). + - **Calendars**: `Persönlich`, `DLR`, `Einkaufsliste`, `Wunschliste` (Nextcloud Tasks is calendar-backed, so the names are shared; Tasks lives on the calendar of the same name). + - Mapping in classification: `DLR` = Work-Kontext, `Persönlich` = Personal-Kontext, `Einkaufsliste` = Shopping, `Wunschliste` = Wishes. + + **Self-healing requirement (the user explicitly noted lists may change/grow):** Nexa must not cache IDs forever. The `#nexa:config` workflow runs (a) on demand, (b) once daily as a scheduled refresh, and (c) automatically as a retry whenever a list/calendar lookup returns 404 or "not found". The runtime-config record in Qdrant's `_config` namespace stores `{name → id, discovered_at}` and gets invalidated on cache-miss. New lists added in Nextcloud surface in the next scheduled refresh and Nexa starts honouring `#einkaufsliste` / `#dlr` etc. without code changes. + +- [x] **Q8 — Mail via Nextcloud, single account `fkrebs@nucli.de`.** No extra IMAP entry. n8n's IMAP node uses the same server credentials Nextcloud Mail already holds for that account; Nexa never sees a second password. Folders confirmed: `Posteingang` (default), `Archiv`, `Junk` (61 — auto-filtered, ignored by Nexa), `Papierkorb`, `Waiting`. The `Waiting` folder is a useful manual signal — items moved there by the user are skipped from digest (treat as "in flight"). + +- [x] **Q9 — Pocket-ID SSO is configured everywhere.** No separate auth for Nexa: Memos / n8n / future Nexa dashboard authenticate humans via Pocket-ID at the Zoraxy layer. Machine-to-machine calls (n8n → Memos webhook, n8n → Nextcloud, n8n → Qdrant, n8n → LiteLLM) keep using API keys / app passwords — SSO is for human UIs only. Don't add basic-auth or per-app login screens. + + This **resolves Q10 too** (the proposed "Pocket-ID SSO in front of n8n / Memos" is already done; Nexa just inherits it). + +- [x] **Q10 — folded into Q9.** + +## Hardware / capacity + +- [x] **Q11 — Capacity confirmed from Proxmox node + LXC summaries.** + - **Host (NUC 14 Pro)**: 22 threads (Intel Core Ultra 7 155H, 1 socket), 62.12 GiB RAM, 1.64 TiB disk (0.35% used). Steady state: 32.4 GiB used (≈24 GiB of which is ZFS ARC), load avg 1.99 / 1.39 / 1.15, IO delay 0.04%. + - **LXC 104 (docker, unprivileged)**: 16 CPU, **31.25 GiB RAM cap (7.86 GiB used = 25%)**, 8 GiB swap (idle), **200 GiB boot disk at 47.7% used (~95 GiB)**. Tag `proxmox-helper-scripts`. Intel iGPU passthrough configured but currently failing — see [12/#26](./12-optimization-opportunities.md#26). + - **RAM headroom is comfortable**: Phase-3 budget (~7 GiB additional — Qdrant ~1.5 + TEI/bge-m3 ~1.1 + GraphDB 4-GB heap + later +infinity/jina-clip-v2 ~1) fits inside the existing LXC cap with ~16 GiB still free. Levers if needed: (1) raise the LXC cap (host has plenty), (2) cap `zfs_arc_max` lower (ARC currently 24 GiB). Either is one line of config. + - **Disk is fine** — the LXC also NFS-mounts the UNAS at `/mnt/pve/unas` (19.4 TiB total, 16.3 TiB free). Large/persistent volumes (Qdrant data, GraphDB repo, snapshots, image bytes if we ever stage them) bind-mount there; the 200 GiB local boot disk only carries images and small ephemeral state. See [12/#27](./12-optimization-opportunities.md#27). + - **Verdict: no blockers for Phase-3.1.** Revisit once 3.4 GraphDB lands. +## Backups (deferred) + +Backup-tier decisions are intentionally parked while Phase-3.1/3.2/3.4 are built. UNAS RAID 6 + Backrest already cover file-level recovery; pool-snapshot configuration ([12/#38](./12-optimization-opportunities.md#38)) remains the single highest-leverage data-protection action and doesn't depend on either question below. Re-open both when [Phase 3.3](./02-roadmap.md#phase-3--memory-qdrant--rag) becomes the next-up item. + +- [ ] **Q12 — *deferred.*** S3 archive bucket on `s3.nuclide.systems` (warm tier). Pick when Phase 3.3 is up next. +- [ ] **Q20 — *deferred.*** Off-site cold-tier provider (Jottacloud is the leading candidate). Pick when Phase 3.3 is up next. + +--- + +## Speed budget (Q3 follow-up) + +Workload measured against the **LXC 104 cap** (16 CPU, 31.25 GiB RAM — host has 22 threads / 62 GiB if we ever raise the cap): + +| Task | Volume | Latency target | Achievable on CPU with `bge-m3` | Achievable with `nomic-embed-text` | +|------|--------|----------------|---------------------------------|------------------------------------| +| Real-time memo embed | 1 doc | <500 ms incl. n8n round-trip | ✅ ~50–100 ms | ✅ ~20 ms | +| Daily ingest | ~70 docs | <60 s | ✅ ~5–10 s | ✅ ~2 s | +| Obsidian backfill (one-shot) | ~2 000 docs | <15 min | ✅ ~2–4 min | ✅ <1 min | +| RAG query embed (`#nexa:ask`) | 1 doc | <300 ms | ✅ ~50 ms | ✅ ~20 ms | + +Conclusion: CPU-only TEI is sufficient — no GPU needed for current scope. Bottleneck is SAIA chat (already remote), not embeddings. SAIA's own embedding endpoint is rate-limited to **10 req/min** which would block real-time embed; self-hosted TEI side-steps that completely. diff --git a/stacks/nexa/docs/12-optimization-opportunities.md b/stacks/nexa/docs/12-optimization-opportunities.md new file mode 100644 index 0000000..816f444 --- /dev/null +++ b/stacks/nexa/docs/12-optimization-opportunities.md @@ -0,0 +1,149 @@ +# 12 — Optimization Opportunities + +Observations from the running infrastructure. Each item is independent — accept, defer, or reject. + +## For Nexa directly + +1. **Reuse, don't redeploy.** The earlier `DEPLOYMENT.md` would have spun a second Memos / n8n / Qdrant. The current homelab already runs all three. The new [09-deployment](./09-deployment.md) treats these as pre-existing — keeps the config minimal and avoids port collisions. +2. **Use LiteLLM virtual keys per logical caller.** Today there's one SAIA key. Issuing one key per workflow (`nexa-router`, `nexa-embed`, `nexa-digest`) lets you set different per-key rate/cost limits and disable a single workflow without rotating everything. +3. **Use n8n's *credential* objects, never inline secrets.** The current workflows under `nexa-core/n8n-workflows/phase-1/*.json` should be reviewed — if any header `Authorization` is hardcoded, replace with credential references before importing. +4. **Centralise system alerts on a single ntfy topic** (`nexa.system`). Backrest, Proxmox notifications, n8n failure-webhook and the Octoprint Exited state all go to that topic; one Memos system memo aggregates them. +5. **Defer graph DB until Phase 3.4.** Qdrant alone covers ~80% of the assistant's daily value. The graph DB is justified once you actually need dependency analysis or critical-path queries. +6. **Auto-export n8n workflows.** `nexa-core/scripts/backup_workflows.sh` already exists. Schedule it inside the n8n container (cron) and let it `git commit && git push` — this is the cheapest disaster recovery. + +## For the wider homelab (out of scope but worth noting) + +7. **AI gateway naming.** `ai.nuclide.systems` currently proxies LobeHub (a chat UI on `:3210`), while the LiteLLM API lives on `:4000`. For Nexa, point n8n directly at LiteLLM (`http://192.168.1.40:4000` over the docker net — no public TLS hop needed) to save latency and isolate from UI restarts. +8. **MCP servers consolidation.** Dozzle shows `crawl4ai-mcp`, `markitdown-mcp`, `papersearch-mcp` running individually. They're all MCP servers — Nexa Phase-3 could pull from these via LiteLLM's MCP support to enrich the embedding pipeline (e.g. fetch + markitdown a Karakeep link before embedding). +9. **Backup the n8n SQLite file** — Backrest covers `/home/node/.n8n` if added; today the only "backup" is the workflow JSON which omits credentials and execution history. +10. **Pocket-ID SSO in front of n8n** would let you remove n8n basic-auth and unify session management across the whole stack. One-time setup, large UX win. +11. **Vaultwarden as the secret store** for Nexa secrets (`SAIA_API_KEY`, `MEMOS_API_KEY`, …) — read at bootstrap via the Bitwarden CLI from inside the docker host. Removes the need for a `.env` on disk. +12. **AdGuard as DNS-based control plane.** Since AdGuard is the resolver for the LAN, you can rewrite `*.nuclide.systems` to `192.168.1.4` (Zoraxy) internally and avoid a hairpin via the WAN — already the case if AdGuard rewrite rules are set, worth verifying. +13. **Disk usage on LXC 104** is 47.7 % (Proxmox). Monitor; n8n execution logs and Dozzle history are the usual culprits. Setting `EXECUTIONS_DATA_PRUNE=true` and `EXECUTIONS_DATA_MAX_AGE=168` (7 days) on n8n keeps it bounded. +14. **Vector-store sprawl in the homelab — three competing indexes today.** Nexa is about to be the fourth. Track for eventual consolidation: + - **Obsidian Copilot** (`.copilot`, `.copilot-index`) and **Smart Connections / Smart Composer** (`.smart-env`, ~13 MB) — embed the Obsidian vault into two separate vector stores inside the vault. + - **Paperless-AI ChromaDB** (`services/paperless-ai/chromadb/`) — embeds scanned documents from Paperless-ngx for Q&A. + - **Nexa Qdrant** (`nexa_knowledge_text`, soon `_visual`) — embeds the cross-source corpus. + + Long-term, Nexa is the natural single source of truth (it sees memos + mail + obsidian + RDF graph). Once its RAG is satisfying, retire the Obsidian-plugin indexes and consider letting Nexa **read from** Paperless-AI's ChromaDB rather than re-embed PDFs (one-line ChromaDB query, much cheaper than redoing OCR-to-vector). Track but don't act yet. + +15. **Retire `open-webui`.** Confirmed stale by the user — only LobeHub is in active use as the LiteLLM chat front-end (`ai.nuclide.systems`). Stop the container, tar `services/open-webui/` into `backup/open-webui/`, then remove the stack. Frees ~500 MB RAM + a couple of GB of model cache. + +16. **`services/siyuan/workspace/` is dead data.** SiYuan retired, content migrated to Obsidian (Nextcloud `Notizen/`). Keep a final tar in `backup/siyuan/`, then `rm -rf services/siyuan/`. Frees disk + removes a "is this still authoritative?" question for future agents (and for Nexa's classifier if it ever sees the path). +17. **Real-time Obsidian sync via `notify_push`.** Phase 3.1 polls WebDAV every 15 min (Q4 resolution). Once that works, swap to Nextcloud's `notify_push` app for sub-second propagation. One-line workflow change in n8n. +18. **`assets/` is 186 MB of binaries** in the Obsidian vault — worth a glance to confirm it's mostly images (Phase-3.2 visual queue) rather than something that should live in Nextcloud Files proper. +19. **Paperless-AI as a Phase-2.x triage helper for `_sortMe/Downloads/`.** UNAS shows ~230 PDFs/docs in `_sortMe/Downloads/` plus another batch under `_sortMe/Anne/`. Paperless-AI (already running) can ingest, OCR, classify and route them; Nexa's role is to *delegate* — wire a workflow that posts a batch to Paperless-AI and reports the result back via Memos. +20. **`media/Recipes/`** has ~300 individually-named recipe folders. Once Phase-3.1 text indexing works, this becomes a high-quality test corpus for `#nexa:ask` (e.g. "what was that gochujang noodle recipe with no anchovies?"). Out-of-scope for the deployment plan, but a satisfying first user-facing win. + +## Proxmox host (NUC 14 Pro) tuning + +Observed from the node summary: 22 threads, 62 GiB RAM (32 GiB used, ~24 GiB of which is ZFS ARC), 1.64 TiB disk (0.35% used), load avg <2.0, IO delay 0.04%, kernel 6.17.13-4-pve, PVE 9.1.9, EFI. Suggestions in priority order: + +21. **Cap ZFS ARC.** Default is 50% of RAM (~31 GiB); current actual ~24 GiB. For a node that runs services rather than a pure storage box, capping at 8–12 GiB frees ~12–16 GiB for guests without measurable IO impact (disk is 1.64 TiB and 0.35% used — there's nothing hot to cache): + ``` + echo 'options zfs zfs_arc_max=8589934592' > /etc/modprobe.d/zfs.conf # 8 GiB + update-initramfs -u + ``` + Reboot or `echo 8589934592 > /sys/module/zfs/parameters/zfs_arc_max` to apply live. +22. **Enable KSM (Kernel Same-page Merging).** With ~40 docker containers + several LXCs, KSM typically frees 1–3 GiB by deduplicating identical memory pages. Currently `KSM sharing: 0 B` in the summary. PVE has `ksmtuned` available — `systemctl enable --now ksmtuned`. +23. **Suppress the `pve-no-subscription` repository warning** — either accept it (it's a homelab) and apply the `pve-no-subscription-warning` polyfill, or move to the enterprise repo. Pure cosmetic, but the orange banner in the UI is noise. +24. **Swap is 31 GiB on a 62 GiB box with ZFS root** — almost certainly oversized. Drop `vm.swappiness` to 10 (`sysctl -w vm.swappiness=10` + persist) so swap is only used under genuine pressure, and consider shrinking the swap volume if disk-layout permits. +25. **Verify scheduled ZFS scrub is enabled.** PVE ships `zfs-scrub-monthly@.timer` — `systemctl list-timers | grep zfs` to confirm. Cheap insurance on a 1.6 TiB pool. +26. **SMART monitoring on the NVMe.** `smartctl -a /dev/nvme0` should be regularly polled; PVE's notification target can ntfy on degradation. Combine with the existing `nexa.system` ntfy topic (optimization #4) so disk-health alerts land in the same Memos system feed as everything else. +27. **NTP source via AdGuard.** AdGuard already resolves DNS for the LAN; pointing the host's `systemd-timesyncd` at `pool.ntp.org` resolved through AdGuard avoids any external dependency for time. One-line change in `/etc/systemd/timesyncd.conf`. +28. **`fstrim.timer` enabled** for the SSD pool — verify with `systemctl status fstrim.timer`. Default-on in modern PVE, but quick to confirm. +29. **Watchdog config** is irrelevant for a single-node setup (HA is the use case), so leave the default. Mentioned only so future agents don't add it speculatively. + +## Docker LXC (104) — observations & wins + +Confirmed allocation: 16 CPU, 31.25 GiB RAM (7.86 GiB used / 25%), 8 GiB swap (idle), **200 GiB boot disk at 47.7% used** — disk pressure outranks RAM pressure. + +30. **Fix the Intel iGPU passthrough.** The container's own notes flag the binding as "likely failing". Once `/dev/dri/{card0,renderD128}` is visible inside the LXC, both TEI and `infinity` can run embeddings on the Arc iGPU via OpenVINO / IPEX-LLM — typically **5–10× faster** than CPU. Equally, Immich's CLIP can be GPU-accelerated. The required config in `/etc/pve/lxc/104.conf`: + ``` + lxc.cgroup2.devices.allow: c 226:0 rwm + lxc.cgroup2.devices.allow: c 226:128 rwm + lxc.cgroup2.devices.allow: c 29:0 rwm + lxc.mount.entry: /dev/dri/card0 dev/dri/card0 none bind,optional,create=file + lxc.mount.entry: /dev/dri/renderD128 dev/dri/renderD128 none bind,optional,create=file + lxc.idmap: u 0 100000 65536 + lxc.idmap: g 0 100000 65536 + lxc.idmap: g 44 44 1 # video group on host + lxc.idmap: g 104 104 1 # render group on host + ``` + Then inside the container: `usermod -aG video,render ` and run TEI with `--device cuda` replaced by the OpenVINO build (`text-embeddings-inference:cpu-1.5-openvino`). Not needed for Phase-3.1's volume but a clean upgrade path. +31. **Match the homelab-wide UNAS storage convention — host bind-mount of `/mnt/pve/unas/services//`.** Verified via the running Karakeep stack (Q19): every container in LXC 104 binds its persistent data into `/mnt/pve/unas/services//...` directly, no `driver_opts`, no CIFS, no credentials in compose. Nexa follows the same pattern. Targets: + - `/mnt/pve/unas/services/nexa/qdrant/` — Qdrant data dir. + - `/mnt/pve/unas/services/nexa/tei-cache/` — embedding model cache (keeps multi-GB models off the boot disk). + - `/mnt/pve/unas/services/nexa/graphdb/` — Phase-3.4 RDF store. + - `/mnt/pve/unas/backup/nexa/snapshots/qdrant//` — daily Qdrant snapshots before rsync to S3 (Phase 3.3 — Q12 open). Mirrors the `backup//` pattern used by Home Assistant / Immich. + - **n8n executions retention**: `EXECUTIONS_DATA_PRUNE=true`, `EXECUTIONS_DATA_MAX_AGE=168`. + - **Monthly `docker image prune --all --filter "until=720h"`** to clear dangling layers. + + Pre-deploy step: `mkdir -p /mnt/pve/unas/services/nexa/{qdrant,tei-cache,graphdb}` and `mkdir -p /mnt/pve/unas/backup/nexa/snapshots/{qdrant,graphdb}` on the docker LXC. Then the compose volumes block is just `- /mnt/pve/unas/services/nexa/:/`. Concrete example in [09-deployment §Step 3](./09-deployment.md#step-3--self-hosted-embeddings-tei). + + Secrets go in the stack's `.env` next to the compose (Karakeep convention: `env_file: .env`). Vaultwarden is the long-term source-of-truth for those secrets ([12/#11](#11)) but isn't required day-one. The earlier proposal of mounting SMB as a docker volume is **withdrawn** ([12/#27](#27)) — it was an over-fitting to the Nextcloud-specific NFS issue. +32. **The LXC has its own 8 GiB swap.** Combined with the host's 31 GiB, that's a lot of swap for guests that should never page. Drop the LXC swap allocation to 1–2 GiB (`pct set 104 -swap 2048`) — frees disk on the LVM-thin pool and forces issues to surface earlier rather than silently swap. +33. **Stacks are managed in Arcane; Dockge is stale.** The LXC was originally provisioned with the Dockge helper-script template, but the user moved on to **Arcane** as the day-to-day docker manager. Deployment of the Nexa stack therefore goes through Arcane, not Dockge. Cleanup task: tar `services/dockge/` (if it exists) into `backup/dockge/`, retire the Dockge container, drop the stale data dir. **Dozzle** stays — it's the log viewer, not a manager, so it isn't redundant with Arcane. +34. **Untriaged local mail on the LXC.** Console shows `You have new mail.` at login — the system mail spool on `/var/mail/root` has unread messages, almost always cron job failures. `mailx` or `mutt` to inspect, then either fix the failing job or send the spool to `nexa.system` ntfy via a tiny aliases entry (`root: |/usr/local/bin/spool-to-ntfy.sh`). + +## Housekeeping campaign (cross-cutting, do together) + +These three are inter-related — picking them up as one campaign is cheaper than chasing each individually, because the audit step is the same. They also map cleanly onto **[Phase 6 (Nexa as homelab steward)](./02-roadmap.md#phase-6--nexa-as-homelab-steward)**: once Nexa can poll Arcane, diff against the documented state and emit actionable findings, items #35–37 become semi-automatic — Nexa proposes the migrations rather than us hunting them down. + +35. **Consolidate Postgres instances.** Today there are at least four independent Postgres containers running — visible from Arcane: `immich_postgres`, `lobe-postgres`, `litellm_db`, `paperless-ngx-db-1` (Karakeep uses Meilisearch + maybe SQLite, separate). Each idles around 100–300 MB RAM and has its own backup story. Two paths: + - **Single shared cluster** (`pgsql` container with one role per app, one database per app). Most modern apps support `DATABASE_URL` → just point them at the shared instance. Saves ~600 MB–1 GB RAM and consolidates backups to one pg_dump cron. + - **Or** migrate the smallest ones to SQLite (LiteLLM has SQLite mode; Paperless does too) and keep Postgres only for Immich. Same outcome, less migration risk. + + Pre-step: list every running container with `docker ps --format '{{.Names}}\t{{.Image}}' | grep -i 'postgres\|mariadb\|mysql'` to inventory exactly what's running. + +36. **UNAS-integration audit.** `/mnt/pve/unas/services//` is the homelab convention (host bind-mount, no driver opts — see #31). Not every container follows it yet. Walk every Arcane stack and check the `volumes:` block: any `- /var/lib/docker/...` or anonymous-volume entry is non-compliant. Suspected non-compliant (need verification): + - `qdrant_scientific` — vector data possibly on local boot disk; **critical to verify** before Phase-3.1 piles on a `nexa_knowledge_text` collection. + - `litellm_db`, `lobe-postgres`, `lobe-redis` — DB containers, persistent state. + - `memos` — the user-facing notes app; loss = data loss. + - `n8n` — workflows + executions DB. + - `audiobookshelf` — listening progress + library metadata. + - the `*-mcp` containers — likely stateless (cache only), low priority. + + Output: a one-page table `service | persistent? | currently bound to | should be bound to`. Then migrate the non-compliant ones one-by-one (stop → rsync data to `services//` on UNAS → re-create stack with the new bind → verify → keep the old volume for 7 days as a rollback). Lock the convention for any new stack going forward. + +37. **3-2-1 backup tiering: warm on-site (S3) + cold off-site.** + `s3.nuclide.systems` is already up but it's **on-site** — same building, same power, same (in-)susceptibility to fire / theft / ransomware. By itself it's a *warm* tier, not a disaster-recovery copy. The full design is two tiers: + + **Warm tier — `s3.nuclide.systems` (already exists, just needs a bucket):** + - **Restic / borg** repos for `backup/home-assistant/`, `backup/immich/`, `backup/nextcloud/`, future `backup/nexa/`. Backrest already orchestrates Borg — just add the S3 destination. + - **Qdrant + GraphDB snapshots** flow UNAS → S3 daily. + - Resolves [Q12](./11-open-questions.md#q12) once the bucket name + access key are set. + + **Cold tier — off-site provider (decision pending — see [Q20](./11-open-questions.md#q20)):** + - **Jottacloud "Unlimited"** (~€9.50/mo, EU/Norway, soft-cap ~5 TB) — user's stated candidate. Best price/value at current 2 TB; reach soft cap in ~10 years. Use `rclone jottacloud:` or native `jotta-cli`. + - **Hetzner Storage Box BX21** (€13/mo, 5 TB EU/DE) — predictable quota, native Borg/Restic/SFTP. Cheapest predictable EU alternative. + - **Backblaze B2** (~$12/mo for 2 TB, S3-compatible, US) — cheapest with the widest tooling support; egress is paid (~$10/TB) which only bites during *full* restores. + - **rsync.net** (~$30/mo) — ZFS send/recv natively; overkill unless you want pool replication. + - **Storj DCS** ($4/TB/mo, decentralized, S3-compatible) — newer ecosystem. + + **Always encrypt before upload** regardless of provider — `restic` or `rclone crypt` over the chosen remote. The provider sees only ciphertext blobs. Key material lives in Vaultwarden + a printed-paper offline copy. + + **Tiering & cadence:** + ``` + [Source] UNAS RAID-6 + native snapshots (#38) + │ daily restic/borg + ▼ + [Warm] s3.nuclide.systems (on-site, fast restore) + │ weekly rclone copy + crypt + ▼ + [Cold] Off-site (Jottacloud / B2 / Hetzner) ← satisfies the "1" in 3-2-1 + ``` + + **What goes off-site (priority order):** + 1. Immich photo originals — irreplaceable. + 2. `media/documents/`, `_sortMe/Downloads/`, Nextcloud user data — irreplaceable. + 3. Memos DB, n8n workflows, Nexa Qdrant + GraphDB snapshots — replicable from sources but expensive to redo. + 4. Vaultwarden DB — *cryptographically* sensitive but small; explicitly include. + + **What does not need off-site:** movies, music, audiobooks, ROMs, derived caches (Immich thumbnails, encoded-video). Those re-download or regenerate. + + Pre-deploy step: choose one provider, create a test bucket/folder, dry-run a `restic init` + `restic backup` of `backup/nexa/` first as a smoke test before pointing the heavyweight repos at it. + +38. **Configure UNAS Pool snapshots.** The UniFi Drive dashboard shows the storage-pool snapshot schedule as "Click to Setup" — i.e. **not configured**. RAID 6 protects against drive failure, not against an `rm -rf` from a misbehaving container or a Nextcloud user mass-delete. Cheap fix: set a daily snapshot with 7-day retention + a weekly with 4-week retention on Pool 1. Native to UniFi Drive, no agent needed. **Highest-leverage data-protection change in the homelab right now** — costs nothing, recovers everything. + +39. **Use UNAS native snapshots for fast Nexa rollback** once #38 is set. Nexa workflows that mutate large state (re-embedding the whole vault, graph rebuild, list-config rewrite) can `pre-snapshot → operate → verify → release` against the same UniFi Drive snapshots. Cheaper and faster than restoring from S3. diff --git a/stacks/nexa/docs/13-information-wishlist.md b/stacks/nexa/docs/13-information-wishlist.md new file mode 100644 index 0000000..7999bba --- /dev/null +++ b/stacks/nexa/docs/13-information-wishlist.md @@ -0,0 +1,96 @@ +# 13 — Information Wishlist + +What additional system inventory would sharpen future decisions, packaged as the **smallest set of paste-and-run tasks** that still answers everything material. Each task is independent — run any subset, in any order. + +Each task header tells you **exactly where to run it** (which shell or which UI). Output goes into a Memos draft, a comment in this conversation, or `docs/inventory/-.txt` — whichever is easiest. + +--- + +## ~~Task 1 — Stack inventory + reference compose~~ ✅ DONE + +User pasted the Karakeep compose. Q19 resolved → host bind-mount of `/mnt/pve/unas/services//` is the convention. Nexa stack updated in [docs/09 §Step 3](./09-deployment.md#step-3--self-hosted-embeddings-tei). The `docker ps -a` half (running-container inventory) is still useful when we get to housekeeping #36 (UNAS-integration audit) — but that's not blocking now. + +--- + +## Task 2 — Storage map (Proxmox + docker) + +📍 **Where:** *Two shells* — first the **Proxmox host** (Datacenter → `nuc` → **Shell**, or `ssh root@192.168.1.20`), then back into the **LXC 104 console**. + +Tells us which Proxmox storage backs what, and which docker volumes are local vs. SMB. + +```bash +# (a) on the Proxmox host (192.168.1.20): +cat /etc/pve/storage.cfg +cat /etc/pve/lxc/104.conf +``` + +```bash +# (b) on LXC 104 (192.168.1.40): +docker volume ls +ls -l /dev/dri/ # confirms whether iGPU passthrough actually works +``` + +**Unblocks:** optimization #30 (iGPU), #36 (UNAS audit), B6/B7/C17. + +--- + +## Task 3 — Postgres / DB workload inventory + +📍 **Where:** **LXC 104 console** (same shell as Task 1). + +Direct input to housekeeping #35 (consolidate postgres instances). + +```bash +for c in $(docker ps --format '{{.Names}}' | grep -iE 'postgres|mariadb|mysql|_db$'); do + echo "=== $c ===" + docker exec "$c" sh -c 'psql -U postgres -l 2>/dev/null || mysql -e "show databases" 2>/dev/null' +done +``` + +**Unblocks:** #35 (consolidate vs. migrate-to-SQLite decision). + +--- + +## Task 4 — LiteLLM model list for the Nexa key + +📍 **Where:** *Either* the **LiteLLM admin UI** (one screenshot) *or* any shell with `$SAIA_API_KEY` exported. + +- **Screenshot path:** open `https://ai.nuclide.systems` (Lobehub) → Settings → Model List, filtered to the Nexa virtual key. Screenshot the model rows. +- **CLI path:** any shell on your PC or the LXC, with `SAIA_API_KEY` set: + ```bash + curl -sH "Authorization: Bearer $SAIA_API_KEY" \ + https://ai.nuclide.systems/v1/models | jq '.data[].id' + ``` + +**Unblocks:** Q18 follow-up — which embedding model SAIA proxies and at what dim, so we can decide if any flow can short-circuit TEI. + +--- + +## Task 5 — S3 archive credentials + +📍 **Where:** **`s3.nuclide.systems` admin UI** in your browser (it's already proxied via Zoraxy — same login as the rest of the homelab). + +Walk to: **Buckets** → either pick an existing Nexa-suitable bucket or create one called `nexa` → note the bucket name. Then **Access Keys** → create a key named `nexa-snapshots` with read/write on that bucket → drop the access-key + secret into Vaultwarden under "Nexa S3", and reply here with **just the bucket name** (the secret stays in Vaultwarden). + +**Unblocks:** Q12, housekeeping #37 (S3 archive tier), Phase-3.3. + +--- + +## Task 6 — Reverse-proxy + DNS authority *(only if needed)* + +📍 **Where:** two browser UIs. + +- **Zoraxy** (`http://192.168.1.4:8000`) → HTTP Proxy → either click "Export" if available, or take a full screenshot of the table. +- **AdGuard** (`http://192.168.1.20` LXC 102 web UI) → Filters → DNS Rewrites → screenshot. + +Skip unless we hit a routing surprise during Phase 1. + +--- + +## When something else is needed + +The smaller items (n8n credentials list, Memos webhook config, smartctl, sample of `_sortMe/`, what cron is failing) only matter when we touch that specific area. The agent will ask for them at the moment they're needed, with the same "📍 Where" framing. + +## Self-serve once Phase 6 ships + +Once [Phase 6 — Nexa as homelab steward](./02-roadmap.md#phase-6--nexa-as-homelab-steward) lands, Nexa runs Tasks 1–3 itself on a schedule and folds the results into a daily drift report. This wishlist becomes a Nexa-managed surface (`#nexa:wishlist-status`) rather than something the user has to remember. diff --git a/stacks/nexa/docs/index.md b/stacks/nexa/docs/index.md new file mode 100644 index 0000000..d7d42c0 --- /dev/null +++ b/stacks/nexa/docs/index.md @@ -0,0 +1,51 @@ +# NEXA Documentation + +**Neural Nexus for Information & Automation** — the central nervous system that ties Memos, n8n, SAIA (LiteLLM), Nextcloud and a vector store into one assistant. + +This is the documentation entry point. Read top-to-bottom for first-time setup, or jump to the section you need. + +--- + +## Table of Contents + +| # | Document | Read when… | +|---|----------|------------| +| 01 | [Vision & Scope](./01-vision-and-scope.md) | You want to understand *what* Nexa is and isn't. | +| 02 | [Roadmap & Phases](./02-roadmap.md) | You want to know the implementation order. | +| 03 | [Architecture Overview](./03-architecture.md) | You need a one-page mental model. | +| 04 | [Integration Matrix](./04-integration-matrix.md) | You're wiring up a new data source or mapping work-vs-personal flows. | +| 05 | [Command System](./05-command-system.md) | You want to know what `#nexa:*` commands do. | +| 06 | [Classification Logic](./06-classification-logic.md) | You're tuning the work/personal router. | +| 07 | [Workflow Spec — Task Router](./07-workflow-spec.md) | You're building the Phase-2 router workflow. | +| 08 | [GraphRAG Architecture](./08-graphrag-architecture.md) | You're working on Phase 3 (Qdrant + Graph). | +| 09 | [Deployment](./09-deployment.md) | You're bringing Nexa up on the real infrastructure. | +| 10 | [Operations](./10-operations.md) | You need backup, monitoring or troubleshooting. | +| 11 | [Open Questions (user-info-required)](./11-open-questions.md) | Items the user still has to answer before progress. | +| 12 | [Optimization Opportunities](./12-optimization-opportunities.md) | Ideas worth considering for the wider homelab. | +| 13 | [Information Wishlist](./13-information-wishlist.md) | What additional system inventory would sharpen future decisions. | + +--- + +## Reading paths + +- **First-time installer:** 01 → 03 → 09 → 10 → 11 +- **Workflow author:** 04 → 05 → 07 → 06 +- **Phase-3 work (memory):** 08 → 09 (§ Phase 3 add-on) +- **Stakeholder / quick read:** 01 → 02 → 12 + +--- + +## Repository layout + +``` +nexa/ +├── README.md → points here +├── docs/ → you are here +└── nexa-core/ → the actual project + ├── ai-prompts/ → SAIA system prompts + ├── config/ → runtime config (YAML, JSON schema) + ├── n8n-workflows/ → exported workflows, version-controlled + └── scripts/ → automation helpers +``` + +Source-of-truth for *runnable* config stays in `nexa-core/`. Documentation lives here in `docs/`. diff --git a/stacks/storage.md b/stacks/storage.md new file mode 100644 index 0000000..af25025 --- /dev/null +++ b/stacks/storage.md @@ -0,0 +1,409 @@ +# Storage Layout — NUC 14 Docker Stacks + +## Data Location Architecture + +All persistent service data is being consolidated to `/mnt/pve/unas/services/{service}/` +on the UNAS NFS share, with media content in `/mnt/pve/unas/media/`. + +**PostgreSQL databases stay on local NVMe** — NFS ownership checks (`stat(2)` UID match) prevent +PostgreSQL from running on network filesystems. PostgreSQL is backed up via `pg_dump` to UNAS instead. + +| Mount | Source | Content | +|---|---|---| +| `/mnt/pve/unas/services/` | UNAS NFS | SQLite DBs, configs, files, S3 data | +| `/mnt/pve/unas/media/` | UNAS NFS | User media (photos, docs, books, audio) | +| `/mnt/pve/unas/backup/` | UNAS NFS | Service backups | +| `/var/lib/docker/volumes/` | Local NVMe | PostgreSQL data, Redis data | +| `/opt/stacks/` | Local NVMe | Compose files, small configs | + +## Migration Status + +### ✅ Already on UNAS — `/mnt/pve/unas/services/` + +| Service | UNAS Path(s) | Container Mount | Data Type | +|---|---|---|---| +| gluetun | `gluetun/data/` | `/gluetun` | VPN config | +| immich | `immich/{encoded-video,profile,thumbs}/` | `/usr/src/app/upload/{encoded-video,profile,thumbs}` | App derived data | +| karakeep | `karakeep/data/`, `karakeep/meilisearch/` | `/data`, `/meili_data` | Bookmarks DB + Meilisearch index | +| ntfy | `ntfy/{config,cache}/` | `/etc/ntfy`, `/var/lib/ntfy` | Notification config + cache | +| paperless-ai | `paperless-ai/` | `/app/data` | AI processing data | +| pocketid | `pocketid/` | `/app/data` | Auth data (SQLite) | +| stremio | `stremio/` | `/root/.stremio-server` | Streaming server data | +| traccar | `traccar/{data,logs,traccar.xml}` | `/opt/traccar/{data,logs,conf/traccar.xml}` | GPS tracking data | +| vaultwarden | `vaultwarden/` | `/data` | Password vault (SQLite) | +| garage | `shared-db/garage/{data,meta}/` | `/var/lib/garage/{data,meta}` | S3 object storage (Garage) | + +### 🆕 Shared Infrastructure — `/opt/stacks/shared-db/` + +| Service | Hostname | Type | Data Location | +|---|---|---|---| +| shared-postgres | `shared-postgres` | PostgreSQL 16 (local volume) | `shared-db_shared-pgdata` volume | +| garage | `garage` | S3-compatible object storage (on UNAS) | `/mnt/pve/unas/services/shared-db/garage/` | + +### ✅ Migrated (Phase 1) — SQLite to UNAS + +| Service | Old Path | New Path | Status | +|---|---|---|---| +| memos | `/opt/stacks/memos/data` | `/mnt/pve/unas/services/memos` | ✅ Healthy | +| arcane | volume `arcane_arcane-data` | `/mnt/pve/unas/services/arcane` | ✅ Running | +| n8n | volume `n8n_n8n_storage` | `/mnt/pve/unas/services/n8n` | ✅ Healthy | + +### ⏳ Still Local — `/opt/stacks/` or Docker Volumes + +| Service | Current Location | Data Type | Estimated Size | UNAS Target | +|---|---|---|---|---| +| arr-stack | `/opt/stacks/arr-stack/` | Configs + media | ~55 GB | `arr-stack/{config,audiobooks,ebooks,podcasts,torrents}` | +| qdrant | `/opt/stacks/qdrant/qdrant_storage` | Vector DB | ~7 MB | `qdrant/` | +| dozzle | `/opt/stacks/dozzle/dozzle_data` | Log viewer state | tiny | *(ephemeral, optional)* | +| homepage | `/opt/stacks/homepage/{config,icons}` | Dashboard config | ~164 KB | *(keep in stack dir)* | + +### 🗄️ Databases Requiring Migration + +| Service | Engine | Current Path | Target | Migration Method | Status | +|---|---|---|---|---|---|---| +| immich | PostgreSQL | `/opt/stacks/immich/postgres` | → shared-postgres | `pg_dump` → `pg_restore` | ⏳ Pending | +| litellm | PostgreSQL | `/opt/stacks/ai/postgres_data` | → shared-postgres | `pg_dump` → `pg_restore` | ⏳ Pending | +| lobehub | PostgreSQL | `/opt/stacks/ai/lobehub/data` | → shared-postgres | `pg_dump` → `pg_restore` | ⏳ Pending | +| daytona | PostgreSQL | volume `daytona-minimal_db_data` | → shared-postgres | `pg_dump` → `pg_restore` | ⏳ Pending | +| paperless-ngx | PostgreSQL | volume `paperless-ngx_pgdata` | → shared-postgres | `pg_dump` → `pg_restore` | ⏳ Pending | +| memos | SQLite | `/opt/stacks/memos/data` | `/mnt/pve/unas/services/memos` | `cp` + `chown` | ✅ Done | +| arcane | SQLite | volume `arcane_arcane-data` | `/mnt/pve/unas/services/arcane` | `cp` + `chown` | ✅ Done | +| n8n | SQLite | volume `n8n_n8n_storage` | `/mnt/pve/unas/services/n8n` | `cp` + `chown` | ✅ Done | +| pocketid | SQLite | `pocketid/` | *(already on UNAS)* | — | ✅ Done | +| vaultwarden | SQLite | `vaultwarden/` | *(already on UNAS)* | — | ✅ Done | +| karakeep | SQLite + Meilisearch | `karakeep/{data,meilisearch}` | *(already on UNAS)* | — | ✅ Done | + +### 🧠 ML Models & Caches (Regenerable — Can Stay Local) + +| Resource | Path | Decision | +|---|---|---| +| Immich ML models | volume `immich_model-cache` | Cache — stay | +| Paperless Redis | volume `paperless-ngx_redisdata` | Ephemeral — stay | +| AI stack Redis | volume `nuc-ai-core_redis_data` | Ephemeral — stay | +| SearXNG search cache | volume (hash) at `/var/lib/docker/volumes/bb5bb...` | Ephemeral — stay | +| Daytona runner Docker | volume `daytona-minimal` runner data | Can stay | + +## Media Content (Already on UNAS) + +| Mount | Content | +|---|---| +| `/mnt/pve/unas/media/images/` | Immich photo library | +| `/mnt/pve/unas/media/documents/public/paperless-ngx/` | Paperless documents (consume, export, library) | +| `/mnt/pve/unas/media/audiobooks/` | Audiobookshelf media | +| `/mnt/pve/unas/media/ebooks/` | Audiobookshelf ebooks | +| `/mnt/pve/unas/media/podcasts/` | Audiobookshelf podcasts | +| `/mnt/pve/unas/backup/immich/` | Immich DB backups | + +## Prospective UNAS Layout + +``` +/mnt/pve/unas/ +├── services/ +│ ├── gluetun/ +│ │ └── data/ ← VPN config +│ ├── immich/ +│ │ ├── db/ ← PostgreSQL (to migrate) +│ │ ├── encoded-video/ ← Already migrated +│ │ ├── profile/ ← Already migrated +│ │ └── thumbs/ ← Already migrated +│ ├── karakeep/ +│ │ ├── data/ ← Already migrated +│ │ └── meilisearch/ ← Already migrated +│ ├── ntfy/ +│ │ ├── cache/ ← Already migrated +│ │ └── config/ ← Already migrated +│ ├── paperless-ai/ ← Already migrated +│ ├── paperless-ngx/ +│ │ └── pgdata/ ← PostgreSQL (to migrate) +│ ├── pocketid/ ← Already migrated +│ ├── stremio/ ← Already migrated +│ ├── traccar/ +│ │ ├── data/ ← Already migrated +│ │ ├── logs/ ← Already migrated +│ │ └── traccar.xml ← Already migrated +│ ├── vaultwarden/ ← Already migrated +│ ├── memos/ ← SQLite ✅ Done +│ ├── qdrant/ ← Vector DB (to migrate) +│ ├── arcane/ ← SQLite ✅ Done +│ ├── n8n/ ← SQLite ✅ Done +│ ├── shared-db/ +│ │ ├── postgres/ ← ⚠️ Local volume, NOT on UNAS +│ │ └── garage/ +│ │ ├── data/ ← S3 data (on UNAS) +│ │ └── meta/ ← S3 metadata (on UNAS) +│ ├── arr-stack/ +│ │ ├── config/ ← Bind mounts (to migrate) +│ │ ├── audiobooks/ ← Media (to migrate) +│ │ ├── ebooks/ ← Media (to migrate) +│ │ ├── podcasts/ ← Media (to migrate) +│ │ └── torrents/ ← Downloads (to migrate) +│ +├── media/ +│ ├── images/ ← Immich photo library +│ └── documents/ +│ └── public/ +│ └── paperless-ngx/ +│ ├── consume/ +│ ├── export/ +│ └── library/ +│ +├── backups/ +│ └── immich/ ← Immich DB backups +│ +└── (potential future: torrents/, cache/) +``` + +## Garage S3 — `/opt/stacks/shared-db/` + +S3-compatible object storage using [Garage](https://garagehq.deuxfleurs.fr/) v1.0.0. +Replaces RustFS from the LobeHub AI stack. Single-node, replication_factor=1. + +| Detail | Value | +|---|---| +| Container | `garage` (dxflrs/garage:v1.0.0) | +| Data | `/mnt/pve/unas/services/shared-db/garage/{data,meta}` | +| S3 API | `garage:3900` (internal, shared_backend) | +| Admin API | `127.0.0.1:3901` (localhost only) | +| Regional | No auth (internal network) | +| Bucket | `lobe-files` | +| Key ID | `GK55210bd6b56d92891db26fac` | +| Secret Key | `d0af9e4eb69d7773ffa17950feee27925dbc70eb3efd4ee23a3e1324672c869a` | + +## Shared PostgreSQL — `/opt/stacks/shared-db/` + +Centralized PostgreSQL 16 instance for all non-immich services. + +| Detail | Value | +|---|---| +| Container | `shared-postgres` (postgres:16) | +| Data | `shared-db_shared-pgdata` (local Docker volume, NOT NFS) | +| Networks | `shared_backend` | +| Auth | `postgres` superuser with password from `.env` | + +### Databases & Users (pre-created on init) + +| Database | Owner | Password | Used By | +|---|---|---|---| +| `paperless` | `paperless` | `paperless` | paperless-ngx | +| `daytona` | `user` | `pass` | daytona | +| `litellm` | `litellm` | `litellm_password_here` | LiteLLM | +| `lobechat` | `lobe` | `uWNZugjBqixf8dxC` | LobeHub | +| `paradedb` | `lobe` | `uWNZugjBqixf8dxC` | LobeHub | + +## Key Architectural Decision: PostgreSQL Stays Local + +**Problem:** PostgreSQL checks data directory ownership via `stat(2)` and refuses to start +if the process UID doesn't match the file owner UID. NFS with `all_squash` maps all UIDs +to `65534` (nobody), breaking this check. `chown` also fails on NFS. + +**Solution:** All PostgreSQL instances use local Docker volumes on NVMe: + +``` +Local NVMe (/var/lib/docker/volumes/) UNAS NFS (/mnt/pve/unas/) +├── shared-postgres data ├── pg_dump backups ← scheduled +├── immich_postgres data ├── SQLite DBs (memos, arcane, n8n) +├── paperless-ngx pgdata ├── Garage S3 data +├── daytona-minimal db_data ├── Media files (immich, paperless) +├── litellm-db / ai postgres_data ├── Configs / service data +├── lobehub postgres data +└── all Redis / cache volumes +``` + +**Backup strategy:** `pg_dump` each database to UNAS on a schedule (e.g. daily cron). + +## Migration Priority + +| Priority | Service | Effort | Risk | Status | Reason | +|---|---|---|---|---|---| +| P0 | memos | Minimal (SQLite copy) | Low | ✅ Done | Simple file move | +| P0 | n8n | Minimal (SQLite copy) | Low | ✅ Done | Simple file move | +| P0 | arcane | Minimal (SQLite copy) | Low | ✅ Done | Simple file move | +| P0 | qdrant | Minimal (file copy) | Low | ✅ Done | Fresh start on UNAS | +| P1 | daytona → shared-postgres | Medium (pg_dump) | Low | ✅ Done | pg_dump → restore → running | +| P1 | litellm → shared-postgres | Medium (pg_dump) | Low | ✅ Done | pg_dump → restore → running | +| P1 | lobehub → shared-postgres | Medium (pg_dump) | Low | ⛔ Skipped | Uses ParadeDB (pg17+extensions), incompatible with shared-postgres (pg16) | +| P1 | paperless-ngx → shared-postgres | Medium (pg_dump) | Medium | ✅ Done | pg_dump → restore → running | +| P1 | immich → shared-postgres | Medium (pg_dump) | Medium | ⏳ | Custom pgvector image | +| P2 | arr-stack | Large (55 GB copy) | Low | ⏳ | Simple copy, time-consuming | + +## Mail Archive (Potential Savings) + +Unattached anonymous volumes from old/stopped containers: + +| Volume | Size | Likely From | Action | +|---|---|---|---| +| `0b6d07ff...` | ? | old daytona runner | prune | +| (100+ hex-named volumes) | ? | ephemeral build cache | prune | + +Recommend: `docker system df` then `docker volume prune` to clean. + +--- + +## Database Consolidation Plan + +### Current: 10 Database Instances + +| # | Container | Type | Image | Used By | Networks | +|---|---|---|---|---|---| +| 1 | `immich_postgres` | PostgreSQL | `immich-app/postgres:14-vectorchord` (pgvector) | immich | `immich_default` | +| 2 | `paperless-ngx-db-1` | PostgreSQL | `postgres:16` | paperless-ngx | `paperless-ngx_default` | +| 3 | `daytona-minimal-db-1` | PostgreSQL | `postgres:16-alpine` | daytona | `daytona-minimal_daytona-network` | +| 4 | `nuc-ai-core-litellm-db-1` | PostgreSQL | `postgres:16-alpine` | litellm | `ai-internal` | +| 5 | `lobe-postgres` | PostgreSQL | `postgres:16-alpine` | lobehub (lobechat, paradedb) | `ai-internal` | +| 6 | `paperless-ngx-broker-1` | Redis 7 | `redis:7` | paperless broker | `paperless-ngx_default` | +| 7 | `immich_redis` | Valkey 9 | `valkey/valkey:9` | immich cache | `immich_default` | +| 8 | `daytona-minimal-redis-1` | Redis Alpine | `redis:alpine` | daytona | `daytona-minimal_daytona-network` | +| 9 | `lobe-redis` | Redis 7 | `redis:7-alpine` | lobehub | `ai-internal` | +| 10 | `redis-searxng` | Redis 7 | `redis:7-alpine` | searxng | `ai-internal` | + +### Target: 6 Database Instances (remove 4) + +| # | Container | Type | DBs Hosted | Saved | +|---|---|---|---|---| +| 1 | `immich_postgres` | PostgreSQL immich-app (pgvector) | immich | 🔒 Must stay separate | +| 2 | **`shared-postgres`** 🆕 | **PostgreSQL 16** | **paperless, daytona, litellm, lobechat, paradedb** | **−3 containers** | +| 3 | `paperless-ngx-broker-1` | Redis 7 | paperless broker | 🔒 Must stay separate | +| 4 | `immich_redis` | Valkey 9 | immich cache | 🔒 Must stay separate | +| 5 | `daytona-minimal-redis-1` | Redis Alpine | daytona | 🔒 Stay separate | +| 6 | `lobe-redis` | Redis 7 | lobehub | 🔒 Stay separate | +| 7 | `redis-searxng` | Redis 7 | searxng | 🔒 Stay separate | +| 7 | ~~`nuc-ai-core-litellm-db-1`~~ | 🗑️ removed | → shared-postgres | ✅ | +| 8 | ~~`lobe-postgres`~~ | 🗑️ removed | → shared-postgres | ✅ | +| 9 | ~~`paperless-ngx-db-1`~~ | 🗑️ removed | → shared-postgres | ✅ | +| 10 | ~~`daytona-minimal-db-1`~~ | 🗑️ removed | → shared-postgres | ✅ | + +**Result: 10 → 6 instances, removing 4 containers, saving ~160MB RAM.** + +### Why Redis Stays Separate + +| Reason | Detail | +|---|---| +| **Valkey ≠ Redis** | Immich uses Valkey 9 which has diverged from Redis 7 | +| **`SELECT` collisions** | Apps may use specific DB numbers that clash | +| **Password mismatch** | Each app has its own `requirepass` / `REDIS_PASSWORD` | +| **Ephemeral = low value** | Caches/brokers rebuild on restart, savings don't justify risk | +| **`FLUSHALL` risk** | One misconfigured app could wipe another's cache | + +### Why PG Consolidation Works + +| Reason | Detail | +|---|---| +| **Same major version** | All target DBs use PostgreSQL 16 | +| **Native compatibility** | No conflicting extensions (only immich has pgvector) | +| **Isolation via databases** | Each service gets its own database + user within shared PG | +| **Simple restore** | `pg_dump` + `pg_restore` into shared instance | +| **Single backup** | One `pg_dumpall` covers all 5 databases | + +### Implementation: Shared PostgreSQL + +Create `/opt/stacks/shared-db/docker-compose.yml`: + +```yaml +services: + postgres: + image: postgres:16 + container_name: shared-postgres + restart: unless-stopped + volumes: + - /mnt/pve/unas/services/shared-db:/var/lib/postgresql/data + environment: + POSTGRES_PASSWORD: ${SHARED_DB_PASSWORD} + ports: + - "127.0.0.1:5432:5432" + deploy: + resources: + limits: + cpus: "2.0" + memory: 1G + reservations: + cpus: "1.0" + memory: 512M + healthcheck: + test: ["CMD-SHELL", "pg_isready -U postgres"] + interval: 10s + timeout: 5s + retries: 5 + +networks: + default: + external: true + name: shared_backend +``` + +### Migration Order (Phase 2) + +``` +Step 0: Create shared-postgres container +Step 1: Dump ALL 4 databases from their current instances +Step 2: Create users + databases on shared-postgres +Step 3: Restore each dump into shared-postgres +Step 4: One by one, update each app's compose to point to shared-postgres +``` + +#### Detailed: Paperless-ngx + +| Change | Before | After | +|---|---|---| +| DB service | `db` (in paperless compose) | remove entirely | +| `PAPERLESS_DBHOST` | `db` | `shared-postgres` | +| DB credentials | defaults (paperless/paperless) | keep same | +| Network | `paperless-ngx_default` | `shared_backend` | + +#### Detailed: Daytona + +| Change | Before | After | +|---|---|---| +| DB service | `db` (in daytona compose) | remove entirely | +| `DB_HOST` | `db` | `shared-postgres` | +| `DB_USERNAME` | `user` | keep same | +| `DB_DATABASE` | `daytona` | keep same | +| Network | `daytona-minimal_daytona-network` | `shared_backend` | + +#### Detailed: LiteLLM + +| Change | Before | After | +|---|---|---| +| DB service | `litellm-db` (in ai compose) | remove entirely | +| `DATABASE_URL` | `postgresql://litellm:...@litellm-db:5432/litellm` | `postgresql://litellm:...@shared-postgres:5432/litellm` | +| Network | `ai-internal` | `shared_backend` | + +#### Detailed: LobeHub + +| Change | Before | After | +|---|---|---| +| DB service | `postgresql` (in lobehub.yml) | remove entirely | +| `DATABASE_URL` | `postgresql://postgres:...@postgresql:5432/lobechat` | `postgresql://postgres:...@shared-postgres:5432/lobechat` | +| Lobe also needs `paradedb` DB | created manually on initial setup | created via SQL script | +| Network | `ai-internal` | `shared_backend` | + +### Total Impact + +| Metric | Before | After | +|---|---|---| +| PostgreSQL instances | 5 | 2 | +| Redis instances | 5 | 5 (unchanged) | +| Total container count | 10 DBs | 7 DBs (−3) | +| RAM saved (estimated) | — | ~160 MB | +| Complexity | Fragmented | Centralized | +| Backup burden | 5 pg_dumps | 2 pg_dumps | + +### Timing Estimates + +| Service | Downtime | Cumulative | +|---|---|---| +| Setup shared-postgres + dump all | 0 (live dump) | 0 | +| daytona → shared | ~2 min | 2 min | +| litellm → shared | ~2 min | 4 min | +| lobehub → shared | ~3 min | 7 min | +| paperless-ngx → shared | ~5 min | 12 min | +| immich DB → UNAS bind mount | ~10 min | 22 min | +| **Total** | **~22 min** | | + +### Notes + +- All compose files are version-controlled in `/opt/stacks/{service}/docker-compose.yml` +- Always test with `docker compose config` before `up -d` +- PostgreSQL stays on local NVMe — backed up via `pg_dump` to UNAS +- Garage S3 replaces RustFS from LobeHub stack (RustFS will be removed when LobeHub is switched) +- LobeHub's `paradedb` extension requires `CREATE EXTENSION IF NOT EXISTS paradedb;` on the shared instance +- All PostgreSQL dbs consolidated into shared-postgres except immich (pgvector requirement) diff --git a/stacks/todo.md b/stacks/todo.md new file mode 100644 index 0000000..e11277b --- /dev/null +++ b/stacks/todo.md @@ -0,0 +1,171 @@ +# Todo + +## Goal +Consolidate databases, migrate storage to UNAS, integrate all services with Homepage. + +--- + +# Ideas (user) +- Claude Max sub integration with litellm (how to handle auth, rate limits, etc.?) +- MCP deployment process ideas? Current setup seems cumbersome — can litellm host directly? + (how to integrate npx, uvx, and maybe docker-based containers?) +- Select offsite backup hosting provider, set up backup process (maybe restic to Backblaze B2 + or similar), then set up rclone for off-site sync of Garage buckets to remote S3 + (Backblaze B2, Jottacloud, or Hetzner) +- Research the LiteLLM MCP `source_url` column problem across GitHub issues, PRs, and docs + to understand the full picture +- Add Google Gemini (vision + 1M ctx + free text-embedding-004; needs GEMINI_API_KEY) +- Add fal.ai MCP server (`mcp/fal`) for fast cloud image gen < 5 s (vs ~150 s local ComfyUI) +- Add Cerebras (Llama 3 at ~1M tok/s, generous free tier; CEREBRAS_API_KEY) + +# Next Steps / Todos +- [x] **ComfyUI added & fixed (2026-05-17)** — root-caused the self-referential + bind-mount, repaired mounts/healthcheck, migrated CPU→**Intel Arc iGPU** + (torch 2.7.1+xpu, no IPEX), pinned v0.21.1. FLUX.1-schnell GGUF, legible + text, ~94 s/image. Canonical workflow in `ai/comfyui/workflows/`. +- [x] **ComfyUI→LobeChat via MCP (2026-05-17)** — purpose-built `comfyui-mcp` + server, proven end-to-end. See `docs/comfyui-mcp.md`. +- [x] **MCP gateway finished (2026-05-17)** — DinD, real streaming MCP + reverse-proxy, Pocket ID bearer + RFC9728 discovery, OAuth callback fixed, + secrets from `.env`, no port collisions. See `docs/mcp-gateway.md`. +- [x] **LobeChat S3 images fixed (2026-05-17)** — added `S3_PUBLIC_DOMAIN=http://192.168.1.40:10004` + to `lobehub.yml`; pre-signed URLs now use the LAN-accessible Garage port instead of + the internal `garage:3900` hostname. +- [x] **MCP gateway groups + combined endpoints (2026-05-17)** — servers now grouped + (research/dev/personal/image); `/group/{name}/mcp` aggregates all member tools with + `server__tool` namespacing; UI shows grouped cards with copy-URL buttons; `/` → `/ui` + redirect added. +- [x] **LiteLLM Redis caching (2026-05-17)** — `cache: true` with Redis backend + (`lobe-redis`, already on `ai-internal`); 10-min TTL on completion + embedding calls. +- [x] **Cohere added to LiteLLM (2026-05-17)** — `command-a-03-2025` (20 RPM), + `cohere-embed-english/multilingual/light-v3` (2000 RPM); `COHERE_API_KEY` in `.env`. + Note: command-r / command-r-plus retired 2025-09-15. +- [x] **LobeChat MCP import done** — `comfyui-flux` server imported via UI. +- [x] **ComfyUI MCP async queue (2026-05-17)** — full rewrite: `generate_image` + returns job_id in <0.1 s; background polling thread updates state; `img2img` + accepts many images (URL/base64/`job:` chaining); `get_job_status`, + `list_queue`, `cancel_job`, `list_recent_images`. New `flux-img2img-api.json` + workflow (VAEEncode path). See `docs/comfyui-mcp.md`. +- [ ] Nextcloud MCP: bearer token enabled in NC admin (user_oidc provider settings). + Gateway now provisions NC app-password at SSO login time (PocketID token live). + **Needs end-to-end test**: log in via mcp.nuclide.systems, then call /nextcloud/mcp. +- [ ] Observability LXC: deploy OTEL collector + VictoriaMetrics + Grafana + Loki + + Uptime Kuma on a Proxmox LXC; see `docs/stack-ideas.md §8`. +- [ ] Arcane: migrate central server to Proxmox LXC; deploy arcane-agent (headless) + on NUC + .49; connect via TCP Docker socket or SSH Docker context. +- [ ] Docling + n8n ingest pipeline: Nextcloud/Paperless → Docling → LiteLLM embed + → LobeChat knowledge base; see `docs/stack-ideas.md §1`. +- [x] **Gateway spawn servers (2026-05-17)** — papersearch, crawl4ai, markitdown + all running on `ai-internal`. Fixed papersearch command (`uvx --with` not + `uv tool run --from`). nextcloud deferred: Nextcloud OIDC discovery not + enabled; needs `NEXTCLOUD_USERNAME`/`NEXTCLOUD_PASSWORD` in `ai/.env` for + BasicAuth mode (server.py already updated). +- [x] **ComfyUI inpaint smoke-tested (2026-05-17)** — Kodak #23 (768×512) + center + alpha-mask, submitted via API, output in 290s. Workflow confirmed working end-to-end. +- [x] **garage external S3 working** — `s3.nuclide.systems` returns 403 (AccessDenied + on root, expected); Zoraxy routes correctly to `10004→3900`. Was already fixed by + the earlier port binding change (`0.0.0.0:10004`); todo was stale. +- [x] **syncstack gone** — container no longer present; one-shot job completed/pruned. +- [x] **immich_machine_learning healthcheck fixed** — already on :3003/ping. + Container shows healthy. +- [x] **LiteLLM healthcheck fixed** — was hitting host port 14000 inside container; + corrected to container port 4000. Now healthy. +- [x] **immich_power_tools / karakeep / karakeep_chrome** — all healthy (slow warmup + self-resolved). + +## In Progress +- None + +## Blocked +- None + +--- + +# Reference / History + +## Done +- Created `/opt/stacks/shared-db/` with shared-postgres (pg16), Garage S3, pgAdmin +- Consolidated daytona, litellm, paperless-ngx → shared-postgres +- LobeHub kept on own pg17 (ParadeDB) +- Switched LobeHub from RustFS → Garage S3 (`lobe-files` bucket) +- Migrated arr-stack configs + media to UNAS via rsync +- Migrated qdrant to UNAS (fresh start) +- Added 29 selfh.st icons, updated all homepage labels to `sh-` prefix +- Homepage: instance hrefs (192.168.1.40:PORT) + public hrefs (*.nuclide.systems) +- Homepage widgets for 11 services +- Groups consolidated from 14 → 6 (infrastructure, security, media, ai, productivity, diagnostic) +- Configured memos with Garage S3 +- Fixed Garage port binding: `127.0.0.1:10004` → `0.0.0.0:10004` (Zoraxy couldn't reach it) +- Fixed memos S3 endpoint: `https://s3.nuclide.systems` → `http://garage:3900` (TLS cert mismatch) +- Cleaned up old local data remnant at `/opt/stacks/memos/data/` +- traccar GPS port fixed: 15001 → 19001 (paperless-ai port conflict) +- Updated PORTMAP.md, created storage.md and volumes.md +- WAL-G v3.0.8 setup for all 3 PG instances (shared-postgres, immich, lobe) +- Garage S3 buckets + per-instance keys for PG backups +- WAL archiving enabled and flowing to Garage (`wal-g wal-push`) +- Daily full backup cron at 3am (`/etc/cron.d/pg-backup`) +- Home Assistant added to Homepage (infrastructure, widget token) +- ShelfArr added to Homepage (diagnostic, MDI icon) + +## Proxy migration → Zoraxy + clean-state recreate — audit 2026-05-16 + +### Done +- Reverse-proxy plan moved Traefik → **Zoraxy** v3.3.2 (`192.168.1.4:8000`, noauth, internal). + Traefik abandoned: it runs on a different host than the containers, so Docker labels + can't work cross-host; Zoraxy already had the wildcard `*.nuclide.systems` cert + rules. +- Source of truth: `proxy/zoraxy/routes.json` + idempotent CSRF-aware + `scripts/zoraxy_sync.py` (`--dry-run` default / `--apply` / `--prune` / `--list`). + Self-service: edit the JSON, run `--apply`. +- Zoraxy reconciled: fixed `ai`→14000 (litellm), `mcp`→8080 (mcp-gateway), + added `chat`→14001 (lobehub). Pruned (not exposed): dozzle, immich-tools, + mcp-auth. Internal-only: pgadmin, paperless, paperless-ai. +- Fixed AI URL configs (`litellm.nuclide.systems` had no route): + `ai/litellm-config/config.yaml`, `ai/.env`, `karakeep/.env` → `ai.nuclide.systems`; + `paperless_ai/.env` CUSTOM_BASE_URL `mcp`→`ai`. +- paperless-ngx host port `15000:8000` published (internal LAN); paperless-ai 15001 already in compose. +- Stripped **all** `traefik.*` labels from every compose file; fixed memos YAML escaping. +- Cleaned `proxy/` (removed dead Traefik artifacts) + added `proxy/zoraxy/README.md`. +- `shared-db/docker-compose.yml` restored from HEAD (WIP had deleted shared-postgres + + `GARBAGE_` typos) and force-recreated cleanly. +- Removed stale `*_old` containers; clean-state force-recreate of all 19 stacks. + +### Verified working (Zoraxy end-to-end, https via .4) +- ai, chat, mcp, arcane, ntfy, id, vault, abs, n8n, memos, hoarder, daytona, traccar — all serving. +- Static/external: ha, nc, ocpp, shepard, shepard-api — reachable. +- shared-postgres healthy; litellm/lobehub/mcp-gateway healthy & proxied; garage S3 working + internally (active wal-g WAL backups). + +### Issues found → fix outcomes (2026-05-16) +- [x] **FIXED — immich_server crash-loop** (`getaddrinfo ENOTFOUND database`): + the `database` service was on `immich_immich_default`+`shared_backend` while + immich-server/redis/ml were on the implicit `immich_default` → no shared net. + Fix: `immich/docker-compose.yml` top-level network is now + `default: {name: immich_default}` and `database` joins `[default, shared_backend]`. + Verified: immich_server **healthy**, `immich.nuclide.systems` → **200** via Zoraxy. +- [x] **FIXED — vpn_gluetun "unhealthy"**: removed the bogus custom healthcheck + (`wget :8000`, 404s) so gluetun's built-in healthcheck (127.0.0.1:9999) runs. + Verified: `vpn_gluetun` **healthy**; arr-stack recreated (it uses + `network_mode: container:vpn_gluetun`) — prowlarr/rdtclient/shelfarr/flaresolverr healthy. +- [ ] **NOT fixed — garage external S3 (`s3.nuclide.systems`)**: tried + `api_bind_addr`/`rpc_bind_addr` `[::]`→`0.0.0.0`, but applying it needs a + force-recreate, and garage's single-node cluster is fragile to restarts — the + recreate caused "could not reach quorum" (RPC). **Reverted `garage.toml` to HEAD + and recreated**; garage is back to the known-good state (HEALTHY node, layout v1, + buckets OK, internal S3 / PG WAL backups working). External `s3.nuclide.systems` + remains non-responsive **as it was before** — tracked in Todos above. +- [ ] `nuc-ai-core-syncstack-1` left `Created` — tracked in Todos above. + +### Other (pre-existing, NOT in the fix-1/2/3 scope, surfaced by clean recreate) +- `immich_machine_learning` shows "unhealthy" but **works** — its healthcheck tests + `localhost:3000/health` while ML serves `:3003` (immich_server logs confirm + "Machine learning server became healthy"). Cosmetic healthcheck-port mismatch in + the immich compose. +- `immich_power_tools`, `karakeep`, `karakeep_chrome` showing "unhealthy" after the + clean recreate — slow/long healthcheck warmup; revisit if still unhealthy later. + +### Notes +- Zoraxy self-service: `python3 scripts/zoraxy_sync.py [--apply|--prune|--list]`. +- The shared-db WIP (garage/postgres rework) was reverted to HEAD — redo it intentionally + if that work was wanted. + + diff --git a/stacks/volumes.md b/stacks/volumes.md new file mode 100644 index 0000000..d380019 --- /dev/null +++ b/stacks/volumes.md @@ -0,0 +1,345 @@ +# Volume Mounts — NUC 14 Docker Stacks + +> Every bind mount and named volume across all running containers. +> Last updated: May 16, 2026 + +## Legend + +| Column | Meaning | +|---|---| +| **Type** | `bind` = host directory, `volume` = docker named volume, `tmpfs` = memory | +| **Source** | Host path (bind) or volume name (volume) | +| **Container** | Mount point inside container | +| **UNAS** | `/mnt/pve/unas/services/` target (migrated or planned) | + +--- + +## Infrastructure + +### Arcane — `arcane` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| bind | `/var/run/docker.sock` | `/var/run/docker.sock` | — | +| bind | `/opt/stacks` | `/app/data/projects` | — | +| **bind** | **`/mnt/pve/unas/services/arcane`** | **`/app/data`** | ✅ **Migrated** | +| ~~volume~~ | ~~`arcane_arcane-data`~~ | ~~`/app/data`~~ | 🗑️ removed | + +### Dozzle — `dozzle` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| bind | `/opt/stacks/dozzle/dozzle_data` | `/data` | (optional, ephemeral) | +| bind | `/var/run/docker.sock` | `/var/run/docker.sock` | — | + +### Homepage — `homepage` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| bind | `/opt/stacks/homepage/config` | `/app/config` | (keep in stack dir) | +| bind | `/opt/stacks/homepage/icons` | `/app/public/icons` | (keep in stack dir) | +| bind | `/var/run/docker.sock` | `/var/run/docker.sock` | — | + +--- + +## Security + +### Pocket ID — `pocketid` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| **bind** | **`/mnt/pve/unas/services/pocketid`** | **`/app/data`** | ✅ **Already on UNAS** | + +### Vaultwarden — `vaultwarden` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| **bind** | **`/mnt/pve/unas/services/vaultwarden`** | **`/data`** | ✅ **Already on UNAS** | +| bind | `/etc/localtime` | `/etc/localtime` | — | +| bind | `/etc/timezone` | `/etc/timezone` | — | + +--- + +## Media — Immich + +### Immich Server — `immich_server` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| **bind** | **`/mnt/pve/unas/services/immich/encoded-video`** | `/usr/src/app/upload/encoded-video` | ✅ **Already on UNAS** | +| **bind** | **`/mnt/pve/unas/services/immich/profile`** | `/usr/src/app/upload/profile` | ✅ **Already on UNAS** | +| **bind** | **`/mnt/pve/unas/services/immich/thumbs`** | `/usr/src/app/upload/thumbs` | ✅ **Already on UNAS** | +| **bind** | **`/mnt/pve/unas/media/images`** | `/usr/src/app/upload` | ✅ **Already on UNAS** | +| **bind** | **`/mnt/pve/unas/backup/immich`** | `/usr/src/app/upload/backups` | ✅ **Already on UNAS** | +| volume | `7d25f4ac...` (anonymous) | `/data` | (unknown, check) | + +### Immich ML — `immich_machine_learning` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| volume | `immich_model-cache` | `/cache` | (cache, regenerable) | +| bind | `/dev/bus/usb` | `/dev/bus/usb` | — | + +### Immich Postgres — `immich_postgres` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| **bind** | **`/opt/stacks/immich/postgres`** | **`/var/lib/postgresql/data`** | ⏳ Plan: `immich/db` | + +--- + +## Media — Downloads / Arr Stack + +All behind gluetun VPN. + +### RDTClient — `rdtclient` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| bind | `/opt/stacks/arr-stack/rdtclient/config` | `/data/db` | ⏳ Plan: `arr-stack/rdtclient` | +| bind | `/opt/stacks/arr-stack/media/Torrents` | `/data/downloads` | ⏳ Plan: `arr-stack/torrents` | + +### Prowlarr — `prowlarr` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| bind | `/opt/stacks/arr-stack/prowlarr` | `/config` | ⏳ Plan: `arr-stack/prowlarr` | + +### Audiobookshelf — `audiobookshelf` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| bind | `/opt/stacks/arr-stack/audiobookshelf` | `/config` | ⏳ Plan: `arr-stack/audiobookshelf` | +| bind | `/opt/stacks/arr-stack/media/audiobooks` | `/audiobooks` | ⏳ Plan: `arr-stack/audiobooks` | +| bind | `/opt/stacks/arr-stack/media/ebooks` | `/ebooks` | ⏳ Plan: `arr-stack/ebooks` | +| bind | `/opt/stacks/arr-stack/media/podcasts` | `/podcasts` | ⏳ Plan: `arr-stack/podcasts` | + +### ShelfArr — `shelfarr` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| bind | `/opt/stacks/arr-stack/shelfarr/storage` | `/rails/storage` | ⏳ Plan: `arr-stack/shelfarr` | +| bind | `/opt/stacks/arr-stack/media/audiobooks` | `/audiobooks` | (shared) | +| bind | `/opt/stacks/arr-stack/media/ebooks` | `/ebooks` | (shared) | +| bind | `/opt/stacks/arr-stack/media/Torrents` | `/downloads` | (shared) | + +### Flaresolverr — `flaresolverr` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| bind | `/var/lib/flaresolver` | `/config` | (cache, regenerable) | + +--- + +## AI Stack + +### LiteLLM — `litellm` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| bind | `/opt/stacks/ai/litellm-config` | `/app/config` | (keep in stack dir) | + +### LiteLLM DB — `nuc-ai-core-litellm-db-1` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| **bind** | **`/opt/stacks/ai/postgres_data`** | **`/var/lib/postgresql/data`** | ⏳ Plan: `ai/litellm-db` | + +### LobeHub — `lobehub` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| bind | `/opt/stacks/ai/lobehub/data` | `/var/lib/postgresql/data` | ⏳ Plan: `ai/lobehub-db` | + +### LobeHub Redis — `lobe-redis` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| volume | `nuc-ai-core_redis_data` | `/data` | (ephemeral, stay) | + +### LobeHub RustFS — `lobe-rustfs` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| volume | `nuc-ai-core_rustfs-data` | `/data` | ⏳ Plan: `ai/rustfs` | + +### SearXNG — `searxng` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| bind | `/opt/stacks/ai/searxng` | `/etc/searxng` | (keep in stack dir) | +| volume | `bb5bb182...` (anonymous) | `/var/cache/searxng` | (cache, regenerable) | + +### SearXNG Redis — `redis-searxng` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| volume | `f9c3c386...` (anonymous) | `/data` | (ephemeral, stay) | + +### SAIA Image Proxy — `nuc-ai-core-saia-image-proxy-1` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| bind | `/opt/stacks/ai/saia-image-proxy` | `/app` | (keep in stack dir) | + +### Crawl4AI — `crawl4ai-mcp` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| tmpfs | `/dev/shm` | `/dev/shm` | — | + +--- + +## Documents + +### Paperless-ngx Webserver — `paperless-ngx-webserver-1` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| **bind** | **`/mnt/pve/unas/media/documents/public/paperless-ngx/consume`** | `/usr/src/paperless/consume` | ✅ **Already on UNAS** | +| **bind** | **`/mnt/pve/unas/media/documents/public/paperless-ngx/export`** | `/usr/src/paperless/export` | ✅ **Already on UNAS** | +| **bind** | **`/mnt/pve/unas/media/documents/public/paperless-ngx/library`** | `/usr/src/paperless/media` | ✅ **Already on UNAS** | +| volume | `paperless-ngx_data` | `/usr/src/paperless/data` | ⏳ Plan: `paperless-ngx/data` | + +### Paperless-ngx DB — `paperless-ngx-db-1` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| volume | `paperless-ngx_pgdata` | `/var/lib/postgresql/data` | ⏳ Plan: `paperless-ngx/pgdata` | + +### Paperless-ngx Broker — `paperless-ngx-broker-1` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| volume | `paperless-ngx_redisdata` | `/data` | (ephemeral, stay) | + +### Paperless AI — `paperless-ai` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| **bind** | **`/mnt/pve/unas/services/paperless-ai`** | **`/app/data`** | ✅ **Already on UNAS** | + +--- + +## Productivity & Bookmarks + +### Memos — `memos` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| **bind** | **`/mnt/pve/unas/services/memos`** | **`/var/opt/memos`** | ✅ **Migrated** | +| ~~bind~~ | ~~`/opt/stacks/memos/data`~~ | ~~`/var/opt/memos`~~ | 🗑️ replaced | + +### Karakeep — `karakeep` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| **bind** | **`/mnt/pve/unas/services/karakeep/data`** | **`/data`** | ✅ **Already on UNAS** | + +### Karakeep Meilisearch — `karakeep_meilisearch` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| **bind** | **`/mnt/pve/unas/services/karakeep/meilisearch`** | **`/meili_data`** | ✅ **Already on UNAS** | + +--- + +## Automation + +### n8n — `n8n` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| **bind** | **`/mnt/pve/unas/services/n8n`** | **`/home/node/.n8n`** | ✅ **Migrated** | +| bind | `/opt/stacks/n8n/hooks.js` | `/home/node/hooks.js` | (keep in stack dir) | +| ~~volume~~ | ~~`n8n_n8n_storage`~~ | ~~`/home/node/.n8n`~~ | 🗑️ removed | + +--- + +## DevOps + +### Daytona API — `daytona-minimal-api-1` + +No DB-specific mounts needed (connects via env vars). + +### Daytona DB — `daytona-minimal-db-1` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| volume | `daytona-minimal_db_data` | `/var/lib/postgresql/data` | ⏳ Plan: `daytona/db` | + +### Daytona Runner — `daytona-minimal-runner-1` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| volume | `9416da14...` (anonymous) | `/var/lib/docker` | (runner state, stay) | +| bind | `/var/run/docker.sock` | `/var/run/docker.sock` | — | + +--- + +## Tracking + +### Traccar — `traccar` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| **bind** | **`/mnt/pve/unas/services/traccar/data`** | `/opt/traccar/data` | ✅ **Already on UNAS** | +| **bind** | **`/mnt/pve/unas/services/traccar/logs`** | `/opt/traccar/logs` | ✅ **Already on UNAS** | +| **bind** | **`/mnt/pve/unas/services/traccar/traccar.xml`** | `/opt/traccar/conf/traccar.xml` | ✅ **Already on UNAS** | + +--- + +## VPN + +### Gluetun — `vpn_gluetun` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| **bind** | **`/mnt/pve/unas/services/gluetun/data`** | **`/gluetun`** | ✅ **Already on UNAS** | + +--- + +## Shared Infrastructure + +### Shared PostgreSQL — `shared-postgres` + +| Type | Source | Container | Notes | +|---|---|---|---| +| volume | `shared-db_shared-pgdata` | `/var/lib/postgresql/data` | 🗄️ Local NVMe (not NFS) | + +### Garage S3 — `garage` + +| Type | Source | Container | Notes | +|---|---|---|---| +| bind | `/mnt/pve/unas/services/shared-db/garage/data` | `/var/lib/garage/data` | S3 object data | +| bind | `/mnt/pve/unas/services/shared-db/garage/meta` | `/var/lib/garage/meta` | S3 metadata (LMDB) | + +--- + +## Stacks Not Running (Compose Config Only) + +### Streamio — `streamio` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| **bind** | **`/mnt/pve/unas/services/stremio`** | **`/root/.stremio-server`** | ✅ **Already on UNAS** | + +### Qdrant — `qdrant_scientific` + +| Type | Source | Container | UNAS | +|---|---|---|---| +| bind | `/opt/stacks/qdrant/qdrant_storage` | `/qdrant/storage` | ⏳ Plan: `qdrant/` | + +--- + +## Summary: Migration Status + +| Status | Count | Services | +|---|---|---| +| ✅ Already on UNAS | 11 | gluetun, immich*(4), karakeep*(2), ntfy*(2), paperless-ai, pocketid, stremio, traccar*(3), vaultwarden, paperless-docs*(3), garage*(2) | +| ✅ Migrated (Phase 1) | 3 | memos, arcane, n8n | +| 🔷 Shared infrastructure | 2 | shared-postgres (local volume), garage (on UNAS) | +| ⏳ Phase 2 planned (PG consolidation) | 5 | immich → shared-postgres, paperless → shared-postgres, daytona → shared-postgres, litellm → shared-postgres, lobehub → shared-postgres | +| ⏳ Phase 3 planned | 1 | arr-stack*(8 mounts) | +| 📋 Keep local | ~5 | homepage, dozzle, litellm-config, searxng-config, saia-image-proxy | +| 🧠 Cache (stay) | ~5 | immich_model-cache, paperless-ngx redis, lobe-redis, searxng-redis, flaresolverr, daytona-runner |