docs: reconcile storage.md + volumes.md with live 2026-07-15 audit

Fixes contradictions found during the CT 104 disk-full incident:
- Arcane: was DECOMMISSIONED 2026-05-26 + local, not "migrated to UNAS"
  (volumes.md and storage.md disagreed).
- arr-stack: fully on UNAS incl. SQLite configs on services/arr-stack/;
  media/downloads on media/ share. Radarr/Sonarr not deployed. Removed
  stale "plan to migrate" rows pointing at local /opt/stacks/arr-stack.
- n8n: local (DB on CT 113), not UNAS.
- Garage: metadata is SQLite (not LMDB); keep off NFS; refreshed bucket
  table with real sizes (82 GB, mostly WAL-G PG backups).
- Marked sunsets: karakeep, memos, litellm, daytona.
- crawl4ai: correct name mcp-crawl4ai, no mounts, cache-growth TODO.
- Added ZFS snapshot note: refquota hides snapshot bloat from container df;
  pre-tier1-backups snapshot held 97.9 GB.
This commit is contained in:
2026-07-15 11:15:53 +02:00
parent ac55f25201
commit 40a02f33ef
2 changed files with 70 additions and 94 deletions
+29 -13
View File
@@ -1,7 +1,7 @@
# Storage layout
> Verified state as of 2026-05-21. Source of truth for which service lives on which storage class.
> Earlier revisions of this doc were stale on multiple items (Garage, n8n, Arcane, Karakeep, Pocket-ID, arr-stack, Nextcloud's storage protocol, missing Coder/Gitea). Reconciled via live audit.
> Verified state as of 2026-07-15 (live audit during the CT 104 disk-full incident). Source of truth for which service lives on which storage class.
> Earlier revisions were stale on Garage, n8n, Arcane, Karakeep/Memos, Pocket-ID, arr-stack, Nextcloud's storage protocol, and missing Coder/Gitea. `volumes.md` was reconciled in the same pass.
## Storage classes
@@ -24,8 +24,8 @@ Bulk + media + non-latency-sensitive app data.
| Paperless-ngx (documents) | 104 | `media/documents/public/paperless-ngx/{consume,export,library}` | none | |
| Paperless-AI | 104 | `services/paperless-ai``/app/data` | none | |
| Traccar (logs, config) | 104 | `services/traccar/{logs,traccar.xml}` | none | data dir reverted to local `/opt/stacks/traccar/data` |
| Memos | 104 | `services/memos``/var/opt/memos` | none | |
| Arr-stack (Audiobookshelf, Prowlarr, RDTClient, Shelfarr) + media | 104 | `services/arr-stack/*` + `media/{audiobooks,ebooks,podcasts,Torrents}` | none | already migrated (older doc claimed "still local") |
| ~~Memos~~ | ~~104~~ | ~~`services/memos`~~ | — | **SUNSET 2026-07-10** — service + data + UNAS dir wiped. See [[CHANGELOG]] |
| Arr-stack (Prowlarr, RDTClient, Shelfarr, Audiobookshelf) | 104 | configs on `services/arr-stack/{prowlarr,rdtclient/config,shelfarr/storage,audiobookshelf}`; library + downloads on `media/{audiobooks,ebooks,podcasts,Torrents}` | none | **fully on UNAS incl. SQLite configs** (verified running 2026-07-15). Radarr/Sonarr not deployed. Stale local `/opt/stacks/arr-stack/media` (27 GB pre-migration copy) deleted 2026-07-15. |
| Gluetun (VPN) | 104 | `services/gluetun/data``/gluetun` | none | |
| Gitea | 111 | `services/gitea``/data` | none | new 2026-05-20 |
| Coder workspace home dirs | 111 | `services/coder/<user>/<workspace>``/home/<user>` | none | new 2026-05-20 |
@@ -38,27 +38,36 @@ Tier-1 state and anything that should NOT be NFS-backed.
| Service | CT | Local path → in container | Backed up? | Notes |
|---|---|---|---|---|
| Pocket-ID | **110** | `/opt/stacks/pocketid/data``/app/data` | none | CT 110 has no NFS mount at all. Earlier doc said UNAS — incorrect. |
| Garage (S3 meta + data) | 104 | `/opt/stacks/shared-db/garage/{data,meta}``/var/lib/garage/*` | none | **moved off NFS** 2026-05-19 after WAL-G outage. Stale copy may still exist on UNAS. |
| Garage (S3 meta + data) | 104 | `/opt/stacks/shared-db/garage/{data,meta}``/var/lib/garage/*` | none | **moved off NFS** 2026-05-19 after WAL-G outage. Metadata is **SQLite** (`db_engine = sqlite`) — do NOT put on NFS. Data 82 GB on disk (2026-07-15), almost all WAL-G PG backups — see bucket table below. |
| Shared Postgres (multi-stack) | 104 | docker named volume `shared-db_shared-pgdata` | **WAL-G → local Garage** (configured; outage 2026-05-19 prompted move) | Garage itself is single-disk local — no off-host copy. |
| n8n | 104 | `/opt/stacks/n8n/data``/home/node/.n8n` | none | **back to local** after PG migration attempt failed. Stale 408 MB `database.sqlite` left on UNAS (May 15) — clean up. n8n now uses `shared-postgres` as its actual DB. |
| Arcane | 104 | `/opt/stacks/arcane/data``/app/data` | none | reverted from UNAS to local 2026-05-19 (undocumented before now) |
| Karakeep + Meilisearch | 104 | `/opt/stacks/karakeep/localdata/{data,meili}` | none | reverted from UNAS to local |
| ~~Arcane~~ | ~~104~~ | ~~`/opt/stacks/arcane/data`~~ | — | **DECOMMISSIONED 2026-05-26** (replaced by Portainer). Compose renamed `.DECOMMISSIONED-2026-05-26`; data was local, not UNAS (volumes.md previously said "migrated" — wrong). |
| ~~Karakeep + Meilisearch~~ | ~~104~~ | ~~`/opt/stacks/karakeep/localdata/{data,meili}`~~ | — | **SUNSET 2026-07-10** — service + local + UNAS data wiped. See [[CHANGELOG]] |
| Immich Postgres (pgvector) | 104 | `/opt/stacks/immich/postgres``/var/lib/postgresql/data` | Immich pg_dump → UNAS | tier-1; PG demands local disk |
| Homepage | 104 | `/opt/stacks/homepage/{config,icons}` | none | |
| Homepage | 104 | `/opt/stacks/homepage/{config,icons}` | none | decommissioned — replaced by Homarr on CT 109 |
| Dozzle | 104 | `/opt/stacks/dozzle/dozzle_data` | none | |
| LiteLLM config | 104 | `/opt/stacks/ai/litellm-config` | none | |
| ~~LiteLLM config~~ | ~~104~~ | ~~`/opt/stacks/ai/litellm-config`~~ | — | **SUNSET 2026-05-26** — replaced by Bifrost |
| SearXNG config | 104 | `/opt/stacks/ai/searxng` | none | |
| Flaresolverr | 104 | `/var/lib/flaresolver` | none | |
| AdGuard / Zoraxy / DNS / Shepard / Backrest binaries | 102/108/103/101 | local zfs only | none | Backrest itself has no self-backup |
| Vaultwarden (attachments + key material) | 104 | `/opt/stacks/vaultwarden/data``/data` | none | **moved off NFS 2026-05-22**; DB is on CT 113 postgres; stale `db.sqlite3` deleted |
| Nextcloud config + sidecars | 105 | local zfs (CT rootfs) | none | app data on NFS — see below |
### Garage S3 buckets (on local zfs, CT 104)
### Garage S3 buckets (on local zfs, CT 104) — verified 2026-07-15
| Bucket | Access key | Use | Public URL |
82 GB on disk, dominated by WAL-G Postgres backups. Trim via WAL-G retention on the source hosts, not by moving Garage to NFS.
| Bucket | Logical size | Objects | Use |
|---|---|---|---|
| `lobe-files` | GK55210… (from `.env`) | LobeHub file uploads + WAL-G PG backups | internal only |
| `chat-artifacts` | GK50bfc… | AI chat output artefacts (images, reports, SVG) | `https://chat-artifacts.s3.nuclide.systems/<key>` |
| `ct113-pg-backup` | 66.7 GiB | 1556 | WAL-G for CT 113 shared Postgres fleet |
| `immich-pg-backup` | 36.3 GiB | 4388 | WAL-G / WAL archive for Immich Postgres (+60 MiB unfinished multipart to clean) |
| `ct109-portainer-backup` | 1.9 GiB | 49 | Portainer backups |
| `shared-pg-backup` | 638 MiB | 4986 | shared-postgres WAL-G |
| `chat-artifacts` | — | — | AI chat output artefacts — `https://chat-artifacts.s3.nuclide.systems/<key>` |
| `owui-files` | 0 B | 0 | Open WebUI file storage |
| ~~`lobe-pg-backup`~~ | 151 MiB | 511 | **stale** — LobeChat decommissioned 2026-05-26; safe to delete |
| ~~`lobe-files`~~ | 0 B | 0 | **stale** — LobeChat |
| ~~`memos`~~ | 3.5 MiB | 4 | **orphan** — Memos sunset 2026-07-10; bucket delete blocked by quorum decode error, key revoked, needs Garage repair pass |
### Docker named volumes
@@ -93,6 +102,13 @@ This is a known gap. Plans:
1. Extend Backrest plans to snapshot tier-1 paths (Pocket-ID sqlite, Vaultwarden data, Gitea repos, Coder workspace homes) to jottacloud nightly.
2. Once the second NVMe lands (see `infra/proxmox-state.md` §6), mirror `rpool` so a single disk death doesn't take everything.
## ZFS snapshots (CT 104)
`rpool/data/subvol-104-disk-0` uses `refquota=200G` (not `quota`) — so **snapshots do not count against the container's 200 G limit, but they DO consume pool space**. A stale snapshot can silently exhaust the pool while the container's own `df` looks fine.
- Watch: `zfs list -t snapshot -o name,used rpool/data/subvol-104-disk-0`
- As of the 2026-07-15 incident, `@pre-tier1-backups-20260521-003708` held **97.9 GB** (pinning everything deleted since 2026-05-21, incl. the Karakeep/Memos sunset and docker prunes). Pool free was down to 57 GB. Destroying stale snapshots is the fastest pool-level reclaim.
## Drift cleanup TODO
- [x] ~~Remove stale `services/n8n/database.sqlite` from UNAS~~ — gone (verified 2026-05-22)