docs: reconcile storage.md + volumes.md with live 2026-07-15 audit
Fixes contradictions found during the CT 104 disk-full incident: - Arcane: was DECOMMISSIONED 2026-05-26 + local, not "migrated to UNAS" (volumes.md and storage.md disagreed). - arr-stack: fully on UNAS incl. SQLite configs on services/arr-stack/; media/downloads on media/ share. Radarr/Sonarr not deployed. Removed stale "plan to migrate" rows pointing at local /opt/stacks/arr-stack. - n8n: local (DB on CT 113), not UNAS. - Garage: metadata is SQLite (not LMDB); keep off NFS; refreshed bucket table with real sizes (82 GB, mostly WAL-G PG backups). - Marked sunsets: karakeep, memos, litellm, daytona. - crawl4ai: correct name mcp-crawl4ai, no mounts, cache-growth TODO. - Added ZFS snapshot note: refquota hides snapshot bloat from container df; pre-tier1-backups snapshot held 97.9 GB.
This commit is contained in:
+29
-13
@@ -1,7 +1,7 @@
|
||||
# Storage layout
|
||||
|
||||
> Verified state as of 2026-05-21. Source of truth for which service lives on which storage class.
|
||||
> Earlier revisions of this doc were stale on multiple items (Garage, n8n, Arcane, Karakeep, Pocket-ID, arr-stack, Nextcloud's storage protocol, missing Coder/Gitea). Reconciled via live audit.
|
||||
> Verified state as of 2026-07-15 (live audit during the CT 104 disk-full incident). Source of truth for which service lives on which storage class.
|
||||
> Earlier revisions were stale on Garage, n8n, Arcane, Karakeep/Memos, Pocket-ID, arr-stack, Nextcloud's storage protocol, and missing Coder/Gitea. `volumes.md` was reconciled in the same pass.
|
||||
|
||||
## Storage classes
|
||||
|
||||
@@ -24,8 +24,8 @@ Bulk + media + non-latency-sensitive app data.
|
||||
| Paperless-ngx (documents) | 104 | `media/documents/public/paperless-ngx/{consume,export,library}` | none | |
|
||||
| Paperless-AI | 104 | `services/paperless-ai` → `/app/data` | none | |
|
||||
| Traccar (logs, config) | 104 | `services/traccar/{logs,traccar.xml}` | none | data dir reverted to local `/opt/stacks/traccar/data` |
|
||||
| Memos | 104 | `services/memos` → `/var/opt/memos` | none | |
|
||||
| Arr-stack (Audiobookshelf, Prowlarr, RDTClient, Shelfarr) + media | 104 | `services/arr-stack/*` + `media/{audiobooks,ebooks,podcasts,Torrents}` | none | already migrated (older doc claimed "still local") |
|
||||
| ~~Memos~~ | ~~104~~ | ~~`services/memos`~~ | — | **SUNSET 2026-07-10** — service + data + UNAS dir wiped. See [[CHANGELOG]] |
|
||||
| Arr-stack (Prowlarr, RDTClient, Shelfarr, Audiobookshelf) | 104 | configs on `services/arr-stack/{prowlarr,rdtclient/config,shelfarr/storage,audiobookshelf}`; library + downloads on `media/{audiobooks,ebooks,podcasts,Torrents}` | none | **fully on UNAS incl. SQLite configs** (verified running 2026-07-15). Radarr/Sonarr not deployed. Stale local `/opt/stacks/arr-stack/media` (27 GB pre-migration copy) deleted 2026-07-15. |
|
||||
| Gluetun (VPN) | 104 | `services/gluetun/data` → `/gluetun` | none | |
|
||||
| Gitea | 111 | `services/gitea` → `/data` | none | new 2026-05-20 |
|
||||
| Coder workspace home dirs | 111 | `services/coder/<user>/<workspace>` → `/home/<user>` | none | new 2026-05-20 |
|
||||
@@ -38,27 +38,36 @@ Tier-1 state and anything that should NOT be NFS-backed.
|
||||
| Service | CT | Local path → in container | Backed up? | Notes |
|
||||
|---|---|---|---|---|
|
||||
| Pocket-ID | **110** | `/opt/stacks/pocketid/data` → `/app/data` | none | CT 110 has no NFS mount at all. Earlier doc said UNAS — incorrect. |
|
||||
| Garage (S3 meta + data) | 104 | `/opt/stacks/shared-db/garage/{data,meta}` → `/var/lib/garage/*` | none | **moved off NFS** 2026-05-19 after WAL-G outage. Stale copy may still exist on UNAS. |
|
||||
| Garage (S3 meta + data) | 104 | `/opt/stacks/shared-db/garage/{data,meta}` → `/var/lib/garage/*` | none | **moved off NFS** 2026-05-19 after WAL-G outage. Metadata is **SQLite** (`db_engine = sqlite`) — do NOT put on NFS. Data 82 GB on disk (2026-07-15), almost all WAL-G PG backups — see bucket table below. |
|
||||
| Shared Postgres (multi-stack) | 104 | docker named volume `shared-db_shared-pgdata` | **WAL-G → local Garage** (configured; outage 2026-05-19 prompted move) | Garage itself is single-disk local — no off-host copy. |
|
||||
| n8n | 104 | `/opt/stacks/n8n/data` → `/home/node/.n8n` | none | **back to local** after PG migration attempt failed. Stale 408 MB `database.sqlite` left on UNAS (May 15) — clean up. n8n now uses `shared-postgres` as its actual DB. |
|
||||
| Arcane | 104 | `/opt/stacks/arcane/data` → `/app/data` | none | reverted from UNAS to local 2026-05-19 (undocumented before now) |
|
||||
| Karakeep + Meilisearch | 104 | `/opt/stacks/karakeep/localdata/{data,meili}` | none | reverted from UNAS to local |
|
||||
| ~~Arcane~~ | ~~104~~ | ~~`/opt/stacks/arcane/data`~~ | — | **DECOMMISSIONED 2026-05-26** (replaced by Portainer). Compose renamed `.DECOMMISSIONED-2026-05-26`; data was local, not UNAS (volumes.md previously said "migrated" — wrong). |
|
||||
| ~~Karakeep + Meilisearch~~ | ~~104~~ | ~~`/opt/stacks/karakeep/localdata/{data,meili}`~~ | — | **SUNSET 2026-07-10** — service + local + UNAS data wiped. See [[CHANGELOG]] |
|
||||
| Immich Postgres (pgvector) | 104 | `/opt/stacks/immich/postgres` → `/var/lib/postgresql/data` | Immich pg_dump → UNAS | tier-1; PG demands local disk |
|
||||
| Homepage | 104 | `/opt/stacks/homepage/{config,icons}` | none | |
|
||||
| Homepage | 104 | `/opt/stacks/homepage/{config,icons}` | none | decommissioned — replaced by Homarr on CT 109 |
|
||||
| Dozzle | 104 | `/opt/stacks/dozzle/dozzle_data` | none | |
|
||||
| LiteLLM config | 104 | `/opt/stacks/ai/litellm-config` | none | |
|
||||
| ~~LiteLLM config~~ | ~~104~~ | ~~`/opt/stacks/ai/litellm-config`~~ | — | **SUNSET 2026-05-26** — replaced by Bifrost |
|
||||
| SearXNG config | 104 | `/opt/stacks/ai/searxng` | none | |
|
||||
| Flaresolverr | 104 | `/var/lib/flaresolver` | none | |
|
||||
| AdGuard / Zoraxy / DNS / Shepard / Backrest binaries | 102/108/103/101 | local zfs only | none | Backrest itself has no self-backup |
|
||||
| Vaultwarden (attachments + key material) | 104 | `/opt/stacks/vaultwarden/data` → `/data` | none | **moved off NFS 2026-05-22**; DB is on CT 113 postgres; stale `db.sqlite3` deleted |
|
||||
| Nextcloud config + sidecars | 105 | local zfs (CT rootfs) | none | app data on NFS — see below |
|
||||
|
||||
### Garage S3 buckets (on local zfs, CT 104)
|
||||
### Garage S3 buckets (on local zfs, CT 104) — verified 2026-07-15
|
||||
|
||||
| Bucket | Access key | Use | Public URL |
|
||||
82 GB on disk, dominated by WAL-G Postgres backups. Trim via WAL-G retention on the source hosts, not by moving Garage to NFS.
|
||||
|
||||
| Bucket | Logical size | Objects | Use |
|
||||
|---|---|---|---|
|
||||
| `lobe-files` | GK55210… (from `.env`) | LobeHub file uploads + WAL-G PG backups | internal only |
|
||||
| `chat-artifacts` | GK50bfc… | AI chat output artefacts (images, reports, SVG) | `https://chat-artifacts.s3.nuclide.systems/<key>` |
|
||||
| `ct113-pg-backup` | 66.7 GiB | 1556 | WAL-G for CT 113 shared Postgres fleet |
|
||||
| `immich-pg-backup` | 36.3 GiB | 4388 | WAL-G / WAL archive for Immich Postgres (+60 MiB unfinished multipart to clean) |
|
||||
| `ct109-portainer-backup` | 1.9 GiB | 49 | Portainer backups |
|
||||
| `shared-pg-backup` | 638 MiB | 4986 | shared-postgres WAL-G |
|
||||
| `chat-artifacts` | — | — | AI chat output artefacts — `https://chat-artifacts.s3.nuclide.systems/<key>` |
|
||||
| `owui-files` | 0 B | 0 | Open WebUI file storage |
|
||||
| ~~`lobe-pg-backup`~~ | 151 MiB | 511 | **stale** — LobeChat decommissioned 2026-05-26; safe to delete |
|
||||
| ~~`lobe-files`~~ | 0 B | 0 | **stale** — LobeChat |
|
||||
| ~~`memos`~~ | 3.5 MiB | 4 | **orphan** — Memos sunset 2026-07-10; bucket delete blocked by quorum decode error, key revoked, needs Garage repair pass |
|
||||
|
||||
### Docker named volumes
|
||||
|
||||
@@ -93,6 +102,13 @@ This is a known gap. Plans:
|
||||
1. Extend Backrest plans to snapshot tier-1 paths (Pocket-ID sqlite, Vaultwarden data, Gitea repos, Coder workspace homes) to jottacloud nightly.
|
||||
2. Once the second NVMe lands (see `infra/proxmox-state.md` §6), mirror `rpool` so a single disk death doesn't take everything.
|
||||
|
||||
## ZFS snapshots (CT 104)
|
||||
|
||||
`rpool/data/subvol-104-disk-0` uses `refquota=200G` (not `quota`) — so **snapshots do not count against the container's 200 G limit, but they DO consume pool space**. A stale snapshot can silently exhaust the pool while the container's own `df` looks fine.
|
||||
|
||||
- Watch: `zfs list -t snapshot -o name,used rpool/data/subvol-104-disk-0`
|
||||
- As of the 2026-07-15 incident, `@pre-tier1-backups-20260521-003708` held **97.9 GB** (pinning everything deleted since 2026-05-21, incl. the Karakeep/Memos sunset and docker prunes). Pool free was down to 57 GB. Destroying stale snapshots is the fastest pool-level reclaim.
|
||||
|
||||
## Drift cleanup TODO
|
||||
|
||||
- [x] ~~Remove stale `services/n8n/database.sqlite` from UNAS~~ — gone (verified 2026-05-22)
|
||||
|
||||
Reference in New Issue
Block a user