6.2 KiB
Case Study — The Nuclide Homelab, built with Claude
Origin story
One Saturday, the owner's wife left him home alone. He got bored, subscribed to
Claude, and started tinkering with a home server. That afternoon of boredom
turned into the /opt/stacks ecosystem documented here — a ~66-container,
~23-stack self-hosted platform with SSO, an MCP/agent gateway, GPU offload,
and a fully audited network. This is that story, kept as a record of what a
curiosity-driven collaboration produced.
This is a personal passion project, not a work deliverable. The tone and scope reflect that: depth and exploration over minimum-viable.
What was built (high level)
- AI/agent core — LiteLLM model gateway (~28 curated models), an MCP gateway that DinD-spawns and OAuth-gates MCP servers, an Agent Operator (cron/event agents), semantic tool retrieval, MCP→LobeChat registration.
- Identity — PocketID as the universal OIDC IdP; every service behind SSO (LiteLLM, LobeChat, n8n, Nextcloud, Daytona via a Keycloak→PocketID adapter).
- Data layer — shared-postgres standard; a full tier-1 "no SQLite on NFS" migration (ntfy, Karakeep, Arcane, traccar → local disk; Memos + Vaultwarden → Postgres via pgloader); WAL-G PITR backups re-established.
- Network — full read-only UniFi audit + the unmanaged D-Link DGS-1210 core (SNTP fixed, topology/FDB mapped, DHCP/LLDP interop verified, SPOF identified); pinned-NTP-source standard adopted.
- Reliability — gateway deep health-check (real MCP tool-call probe, usage-aware backoff) that caught silently-broken servers.
- Cloud burst — on-demand Scaleway L40S GPU offload via WireGuard.
See homelab-architecture.md for the living technical reference and
PORTMAP.md for the authoritative port/route map.
Activity signal
- ~79 commits in the trailing 14 days (
git log --since="14 days ago"), spanning gateway OAuth/health, OIDC bolt-ons, DB migrations, network audit, GPU integration, and docs. - Multi-session, incident-driven: several entries trace to real failures
caught and fixed (WAL-G archiver hung silently ~13 h; n8n
latest-drift outage; SQLite-on-NFS corruption risk; an 8.5-month-stale switch clock).
Productivity estimate (honest framing)
These are rough order-of-magnitude estimates, not measurements. Assumptions are stated so they can be challenged.
- Scope delivered ≈ a small platform: identity, AI gateway, agent runtime, ~23 service stacks, a DB-migration program, a full network audit, backups.
- A solo engineer doing this unaided, part-time, learning the unfamiliar pieces (OIDC internals, pgloader, MCP, UniFi/D-Link internals): a conservative bound is several hundred focused hours (≈ 8–16 part-time weeks). Assumes the owner is competent but not a specialist in every domain touched (identity, Postgres ops, embedded-switch web UIs, MCP).
- With the assistant: compressed into a small number of intensive sessions over ~2 weeks. The leverage is largest where the work is research-heavy but low-novelty — reverse-engineering a D-Link form POST, deriving an OIDC redirect, mapping an FDB table — i.e. tasks that are tedious solo but not conceptually hard. The leverage is smallest on genuine judgment calls (what to prioritise, what risk is acceptable), which stayed with the owner.
- Caveat: estimate excludes the owner's own steering/review time, which was substantial and is the reason the output is coherent rather than just voluminous.
CO2 estimate (honest framing)
Also order-of-magnitude, assumptions explicit.
- LLM inference: a heavy multi-session collaboration of this kind is on the order of a few million tokens. Public estimates put frontier-model inference at roughly 1–5 Wh per ~1k output tokens equivalent (wide error bars). Taking ~3 M tokens × ~2 Wh/1k ≈ ~6 kWh → at a ~0.35 kgCO2e/kWh grid ≈ ~2 kgCO2e. Plausible range 1–5 kgCO2e. Datacenter PUE and exact model size dominate the uncertainty.
- Homelab runtime (the larger ongoing footprint): the NUC 14 Pro draws ~15–45 W under mixed load. At ~30 W average → ~0.72 kWh/day → ~260 kWh/yr → ~90 kgCO2e/yr at the same grid factor. The German grid is cleaner than that average in many hours, so treat as an upper-ish bound. The on-demand Scaleway L40S burst is deliberately on-demand precisely to avoid a 24/7 GPU's footprint.
- Takeaway: the assistant-collaboration carbon is a rounding error next to a year of always-on homelab power. Efficiency wins (CPU-only default, on-demand GPU, idle-aware health probing) matter more than the chat cost.
Handover / current state
Healthy & verified
- Tier-1 SQLite-off-NFS: complete.
- OIDC: n8n (302→PocketID, client
33135ad4) and LobeChat (AUTH_TRUSTED_ ORIGINSfix, sign-in→PocketID) — both verified; LobeChat wants one real browser login as final proof. - D-Link SNTP: fixed (pinned PTB+Cloudflare IPs, clock corrected & synced).
- Gateway deep health-check: live, usage-aware, surfaced in
/api/servers.
Open / pending (see homelab-architecture.md roadmap for detail)
- Broken MCP servers surfaced by the new health-check:
memos(degraded —mcp-memoscan't resolvememoshost; Docker-network isolation),context7/crawl4ai/markitdown(down),nextcloud(probe false-positive — needshealth_check:falseor per-user creds). - D-Link mgmt hardening (bundle, confirm-first): HTTPS, SNMP review,
Trusted-Host allowlist
192.168.1.0/24. Sharedtapirnasepassword reuse (WiFi/LiteLLM/switch) — rotation deferred, noted. - Network: IoT-VLAN segmentation; D-Link is the unmanaged core/SPOF; mgmt-TLS certs for Proxmox + D-Link.
- Platform: env→secret vault; LobeChat external-feature disable; observability LXC; agent-platform evolution (memory/teams/MCP-exposed).
nexaanalysis blocked — private repo; deploy key pending authorization.
Operating rules to preserve
- Confirm + risk-assess before any Proxmox / Ubiquiti / network-gear write.
- Never put DB/SQLite on the UNAS NFS share.
- Only a full pgloader of all tables is a complete DB migration.
- Prefer self-hosted; pin critical container images (no
latestdrift).