Files
docs/stacks/todo.md
T

10 KiB
Raw Blame History

Todo

Goal

Consolidate databases, migrate storage to UNAS, integrate all services with Homepage.


Ideas (user)

  • Claude Max sub integration with litellm (how to handle auth, rate limits, etc.?)
  • MCP deployment process ideas? Current setup seems cumbersome — can litellm host directly? (how to integrate npx, uvx, and maybe docker-based containers?)
  • Select offsite backup hosting provider, set up backup process (maybe restic to Backblaze B2 or similar), then set up rclone for off-site sync of Garage buckets to remote S3 (Backblaze B2, Jottacloud, or Hetzner)
  • Research the LiteLLM MCP source_url column problem across GitHub issues, PRs, and docs to understand the full picture
  • Add Google Gemini (vision + 1M ctx + free text-embedding-004; needs GEMINI_API_KEY)
  • Add fal.ai MCP server (mcp/fal) for fast cloud image gen < 5 s (vs ~150 s local ComfyUI)
  • Add Cerebras (Llama 3 at ~1M tok/s, generous free tier; CEREBRAS_API_KEY)

Next Steps / Todos

  • ComfyUI added & fixed (2026-05-17) — root-caused the self-referential bind-mount, repaired mounts/healthcheck, migrated CPU→Intel Arc iGPU (torch 2.7.1+xpu, no IPEX), pinned v0.21.1. FLUX.1-schnell GGUF, legible text, ~94 s/image. Canonical workflow in ai/comfyui/workflows/.
  • ComfyUI→LobeChat via MCP (2026-05-17) — purpose-built comfyui-mcp server, proven end-to-end. See docs/comfyui-mcp.md.
  • MCP gateway finished (2026-05-17) — DinD, real streaming MCP reverse-proxy, Pocket ID bearer + RFC9728 discovery, OAuth callback fixed, secrets from .env, no port collisions. See docs/mcp-gateway.md.
  • LobeChat S3 images fixed (2026-05-17) — added S3_PUBLIC_DOMAIN=http://192.168.1.40:10004 to lobehub.yml; pre-signed URLs now use the LAN-accessible Garage port instead of the internal garage:3900 hostname.
  • MCP gateway groups + combined endpoints (2026-05-17) — servers now grouped (research/dev/personal/image); /group/{name}/mcp aggregates all member tools with server__tool namespacing; UI shows grouped cards with copy-URL buttons; //ui redirect added.
  • LiteLLM Redis caching (2026-05-17)cache: true with Redis backend (lobe-redis, already on ai-internal); 10-min TTL on completion + embedding calls.
  • Cohere added to LiteLLM (2026-05-17)command-a-03-2025 (20 RPM), cohere-embed-english/multilingual/light-v3 (2000 RPM); COHERE_API_KEY in .env. Note: command-r / command-r-plus retired 2025-09-15.
  • LobeChat MCP import donecomfyui-flux server imported via UI.
  • ComfyUI MCP async queue (2026-05-17) — full rewrite: generate_image returns job_id in <0.1 s; background polling thread updates state; img2img accepts many images (URL/base64/job:<id> chaining); get_job_status, list_queue, cancel_job, list_recent_images. New flux-img2img-api.json workflow (VAEEncode path). See docs/comfyui-mcp.md.
  • Nextcloud MCP: bearer token enabled in NC admin (user_oidc provider settings). Gateway now provisions NC app-password at SSO login time (PocketID token live). Needs end-to-end test: log in via mcp.nuclide.systems, then call /nextcloud/mcp.
  • Observability LXC: deploy OTEL collector + VictoriaMetrics + Grafana + Loki
    • Uptime Kuma on a Proxmox LXC; see docs/stack-ideas.md §8.
  • Arcane: migrate central server to Proxmox LXC; deploy arcane-agent (headless) on NUC + .49; connect via TCP Docker socket or SSH Docker context.
  • Docling + n8n ingest pipeline: Nextcloud/Paperless → Docling → LiteLLM embed → LobeChat knowledge base; see docs/stack-ideas.md §1.
  • Gateway spawn servers (2026-05-17) — papersearch, crawl4ai, markitdown all running on ai-internal. Fixed papersearch command (uvx --with not uv tool run --from). nextcloud deferred: Nextcloud OIDC discovery not enabled; needs NEXTCLOUD_USERNAME/NEXTCLOUD_PASSWORD in ai/.env for BasicAuth mode (server.py already updated).
  • ComfyUI inpaint smoke-tested (2026-05-17) — Kodak #23 (768×512) + center alpha-mask, submitted via API, output in 290s. Workflow confirmed working end-to-end.
  • garage external S3 workings3.nuclide.systems returns 403 (AccessDenied on root, expected); Zoraxy routes correctly to 10004→3900. Was already fixed by the earlier port binding change (0.0.0.0:10004); todo was stale.
  • syncstack gone — container no longer present; one-shot job completed/pruned.
  • immich_machine_learning healthcheck fixed — already on :3003/ping. Container shows healthy.
  • LiteLLM healthcheck fixed — was hitting host port 14000 inside container; corrected to container port 4000. Now healthy.
  • immich_power_tools / karakeep / karakeep_chrome — all healthy (slow warmup self-resolved).

In Progress

  • None

Blocked

  • None

Reference / History

Done

  • Created /opt/stacks/shared-db/ with shared-postgres (pg16), Garage S3, pgAdmin
  • Consolidated daytona, litellm, paperless-ngx → shared-postgres
  • LobeHub kept on own pg17 (ParadeDB)
  • Switched LobeHub from RustFS → Garage S3 (lobe-files bucket)
  • Migrated arr-stack configs + media to UNAS via rsync
  • Migrated qdrant to UNAS (fresh start)
  • Added 29 selfh.st icons, updated all homepage labels to sh- prefix
  • Homepage: instance hrefs (192.168.1.40:PORT) + public hrefs (*.nuclide.systems)
  • Homepage widgets for 11 services
  • Groups consolidated from 14 → 6 (infrastructure, security, media, ai, productivity, diagnostic)
  • Configured memos with Garage S3
  • Fixed Garage port binding: 127.0.0.1:100040.0.0.0:10004 (Zoraxy couldn't reach it)
  • Fixed memos S3 endpoint: https://s3.nuclide.systemshttp://garage:3900 (TLS cert mismatch)
  • Cleaned up old local data remnant at /opt/stacks/memos/data/
  • traccar GPS port fixed: 15001 → 19001 (paperless-ai port conflict)
  • Updated PORTMAP.md, created storage.md and volumes.md
  • WAL-G v3.0.8 setup for all 3 PG instances (shared-postgres, immich, lobe)
  • Garage S3 buckets + per-instance keys for PG backups
  • WAL archiving enabled and flowing to Garage (wal-g wal-push)
  • Daily full backup cron at 3am (/etc/cron.d/pg-backup)
  • Home Assistant added to Homepage (infrastructure, widget token)
  • ShelfArr added to Homepage (diagnostic, MDI icon)

Proxy migration → Zoraxy + clean-state recreate — audit 2026-05-16

Done

  • Reverse-proxy plan moved Traefik → Zoraxy v3.3.2 (192.168.1.4:8000, noauth, internal). Traefik abandoned: it runs on a different host than the containers, so Docker labels can't work cross-host; Zoraxy already had the wildcard *.nuclide.systems cert + rules.
  • Source of truth: proxy/zoraxy/routes.json + idempotent CSRF-aware scripts/zoraxy_sync.py (--dry-run default / --apply / --prune / --list). Self-service: edit the JSON, run --apply.
  • Zoraxy reconciled: fixed ai→14000 (litellm), mcp→8080 (mcp-gateway), added chat→14001 (lobehub). Pruned (not exposed): dozzle, immich-tools, mcp-auth. Internal-only: pgadmin, paperless, paperless-ai.
  • Fixed AI URL configs (litellm.nuclide.systems had no route): ai/litellm-config/config.yaml, ai/.env, karakeep/.envai.nuclide.systems; paperless_ai/.env CUSTOM_BASE_URL mcpai.
  • paperless-ngx host port 15000:8000 published (internal LAN); paperless-ai 15001 already in compose.
  • Stripped all traefik.* labels from every compose file; fixed memos YAML escaping.
  • Cleaned proxy/ (removed dead Traefik artifacts) + added proxy/zoraxy/README.md.
  • shared-db/docker-compose.yml restored from HEAD (WIP had deleted shared-postgres
    • GARBAGE_ typos) and force-recreated cleanly.
  • Removed stale *_old containers; clean-state force-recreate of all 19 stacks.

Verified working (Zoraxy end-to-end, https via .4)

  • ai, chat, mcp, arcane, ntfy, id, vault, abs, n8n, memos, hoarder, daytona, traccar — all serving.
  • Static/external: ha, nc, ocpp, shepard, shepard-api — reachable.
  • shared-postgres healthy; litellm/lobehub/mcp-gateway healthy & proxied; garage S3 working internally (active wal-g WAL backups).

Issues found → fix outcomes (2026-05-16)

  • FIXED — immich_server crash-loop (getaddrinfo ENOTFOUND database): the database service was on immich_immich_default+shared_backend while immich-server/redis/ml were on the implicit immich_default → no shared net. Fix: immich/docker-compose.yml top-level network is now default: {name: immich_default} and database joins [default, shared_backend]. Verified: immich_server healthy, immich.nuclide.systems200 via Zoraxy.
  • FIXED — vpn_gluetun "unhealthy": removed the bogus custom healthcheck (wget :8000, 404s) so gluetun's built-in healthcheck (127.0.0.1:9999) runs. Verified: vpn_gluetun healthy; arr-stack recreated (it uses network_mode: container:vpn_gluetun) — prowlarr/rdtclient/shelfarr/flaresolverr healthy.
  • NOT fixed — garage external S3 (s3.nuclide.systems): tried api_bind_addr/rpc_bind_addr [::]0.0.0.0, but applying it needs a force-recreate, and garage's single-node cluster is fragile to restarts — the recreate caused "could not reach quorum" (RPC). Reverted garage.toml to HEAD and recreated; garage is back to the known-good state (HEALTHY node, layout v1, buckets OK, internal S3 / PG WAL backups working). External s3.nuclide.systems remains non-responsive as it was before — tracked in Todos above.
  • nuc-ai-core-syncstack-1 left Created — tracked in Todos above.

Other (pre-existing, NOT in the fix-1/2/3 scope, surfaced by clean recreate)

  • immich_machine_learning shows "unhealthy" but works — its healthcheck tests localhost:3000/health while ML serves :3003 (immich_server logs confirm "Machine learning server became healthy"). Cosmetic healthcheck-port mismatch in the immich compose.
  • immich_power_tools, karakeep, karakeep_chrome showing "unhealthy" after the clean recreate — slow/long healthcheck warmup; revisit if still unhealthy later.

Notes

  • Zoraxy self-service: python3 scripts/zoraxy_sync.py [--apply|--prune|--list].
  • The shared-db WIP (garage/postgres rework) was reverted to HEAD — redo it intentionally if that work was wanted.