Files
docs/ideas/stack-ideas.md
fkrebs ad0fdd2beb docs: dissolve Stacks chapter → infra/; reconcile portmap, storage, changelog (2026-05-21)
- Move PORTMAP, storage, volumes, docker-networks from stacks/ → infra/
- Remove stacks/todo.md (content migrated to CHANGELOG + ideas)
- portmap: add new MCP servers (18005/07/09/10/11), fix s3.nuclide.systems entry,
  update Daytona OIDC row to DECOMMISSIONED, add Vaultwarden OIDC placeholder,
  add DevOps table row for MCP server Gitea repos
- storage: add Garage S3 buckets subsection (lobe-files + chat-artifacts)
- docker-networks: remove daytona-minimal_daytona-network row + daytona from
  "Stacks Keeping Own Networks" list
- ct-inventory: add CT 112 (secrets/Infisical, planned)
- CHANGELOG: prepend 2026-05-21 entry (docs, S3, MCP, dev env, Gitea repos, planned)
- ideas/stack-ideas: add §11 Renovate Bot, §12 egress firewall, §13 local LLM
- mkdocs.yml: replace nav — dissolve Stacks chapter, add infra/ entries
2026-05-21 07:54:21 +02:00

287 lines
13 KiB
Markdown

# Stack Ideas & Improvements
Hardware baseline: **NUC 14 Pro** — Intel Core Ultra (Meteor Lake), Intel Arc iGPU
(currently used for ComfyUI / XPU image gen), no NVIDIA, UNAS for media/bulk storage.
---
## 1. Document ingestion → LobeChat knowledge base + MCP
**Goal**: ingest PDFs, Word, PPT, XLS → searchable in LobeChat chat + accessible via MCP.
Qdrant is a nice-to-have, not a requirement (no live pipeline uses it yet).
**LobeChat already has a built-in knowledge base** (`knowledge_base_files`, `chunks`,
`embeddings` tables in Postgres). The missing piece is an ingest pipeline that feeds it.
**Recommended architecture**:
```
Nextcloud folder / Paperless webhook
→ n8n trigger (already running)
→ Docling (PDF/DOCX/PPTX/XLSX → Markdown + structure)
→ LiteLLM /v1/embeddings (Mistral-embed / codestral-embed)
→ LobeChat knowledge base API ←— available in chat natively
→ Qdrant sidecar (optional, if multi-app search needed)
```
**Document conversion options**:
| Tool | Docker image | Formats | NUC 14 CPU fit |
|------|-------------|---------|----------------|
| **Docling** (IBM, 2024) | `ghcr.io/ds4sd/docling` | PDF, DOCX, PPTX, XLSX, HTML, Markdown | ✅ CPU-only, 10-60s/doc |
| **MinerU** (OpenDataLab) | `opendatalab/mineru` | PDF (layout-aware, OCR) | ✅ CPU mode; GPU optional for speed |
| **Unstructured** | `quay.io/unstructured-io/unstructured-api` | Very broad (25+ formats) | ✅ lighter than Docling |
| **Markitdown** (already in MCP gateway) | — | Office + PDFs → Markdown | ✅ ad-hoc only, not batch |
**Docling vs MinerU**: Docling is better for structured Office/PDF with tables and
figures. MinerU (OpenDataLab) is better for pure PDF optical layout analysis (e.g.,
academic papers, scanned documents). Both run on NUC 14 CPU. Start with Docling — single
container, REST API, well-documented.
**MCP access**: add a `knowledge-search` MCP tool to the gateway that calls LobeChat's
knowledge base search API. Zero new infra — LobeChat is already on `shared_backend`.
**IBM Granite embedding** (granite-embedding-30m-english) — good quality, tiny (30M params),
runs CPU-only. Alternative to Mistral-embed if you want fully local embeddings. Not on
LiteLLM yet but can add as a custom provider pointing at an Ollama/IPEX-LLM instance.
---
## 2. Document conversion pipeline
**Problem**: No systematic ingestion path from raw docs (PDF, DOCX, HTML) into
structured text for chunking/embedding.
**Tools to consider**:
| Tool | Docker image | Best for |
|------|-------------|----------|
| **Docling** (IBM, 2024) | `ghcr.io/ds4sd/docling` | PDFs with complex layout (tables, columns, figures); outputs Markdown |
| **Unstructured** | `quay.io/unstructured-io/unstructured-api` | Wide format support (HTML, DOCX, PPTX, images with OCR); REST API |
| **Gotenberg** | `gotenberg/gotenberg` | HTML/Office → PDF pre-processing stage; not a text extractor |
| **Apache Tika** | `apache/tika` | Broad format support; outputs plain text; lower quality than Docling for PDFs |
| **Markitdown MCP** | already in gateway | Ad-hoc on-demand conversion; not suitable for batch pipelines |
**Recommended pipeline (n8n-orchestrated)**:
```
Nextcloud / Paperless webhook
→ Gotenberg (Office → PDF)
→ Docling (PDF → Markdown chunks)
→ LiteLLM /v1/embeddings (Mistral-embed)
→ Qdrant upsert
```
Docling runs CPU-only comfortably on the NUC. Unstructured is heavier but has a
managed API if the self-hosted version is too slow.
Paperless-ngx already OCRs documents — its full-text content is accessible via the
REST API. An n8n workflow polling `/api/documents/?added__gt=<last_run>` can extract
and re-embed directly, without re-running OCR.
---
## 3. On-device LLM (Arc iGPU)
The Arc iGPU is currently ComfyUI-only. Small language models can also run on it:
- **IPEX-LLM** (Intel) + **Ollama** — Intel provides a patched Ollama build that uses
IPEX-LLM for Arc acceleration. Supports Phi-3.5-mini, Gemma-2B, TinyLlama.
Useful as a fast/cheap fallback when Groq/Mistral rate limits hit.
- **llama.cpp with Vulkan** — alternative to IPEX-LLM; no Intel-specific driver,
uses Vulkan compute. Less optimised for Arc but simpler to deploy.
- **Practical constraint**: Arc iGPU shares system RAM; sustained LLM inference
competes with other XPU workloads (e.g., ComfyUI). A model serving config that
pauses one workload while the other runs is needed, or deploy them on separate
CPU/GPU budgets.
If a second NUC or small GPU box becomes available, this becomes the primary use case.
---
## 4. Arcane agents on NUC (central instance on Proxmox LXC)
Arcane's central server moves to a **Proxmox LXC** (lightweight, stays responsive even
if the Docker stack has issues). The NUC and any other Docker host runs the **Arcane
agent** (headless worker), which connects back to the central instance.
**Deployment**:
- LXC: central Arcane server + observability stack (see §8)
- NUC (`.40`) and `.49`: `arcane-agent` (headless) in compose, connects to LXC
- Arcane agents reach Docker via either:
- **TCP Docker socket + TLS certs** — most secure; requires cert generation per host
- **SSH Docker contexts** — simpler; gateway mounts socket into agent via SSH tunnel
**Observability co-located on the LXC** (see §8 — lightweight, OTEL-aware stack that
survives Docker restarts).
---
## 5. AI memory / personalization
**Mem0** (`mem0ai/mem0`) — persistent user memory layer that sits in front of LLM calls.
Stores facts extracted from conversations into a vector DB (Qdrant backend supported).
Can integrate with LobeChat or as an MCP tool. Lets the AI remember preferences,
past context, and user-specific facts across sessions.
Alternative: **Letta** (formerly MemGPT) — stateful agent framework with persistent
memory; more opinionated.
---
## 6. Workflow / automation upgrades
- **n8n → AI agent nodes**: n8n 1.x has native AI agent nodes (LangChain under the hood).
Can build RAG, email triage, document processing workflows visually.
- **Activepieces** — lighter-weight n8n alternative; good for simple integrations.
Probably not worth adding since n8n is already running.
- **Temporal** — durable workflow engine for long-running or retry-heavy pipelines
(e.g., large document batch processing). Overkill unless pipelines become complex.
---
## 7. External services worth considering
| Service | Purpose | Notes |
|---------|---------|-------|
| **Backblaze B2** | Off-site backup (already in todos) | rclone sync from Garage; restic for PG dumps |
| **Cloudflare R2** | S3-compatible CDN-backed storage | Free egress; good for Lobe file serving |
| **Resend** | Transactional email | 3K/mo free; better deliverability than self-hosted Postal |
| **ntfy.sh (cloud)** | Push notifications fallback | Already running self-hosted ntfy; cloud as relay |
| **Cloudflare Turnstile** | Bot protection for public endpoints | Free; no JS challenge |
| **Novu** | Notification orchestration | Multi-channel (email, push, Slack); self-hostable |
---
## 8. Observability (on Proxmox LXC, alongside Arcane)
**Runs on the LXC, not the NUC** — stays alive if the Docker stack misbehaves. Lightweight
enough for a 2 vCPU / 4GB LXC.
**Recommended stack** (all OTEL-aware, compose-based):
| Component | Image | Role |
|-----------|-------|------|
| **OpenTelemetry Collector** | `otel/opentelemetry-collector-contrib` | Receives traces/metrics/logs from all services (OTLP gRPC+HTTP); fans out to backends |
| **VictoriaMetrics** | `victoriametrics/victoria-metrics` | Prometheus-compatible TSDB; scrapes NUC exporters + receives from OTEL collector; lighter than Prometheus |
| **Grafana** | `grafana/grafana` | Dashboards; datasource = VictoriaMetrics + Loki |
| **Loki** | `grafana/loki` | Log aggregation; receives from OTEL collector |
| **Uptime Kuma** | `louislam/uptime-kuma` | HTTP/TCP uptime checks for all public endpoints; alerts via ntfy |
**OTEL receivers** from the NUC Docker stack:
- **LiteLLM**: native OTEL export — traces every LLM call with token counts, model, latency
- **n8n**: Prometheus `/metrics` endpoint
- **Garage**: Prometheus `/metrics`
- **mcp-gateway**: add `opentelemetry-sdk` instrumentation to `server.py`
- **Docker host**: `node_exporter` + `cadvisor` on the NUC, scraped by VictoriaMetrics
**NUC → LXC connectivity**: Both are on the same LAN. OTEL collector listens on the LXC's
LAN IP (e.g., `192.168.1.X:4317` gRPC). Services push OTEL directly to it; Prometheus
pull-scraping from VictoriaMetrics goes to NUC exporters over LAN.
---
## 9. Storage / S3 improvements
- **Garage external S3** is currently non-responsive (tracked in todos). Fix or
replace with **MinIO** (single-node, much better tooling/UI, easy migration).
- **Lobe file serving via Cloudflare R2**: Lobe uploads to Garage → sync to R2 →
serve from CDN. Reduces NUC egress for image-heavy sessions.
- **Qdrant data on UNAS**: Mount `qdrant_data` from NAS for persistence across
container re-creates (same pattern as arr-stack).
---
## 9. ComfyUI MCP — async queue + image-to-image + reference
**Current problem**: MCP tool blocks until the image is done (~90-290s). LLMs time out at ~60s.
The fix is a job-queue pattern:
```
generate_image(prompt) → returns {job_id, status: "queued"} immediately
get_image_status(job_id) → returns {status, progress, image_url_when_done}
list_queue() → shows all pending/running jobs
cancel_job(job_id) → cancels a queued job (confirms with user first)
```
ComfyUI's own API is already async (`POST /prompt` → poll `/history/{id}`). The MCP server
just needs to expose this model instead of blocking.
**Image-to-image**: new workflow `flux-schnell-img2img-api.json`. Takes an init image URL +
denoise strength. ComfyUI `LoadImageFromURL` node (or upload + `LoadImage`).
**Reference tool**: `list_previous_images(n=5)` — queries ComfyUI `/history` API,
returns recent job thumbnails + prompts. User can pick one to reference or iterate from.
**Delivery to S3/Garage**: on completion, upload output PNG to Garage `comfyui-outputs` bucket
→ return a permanent URL. Avoids ComfyUI's ephemeral `/view` endpoint.
---
## 10. LiteLLM custom provider / LobeChat provider extension
**Question**: build an adapter in LiteLLM to use models via "quasi API" (non-standard
endpoints, auth, or routing)?
**LiteLLM supports custom providers** via `custom_llm_provider` in config.yaml. You write
a Python class that implements `completion()` and `async_completion()`. This is the right
path for wrapping non-standard APIs (local models, proprietary endpoints, protocol bridges).
Example use cases:
- Wrap a Claude Max subscription via Anthropic's API (different billing model)
- Add a local model served by IPEX-LLM on the Arc iGPU
- Bridge a custom inference server that speaks a different protocol
**LobeChat provider extension**: LobeChat's provider list is compiled into the app.
Adding a new provider requires rebuilding LobeChat from source (fork + add to
`src/config/aiModels/`). High effort; only worth it for a permanent/long-term provider.
For ad-hoc needs, use LiteLLM as the adapter and point LobeChat at it via the existing
`ai.nuclide.systems` OpenAI-compatible endpoint.
---
## 11. Dependency updates via Renovate Bot (Gitea Actions)
Run **Renovate Bot** as a Gitea Actions workflow to automatically open PRs for outdated
dependencies in MCP server repos (`mcp-comfyui`, `mcp-docling`, `mcp-shepard`,
`mcp-upload-artifact`). Targets: `Dockerfile` base image tags + `requirements.txt` /
`pyproject.toml` Python deps.
Renovate supports Gitea natively via `platform: gitea` in `renovate.json`. The Gitea
Actions runner (`ct111-runner`) already exists; add a scheduled workflow calling
`renovate/renovate` Docker image once daily.
---
## 12. Egress firewall on UDM (UniFi)
Enforce outbound allow-list on the UDM / UniFi gateway — block all non-approved egress
by default. Goals:
- Prevent exfiltration from compromised containers
- Audit unexpected outbound connections (model providers, analytics, telemetry)
- Approved: LiteLLM model provider endpoints, Jottacloud, UNAS internal, NTP, DNS
Implementation: UDM firewall rules (WAN_OUT) + Threat Management IDS in monitor mode first.
---
## 13. Local LLM on Arc GPU via vllm / ollama for sensitive workloads
Run a privacy-sensitive LLM locally on the Arc iGPU using **vllm** (with XPU/IPEX backend)
or the Intel-patched **Ollama** build. Use cases: document classification in Paperless
workflows, offline coding assistant, fallback when cloud rate limits hit.
See §3 (On-device LLM) for implementation notes. This item tracks the specific motivation
of *sensitive workload isolation* — i.e., running prompts that should not leave the LAN.
---
## Priority order (rough)
1. **ComfyUI MCP async queue** — fixes timeout; unblocks img2img + reference features
2. **Docling container + n8n ingest pipeline** — Nextcloud/Paperless → LobeChat knowledge base
3. **Arcane LXC + agents on NUC** — when ready to migrate
4. **Observability LXC** — OTEL collector + VictoriaMetrics + Grafana + Uptime Kuma, co-located
5. **Backblaze B2 off-site backup** — already in todos
6. **On-device LLM (Arc)** — depends on IPEX-LLM + Ollama Intel build stability
7. **Renovate Bot** — low-effort automation win for MCP repos
8. **Egress firewall** — security hygiene; plan before adding more external-facing services