249 lines
12 KiB
Markdown
249 lines
12 KiB
Markdown
# Stack Ideas & Improvements
|
|
|
|
Hardware baseline: **NUC 14 Pro** — Intel Core Ultra (Meteor Lake), Intel Arc iGPU
|
|
(currently used for ComfyUI / XPU image gen), no NVIDIA, UNAS for media/bulk storage.
|
|
|
|
---
|
|
|
|
## 1. Document ingestion → LobeChat knowledge base + MCP
|
|
|
|
**Goal**: ingest PDFs, Word, PPT, XLS → searchable in LobeChat chat + accessible via MCP.
|
|
Qdrant is a nice-to-have, not a requirement (no live pipeline uses it yet).
|
|
|
|
**LobeChat already has a built-in knowledge base** (`knowledge_base_files`, `chunks`,
|
|
`embeddings` tables in Postgres). The missing piece is an ingest pipeline that feeds it.
|
|
|
|
**Recommended architecture**:
|
|
```
|
|
Nextcloud folder / Paperless webhook
|
|
→ n8n trigger (already running)
|
|
→ Docling (PDF/DOCX/PPTX/XLSX → Markdown + structure)
|
|
→ LiteLLM /v1/embeddings (Mistral-embed / codestral-embed)
|
|
→ LobeChat knowledge base API ←— available in chat natively
|
|
→ Qdrant sidecar (optional, if multi-app search needed)
|
|
```
|
|
|
|
**Document conversion options**:
|
|
|
|
| Tool | Docker image | Formats | NUC 14 CPU fit |
|
|
|------|-------------|---------|----------------|
|
|
| **Docling** (IBM, 2024) | `ghcr.io/ds4sd/docling` | PDF, DOCX, PPTX, XLSX, HTML, Markdown | ✅ CPU-only, 10-60s/doc |
|
|
| **MinerU** (OpenDataLab) | `opendatalab/mineru` | PDF (layout-aware, OCR) | ✅ CPU mode; GPU optional for speed |
|
|
| **Unstructured** | `quay.io/unstructured-io/unstructured-api` | Very broad (25+ formats) | ✅ lighter than Docling |
|
|
| **Markitdown** (already in MCP gateway) | — | Office + PDFs → Markdown | ✅ ad-hoc only, not batch |
|
|
|
|
**Docling vs MinerU**: Docling is better for structured Office/PDF with tables and
|
|
figures. MinerU (OpenDataLab) is better for pure PDF optical layout analysis (e.g.,
|
|
academic papers, scanned documents). Both run on NUC 14 CPU. Start with Docling — single
|
|
container, REST API, well-documented.
|
|
|
|
**MCP access**: add a `knowledge-search` MCP tool to the gateway that calls LobeChat's
|
|
knowledge base search API. Zero new infra — LobeChat is already on `shared_backend`.
|
|
|
|
**IBM Granite embedding** (granite-embedding-30m-english) — good quality, tiny (30M params),
|
|
runs CPU-only. Alternative to Mistral-embed if you want fully local embeddings. Not on
|
|
LiteLLM yet but can add as a custom provider pointing at an Ollama/IPEX-LLM instance.
|
|
|
|
---
|
|
|
|
## 2. Document conversion pipeline
|
|
|
|
**Problem**: No systematic ingestion path from raw docs (PDF, DOCX, HTML) into
|
|
structured text for chunking/embedding.
|
|
|
|
**Tools to consider**:
|
|
|
|
| Tool | Docker image | Best for |
|
|
|------|-------------|----------|
|
|
| **Docling** (IBM, 2024) | `ghcr.io/ds4sd/docling` | PDFs with complex layout (tables, columns, figures); outputs Markdown |
|
|
| **Unstructured** | `quay.io/unstructured-io/unstructured-api` | Wide format support (HTML, DOCX, PPTX, images with OCR); REST API |
|
|
| **Gotenberg** | `gotenberg/gotenberg` | HTML/Office → PDF pre-processing stage; not a text extractor |
|
|
| **Apache Tika** | `apache/tika` | Broad format support; outputs plain text; lower quality than Docling for PDFs |
|
|
| **Markitdown MCP** | already in gateway | Ad-hoc on-demand conversion; not suitable for batch pipelines |
|
|
|
|
**Recommended pipeline (n8n-orchestrated)**:
|
|
```
|
|
Nextcloud / Paperless webhook
|
|
→ Gotenberg (Office → PDF)
|
|
→ Docling (PDF → Markdown chunks)
|
|
→ LiteLLM /v1/embeddings (Mistral-embed)
|
|
→ Qdrant upsert
|
|
```
|
|
|
|
Docling runs CPU-only comfortably on the NUC. Unstructured is heavier but has a
|
|
managed API if the self-hosted version is too slow.
|
|
|
|
Paperless-ngx already OCRs documents — its full-text content is accessible via the
|
|
REST API. An n8n workflow polling `/api/documents/?added__gt=<last_run>` can extract
|
|
and re-embed directly, without re-running OCR.
|
|
|
|
---
|
|
|
|
## 3. On-device LLM (Arc iGPU)
|
|
|
|
The Arc iGPU is currently ComfyUI-only. Small language models can also run on it:
|
|
|
|
- **IPEX-LLM** (Intel) + **Ollama** — Intel provides a patched Ollama build that uses
|
|
IPEX-LLM for Arc acceleration. Supports Phi-3.5-mini, Gemma-2B, TinyLlama.
|
|
Useful as a fast/cheap fallback when Groq/Mistral rate limits hit.
|
|
- **llama.cpp with Vulkan** — alternative to IPEX-LLM; no Intel-specific driver,
|
|
uses Vulkan compute. Less optimised for Arc but simpler to deploy.
|
|
- **Practical constraint**: Arc iGPU shares system RAM; sustained LLM inference
|
|
competes with other XPU workloads (e.g., ComfyUI). A model serving config that
|
|
pauses one workload while the other runs is needed, or deploy them on separate
|
|
CPU/GPU budgets.
|
|
|
|
If a second NUC or small GPU box becomes available, this becomes the primary use case.
|
|
|
|
---
|
|
|
|
## 4. Arcane agents on NUC (central instance on Proxmox LXC)
|
|
|
|
Arcane's central server moves to a **Proxmox LXC** (lightweight, stays responsive even
|
|
if the Docker stack has issues). The NUC and any other Docker host runs the **Arcane
|
|
agent** (headless worker), which connects back to the central instance.
|
|
|
|
**Deployment**:
|
|
- LXC: central Arcane server + observability stack (see §8)
|
|
- NUC (`.40`) and `.49`: `arcane-agent` (headless) in compose, connects to LXC
|
|
- Arcane agents reach Docker via either:
|
|
- **TCP Docker socket + TLS certs** — most secure; requires cert generation per host
|
|
- **SSH Docker contexts** — simpler; gateway mounts socket into agent via SSH tunnel
|
|
|
|
**Observability co-located on the LXC** (see §8 — lightweight, OTEL-aware stack that
|
|
survives Docker restarts).
|
|
|
|
---
|
|
|
|
## 5. AI memory / personalization
|
|
|
|
**Mem0** (`mem0ai/mem0`) — persistent user memory layer that sits in front of LLM calls.
|
|
Stores facts extracted from conversations into a vector DB (Qdrant backend supported).
|
|
Can integrate with LobeChat or as an MCP tool. Lets the AI remember preferences,
|
|
past context, and user-specific facts across sessions.
|
|
|
|
Alternative: **Letta** (formerly MemGPT) — stateful agent framework with persistent
|
|
memory; more opinionated.
|
|
|
|
---
|
|
|
|
## 6. Workflow / automation upgrades
|
|
|
|
- **n8n → AI agent nodes**: n8n 1.x has native AI agent nodes (LangChain under the hood).
|
|
Can build RAG, email triage, document processing workflows visually.
|
|
- **Activepieces** — lighter-weight n8n alternative; good for simple integrations.
|
|
Probably not worth adding since n8n is already running.
|
|
- **Temporal** — durable workflow engine for long-running or retry-heavy pipelines
|
|
(e.g., large document batch processing). Overkill unless pipelines become complex.
|
|
|
|
---
|
|
|
|
## 7. External services worth considering
|
|
|
|
| Service | Purpose | Notes |
|
|
|---------|---------|-------|
|
|
| **Backblaze B2** | Off-site backup (already in todos) | rclone sync from Garage; restic for PG dumps |
|
|
| **Cloudflare R2** | S3-compatible CDN-backed storage | Free egress; good for Lobe file serving |
|
|
| **Resend** | Transactional email | 3K/mo free; better deliverability than self-hosted Postal |
|
|
| **ntfy.sh (cloud)** | Push notifications fallback | Already running self-hosted ntfy; cloud as relay |
|
|
| **Cloudflare Turnstile** | Bot protection for public endpoints | Free; no JS challenge |
|
|
| **Novu** | Notification orchestration | Multi-channel (email, push, Slack); self-hostable |
|
|
|
|
---
|
|
|
|
## 8. Observability (on Proxmox LXC, alongside Arcane)
|
|
|
|
**Runs on the LXC, not the NUC** — stays alive if the Docker stack misbehaves. Lightweight
|
|
enough for a 2 vCPU / 4GB LXC.
|
|
|
|
**Recommended stack** (all OTEL-aware, compose-based):
|
|
|
|
| Component | Image | Role |
|
|
|-----------|-------|------|
|
|
| **OpenTelemetry Collector** | `otel/opentelemetry-collector-contrib` | Receives traces/metrics/logs from all services (OTLP gRPC+HTTP); fans out to backends |
|
|
| **VictoriaMetrics** | `victoriametrics/victoria-metrics` | Prometheus-compatible TSDB; scrapes NUC exporters + receives from OTEL collector; lighter than Prometheus |
|
|
| **Grafana** | `grafana/grafana` | Dashboards; datasource = VictoriaMetrics + Loki |
|
|
| **Loki** | `grafana/loki` | Log aggregation; receives from OTEL collector |
|
|
| **Uptime Kuma** | `louislam/uptime-kuma` | HTTP/TCP uptime checks for all public endpoints; alerts via ntfy |
|
|
|
|
**OTEL receivers** from the NUC Docker stack:
|
|
- **LiteLLM**: native OTEL export — traces every LLM call with token counts, model, latency
|
|
- **n8n**: Prometheus `/metrics` endpoint
|
|
- **Garage**: Prometheus `/metrics`
|
|
- **mcp-gateway**: add `opentelemetry-sdk` instrumentation to `server.py`
|
|
- **Docker host**: `node_exporter` + `cadvisor` on the NUC, scraped by VictoriaMetrics
|
|
|
|
**NUC → LXC connectivity**: Both are on the same LAN. OTEL collector listens on the LXC's
|
|
LAN IP (e.g., `192.168.1.X:4317` gRPC). Services push OTEL directly to it; Prometheus
|
|
pull-scraping from VictoriaMetrics goes to NUC exporters over LAN.
|
|
|
|
---
|
|
|
|
## 9. Storage / S3 improvements
|
|
|
|
- **Garage external S3** is currently non-responsive (tracked in todos). Fix or
|
|
replace with **MinIO** (single-node, much better tooling/UI, easy migration).
|
|
- **Lobe file serving via Cloudflare R2**: Lobe uploads to Garage → sync to R2 →
|
|
serve from CDN. Reduces NUC egress for image-heavy sessions.
|
|
- **Qdrant data on UNAS**: Mount `qdrant_data` from NAS for persistence across
|
|
container re-creates (same pattern as arr-stack).
|
|
|
|
---
|
|
|
|
## 9. ComfyUI MCP — async queue + image-to-image + reference
|
|
|
|
**Current problem**: MCP tool blocks until the image is done (~90-290s). LLMs time out at ~60s.
|
|
The fix is a job-queue pattern:
|
|
|
|
```
|
|
generate_image(prompt) → returns {job_id, status: "queued"} immediately
|
|
get_image_status(job_id) → returns {status, progress, image_url_when_done}
|
|
list_queue() → shows all pending/running jobs
|
|
cancel_job(job_id) → cancels a queued job (confirms with user first)
|
|
```
|
|
|
|
ComfyUI's own API is already async (`POST /prompt` → poll `/history/{id}`). The MCP server
|
|
just needs to expose this model instead of blocking.
|
|
|
|
**Image-to-image**: new workflow `flux-schnell-img2img-api.json`. Takes an init image URL +
|
|
denoise strength. ComfyUI `LoadImageFromURL` node (or upload + `LoadImage`).
|
|
|
|
**Reference tool**: `list_previous_images(n=5)` — queries ComfyUI `/history` API,
|
|
returns recent job thumbnails + prompts. User can pick one to reference or iterate from.
|
|
|
|
**Delivery to S3/Garage**: on completion, upload output PNG to Garage `comfyui-outputs` bucket
|
|
→ return a permanent URL. Avoids ComfyUI's ephemeral `/view` endpoint.
|
|
|
|
---
|
|
|
|
## 10. LiteLLM custom provider / LobeChat provider extension
|
|
|
|
**Question**: build an adapter in LiteLLM to use models via "quasi API" (non-standard
|
|
endpoints, auth, or routing)?
|
|
|
|
**LiteLLM supports custom providers** via `custom_llm_provider` in config.yaml. You write
|
|
a Python class that implements `completion()` and `async_completion()`. This is the right
|
|
path for wrapping non-standard APIs (local models, proprietary endpoints, protocol bridges).
|
|
|
|
Example use cases:
|
|
- Wrap a Claude Max subscription via Anthropic's API (different billing model)
|
|
- Add a local model served by IPEX-LLM on the Arc iGPU
|
|
- Bridge a custom inference server that speaks a different protocol
|
|
|
|
**LobeChat provider extension**: LobeChat's provider list is compiled into the app.
|
|
Adding a new provider requires rebuilding LobeChat from source (fork + add to
|
|
`src/config/aiModels/`). High effort; only worth it for a permanent/long-term provider.
|
|
For ad-hoc needs, use LiteLLM as the adapter and point LobeChat at it via the existing
|
|
`ai.nuclide.systems` OpenAI-compatible endpoint.
|
|
|
|
---
|
|
|
|
## Priority order (rough)
|
|
|
|
1. **ComfyUI MCP async queue** — fixes timeout; unblocks img2img + reference features
|
|
2. **Docling container + n8n ingest pipeline** — Nextcloud/Paperless → LobeChat knowledge base
|
|
3. **Arcane LXC + agents on NUC** — when ready to migrate
|
|
4. **Observability LXC** — OTEL collector + VictoriaMetrics + Grafana + Uptime Kuma, co-located
|
|
5. **Backblaze B2 off-site backup** — already in todos
|
|
6. **On-device LLM (Arc)** — depends on IPEX-LLM + Ollama Intel build stability
|