3.5 KiB
Document Ingestion Pipeline
Ingest PDFs, Word, PPTX, XLSX from Nextcloud/Paperless into Open WebUI's knowledge base and make them searchable via MCP.
Architecture
Nextcloud folder / Paperless webhook
→ n8n trigger (CT 104)
→ Docling MCP (port 18005) — PDF/DOCX/PPTX/XLSX → Markdown + structure
→ TEI /v1/embeddings — multilingual-e5-base (local, 768d)
→ Qdrant (shared vector store, http://qdrant:6333)
→ Open WebUI knowledge base API ← searchable in chat
Paperless shortcut: Paperless-ngx already OCRs documents. Its full-text content is
available at /api/documents/?added__gt=<last_run>. An n8n workflow can re-embed
directly from Paperless's REST API without re-running Docling for already-OCR'd PDFs.
Deployed components (as of 2026-05-26)
| Component | Location | Endpoint |
|---|---|---|
| Qdrant | CT 104, ai-internal net |
http://qdrant:6333 |
| TEI (Text Embeddings Inference) | CT 104, ai-internal net |
http://tei:80 |
| Open WebUI | CT 104, port 14002 | Uses Qdrant + TEI natively |
| Docling MCP | CT 104, port 18005 | MCP server in gateway |
Embedding stack
- Model:
intfloat/multilingual-e5-base(768d, multilingual DE+EN, ~278 MB ONNX) - Server: HuggingFace TEI
cpu-1.6— OpenAI-compatible athttp://tei:80/v1/embeddings - Auth: none (internal network only)
- Upgrade path: switch to
mistral/mistral-embedvia Bifrost (1024d, better quality) — requires re-index (drop + recreate Qdrant collections)
OWUI RAG config (env-driven)
VECTOR_DB=qdrant
QDRANT_URI=http://qdrant:6333
RAG_EMBEDDING_ENGINE=openai
RAG_OPENAI_API_BASE_URL=http://tei:80
RAG_OPENAI_API_KEY=none
RAG_EMBEDDING_MODEL=intfloat/multilingual-e5-base
CONTENT_EXTRACTION_ENGINE=docling
CHUNK_SIZE=1200
CHUNK_OVERLAP=150
ENABLE_RAG_HYBRID_SEARCH=true
Bifrost embedding models (available for external services / upgrade)
Three embedding models tested and working via http://bifrost:8080/v1/embeddings:
mistral/mistral-embed(1024d) — ✓ production-readymistral/codestral-embed(1024d) — ✓gemini/gemini-embedding-001(768d/1536d) — ✓
Converter comparison
| Tool | Image | Formats | Notes |
|---|---|---|---|
| Docling (deployed) | MCP on 18005 | PDF, DOCX, PPTX, XLSX, HTML | Best for structured Office/PDF with tables |
| MinerU | opendatalab/mineru |
PDF (layout-aware, OCR) | Better for academic papers / scanned PDFs |
| Markitdown (in gateway) | — | Office, PDF | Ad-hoc only; not suitable for batch |
Start with Docling — already deployed. Add MinerU if academic paper OCR quality is needed.
Sharing Qdrant with other services
Qdrant is on ai-internal network — any service on that network can use it:
from qdrant_client import QdrantClient
client = QdrantClient(url="http://qdrant:6333")
n8n, MCP tools, and custom pipelines should use http://tei:80/v1/embeddings for
consistent 768d vectors. Mixing models/dimensions in the same collection will fail.
Intel Arc GPU passthrough (planned)
CT 104 has Intel Core Ultra 7 155H iGPU but no /dev/dri/render* device is
passed through. When passthrough is enabled, switch TEI to the Intel image:
image: ghcr.io/huggingface/text-embeddings-inference:intel-1.6
This uses IPEX and runs ~5–10× faster for embedding batches.
Pending
- n8n workflow: Nextcloud/Paperless → Docling → TEI → Qdrant
- MCP knowledge-search tool pointing at Qdrant (replaces LobeChat search intent)