diff --git a/services/ingest-pipeline.md b/services/ingest-pipeline.md new file mode 100644 index 0000000..79bf61e --- /dev/null +++ b/services/ingest-pipeline.md @@ -0,0 +1,188 @@ +# Ingest Pipeline — Filesystem Scan + +One-shot pipeline that crawls a directory, converts documents to text, embeds them via +the nomic service, and stores them in Qdrant. Run it manually against any mounted path. + +## Architecture + +``` + ┌─────────────────────────┐ + │ ingest.py --path /... │ runs on CT 104 (has direct access to ai-internal network) + └──────────┬──────────────┘ + │ per file + ▼ + ┌────────────────────────────────────────────────────────────────┐ + │ 1. File scanner glob recursively, filter by ext │ + │ 2. Change detection sha256(path + mtime) → skip if seen │ + │ 3. Docling converter POST http://docling:5001/convert │ + │ (PDF/DOCX/PPTX/XLSX/HTML → Markdown) │ + │ Plain text/Markdown read directly │ + │ Images (jpg/png/...) send to nomic vision endpoint │ + │ 4. Chunker split Markdown by headers + size │ + │ chunk_size=1200 overlap=150 (matches OWUI RAG config) │ + │ 5. Nomic embedder POST http://nomic:80/v1/embeddings │ + │ text chunks → 768d images → 768d (same space!) │ + │ 6. Qdrant upsert collection `documents`, named vectors │ + └────────────────────────────────────────────────────────────────┘ + │ + ▼ + ┌──────────────────────────────────────────────────────────────┐ + │ Qdrant: http://qdrant:6333 │ + │ collection: documents │ + │ vector: {size: 768, distance: Cosine} │ + │ payload: {path, title, chunk_idx, text, type, mtime, hash} │ + └──────────────────────────────────────────────────────────────┘ + │ + ▼ + ┌────────────────────────────┐ + │ Open WebUI RAG search │ queries `documents` collection + │ n8n / MCP tools │ same collection, same embedding space + └────────────────────────────┘ +``` + +## Why not n8n for this + +n8n is the right tool for **event-driven** pipelines (webhook on new file, Paperless +webhook, Nextcloud activity, scheduled re-sync). For a **bulk one-shot crawl**, a Python +script is better: +- direct filesystem access with `os.walk` +- proper progress bar and error recovery +- batched Qdrant upserts (n8n does one HTTP call per node) +- easy CLI: `python3 ingest.py /mnt/pve/unas/Notizen` + +n8n handles the _ongoing_ layer (see [Ongoing ingestion](#ongoing-ingestion-n8n)). + +## Supported file types + +| Extension | Handler | Notes | +|-----------|---------|-------| +| `.pdf` | Docling | best quality, preserves tables | +| `.docx`, `.odt` | Docling | | +| `.pptx` | Docling | slides → sections | +| `.xlsx`, `.ods` | Docling | tables → Markdown | +| `.html`, `.htm` | Docling | | +| `.md`, `.txt`, `.rst` | direct read | no conversion needed | +| `.jpg`, `.jpeg`, `.png`, `.webp`, `.gif` | nomic vision | 768d image embedding, no text chunks | + +## Script: `ingest.py` + +Lives at `/opt/stacks/ai/ingest/ingest.py`. Runs directly on CT 104. + +### Usage + +```bash +# Index everything under a path +python3 /opt/stacks/ai/ingest/ingest.py --path /mnt/pve/unas/Notizen + +# Different collection, force re-index +python3 /opt/stacks/ai/ingest/ingest.py \ + --path /mnt/pve/unas/Dokumente \ + --collection work-docs \ + --force + +# Dry run (print files, don't embed) +python3 /opt/stacks/ai/ingest/ingest.py --path /path/to/docs --dry-run +``` + +### State file + +`/opt/stacks/ai/ingest/state/.json` tracks `{path: {hash, indexed_at}}`. +Subsequent runs skip unchanged files. Delete the state file to force full re-index. + +### Qdrant collection schema + +```python +vectors_config = VectorParams(size=768, distance=Distance.COSINE) +# payload per point: +{ + "path": "/mnt/pve/unas/Notizen/someFile.md", + "title": "someFile", # filename without ext + "chunk_idx": 0, # 0-based chunk index within file + "total_chunks": 3, + "text": "…chunk content…", # empty string for images + "type": "markdown", # markdown | pdf | docx | image | … + "mtime": 1716700000.0, + "hash": "a3f…", # sha256 of file content + "source": "filesystem", +} +``` + +### Chunking strategy + +For Markdown output from Docling (and raw .md/.txt): +1. Split on `## ` / `### ` headers first (keep header as first line of chunk) +2. If chunk > 1200 chars, split further on double-newline (`\n\n`) +3. If still > 1200 chars, hard-split with 150-char overlap + +Images: single point per file, no chunking. + +### Error handling + +- Docling timeout (>60s): skip file, log to `ingest_errors.log`, continue +- Qdrant upsert failure: retry 3×, then log and continue +- Re-run is safe: state file prevents double-indexing + +## Building and running + +```bash +# On CT 104 +mkdir -p /opt/stacks/ai/ingest +cd /opt/stacks/ai/ingest + +# Install deps (lightweight — no torch needed, calls services via HTTP) +pip3 install qdrant-client requests tqdm + +# Run +python3 ingest.py --path /mnt/pve/unas/Notizen +``` + +No container needed for the script itself — it runs on CT 104 bare Python and calls +`http://docling:5001`, `http://nomic:80`, `http://qdrant:6333` via `ai-internal` (all +on the same Docker network/host). + +To run it from _outside_ CT 104 (e.g. the PVE host), wrap it in a container later. + +## Open WebUI integration + +OWUI's knowledge base already points at Qdrant (`VECTOR_DB=qdrant`). To surface +documents from the `documents` collection in chat: + +1. Admin → Knowledge → Create Knowledge Base +2. Name: "Local Documents" +3. The collection is populated by `ingest.py` — OWUI will search it on `#`-prefixed + RAG queries or when the knowledge base is enabled in a chat. + +> **Note**: OWUI creates its own internal collection names. To share the same `documents` +> collection between `ingest.py` and OWUI, use the OWUI API to create a knowledge base +> pointing to the pre-populated collection — or let OWUI manage its own collection and +> have `ingest.py` add documents via the OWUI knowledge API (`POST /api/v1/knowledge/{id}/file/add`). +> The OWUI API path is cleaner for OWUI search integration; the direct Qdrant path is +> better for external tools (n8n, MCP). + +## Ongoing ingestion (n8n) + +After the one-shot crawl, wire n8n for continuous ingestion: + +| Trigger | n8n nodes | Notes | +|---------|-----------|-------| +| Nextcloud webhook (file created/modified) | HTTP → SSH → `ingest.py --path ` | Nextcloud admin → Webhooks app | +| Paperless post-consume webhook | HTTP → Docling → nomic → Qdrant upsert | Paperless has `POST_CONSUME_SCRIPT` hook | +| Cron re-scan | Schedule → SSH → `ingest.py --path /mnt/pve/unas/Notizen` | weekly full re-sync | + +The cron re-scan is safe because `ingest.py` skips unchanged files via state hash. + +## Paths available on CT 104 + +| Path | Contents | +|------|---------| +| `/mnt/pve/unas/Notizen/` | Obsidian vault (Markdown) | +| `/mnt/pve/unas/` | full UNAS NFS share | +| `/opt/stacks/*/` | stack configs (already in Git, lower priority) | + +Nextcloud files are on CT 105 (`192.168.1.41`). Access via: +- WebDAV: `https://nc.nuclide.systems/remote.php/dav/files/fkrebs@nucli.de/` +- Or mount the NC data volume — not currently mounted on CT 104 + +## Status + +**Not yet implemented.** Design only. Next step: write `ingest.py`.