docs: ingest-pipeline design doc; update doc-ingestion with nomic service
This commit is contained in:
@@ -0,0 +1,188 @@
|
||||
# Ingest Pipeline — Filesystem Scan
|
||||
|
||||
One-shot pipeline that crawls a directory, converts documents to text, embeds them via
|
||||
the nomic service, and stores them in Qdrant. Run it manually against any mounted path.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
┌─────────────────────────┐
|
||||
│ ingest.py --path /... │ runs on CT 104 (has direct access to ai-internal network)
|
||||
└──────────┬──────────────┘
|
||||
│ per file
|
||||
▼
|
||||
┌────────────────────────────────────────────────────────────────┐
|
||||
│ 1. File scanner glob recursively, filter by ext │
|
||||
│ 2. Change detection sha256(path + mtime) → skip if seen │
|
||||
│ 3. Docling converter POST http://docling:5001/convert │
|
||||
│ (PDF/DOCX/PPTX/XLSX/HTML → Markdown) │
|
||||
│ Plain text/Markdown read directly │
|
||||
│ Images (jpg/png/...) send to nomic vision endpoint │
|
||||
│ 4. Chunker split Markdown by headers + size │
|
||||
│ chunk_size=1200 overlap=150 (matches OWUI RAG config) │
|
||||
│ 5. Nomic embedder POST http://nomic:80/v1/embeddings │
|
||||
│ text chunks → 768d images → 768d (same space!) │
|
||||
│ 6. Qdrant upsert collection `documents`, named vectors │
|
||||
└────────────────────────────────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────────────────────────────────────────────────────┐
|
||||
│ Qdrant: http://qdrant:6333 │
|
||||
│ collection: documents │
|
||||
│ vector: {size: 768, distance: Cosine} │
|
||||
│ payload: {path, title, chunk_idx, text, type, mtime, hash} │
|
||||
└──────────────────────────────────────────────────────────────┘
|
||||
│
|
||||
▼
|
||||
┌────────────────────────────┐
|
||||
│ Open WebUI RAG search │ queries `documents` collection
|
||||
│ n8n / MCP tools │ same collection, same embedding space
|
||||
└────────────────────────────┘
|
||||
```
|
||||
|
||||
## Why not n8n for this
|
||||
|
||||
n8n is the right tool for **event-driven** pipelines (webhook on new file, Paperless
|
||||
webhook, Nextcloud activity, scheduled re-sync). For a **bulk one-shot crawl**, a Python
|
||||
script is better:
|
||||
- direct filesystem access with `os.walk`
|
||||
- proper progress bar and error recovery
|
||||
- batched Qdrant upserts (n8n does one HTTP call per node)
|
||||
- easy CLI: `python3 ingest.py /mnt/pve/unas/Notizen`
|
||||
|
||||
n8n handles the _ongoing_ layer (see [Ongoing ingestion](#ongoing-ingestion-n8n)).
|
||||
|
||||
## Supported file types
|
||||
|
||||
| Extension | Handler | Notes |
|
||||
|-----------|---------|-------|
|
||||
| `.pdf` | Docling | best quality, preserves tables |
|
||||
| `.docx`, `.odt` | Docling | |
|
||||
| `.pptx` | Docling | slides → sections |
|
||||
| `.xlsx`, `.ods` | Docling | tables → Markdown |
|
||||
| `.html`, `.htm` | Docling | |
|
||||
| `.md`, `.txt`, `.rst` | direct read | no conversion needed |
|
||||
| `.jpg`, `.jpeg`, `.png`, `.webp`, `.gif` | nomic vision | 768d image embedding, no text chunks |
|
||||
|
||||
## Script: `ingest.py`
|
||||
|
||||
Lives at `/opt/stacks/ai/ingest/ingest.py`. Runs directly on CT 104.
|
||||
|
||||
### Usage
|
||||
|
||||
```bash
|
||||
# Index everything under a path
|
||||
python3 /opt/stacks/ai/ingest/ingest.py --path /mnt/pve/unas/Notizen
|
||||
|
||||
# Different collection, force re-index
|
||||
python3 /opt/stacks/ai/ingest/ingest.py \
|
||||
--path /mnt/pve/unas/Dokumente \
|
||||
--collection work-docs \
|
||||
--force
|
||||
|
||||
# Dry run (print files, don't embed)
|
||||
python3 /opt/stacks/ai/ingest/ingest.py --path /path/to/docs --dry-run
|
||||
```
|
||||
|
||||
### State file
|
||||
|
||||
`/opt/stacks/ai/ingest/state/<collection>.json` tracks `{path: {hash, indexed_at}}`.
|
||||
Subsequent runs skip unchanged files. Delete the state file to force full re-index.
|
||||
|
||||
### Qdrant collection schema
|
||||
|
||||
```python
|
||||
vectors_config = VectorParams(size=768, distance=Distance.COSINE)
|
||||
# payload per point:
|
||||
{
|
||||
"path": "/mnt/pve/unas/Notizen/someFile.md",
|
||||
"title": "someFile", # filename without ext
|
||||
"chunk_idx": 0, # 0-based chunk index within file
|
||||
"total_chunks": 3,
|
||||
"text": "…chunk content…", # empty string for images
|
||||
"type": "markdown", # markdown | pdf | docx | image | …
|
||||
"mtime": 1716700000.0,
|
||||
"hash": "a3f…", # sha256 of file content
|
||||
"source": "filesystem",
|
||||
}
|
||||
```
|
||||
|
||||
### Chunking strategy
|
||||
|
||||
For Markdown output from Docling (and raw .md/.txt):
|
||||
1. Split on `## ` / `### ` headers first (keep header as first line of chunk)
|
||||
2. If chunk > 1200 chars, split further on double-newline (`\n\n`)
|
||||
3. If still > 1200 chars, hard-split with 150-char overlap
|
||||
|
||||
Images: single point per file, no chunking.
|
||||
|
||||
### Error handling
|
||||
|
||||
- Docling timeout (>60s): skip file, log to `ingest_errors.log`, continue
|
||||
- Qdrant upsert failure: retry 3×, then log and continue
|
||||
- Re-run is safe: state file prevents double-indexing
|
||||
|
||||
## Building and running
|
||||
|
||||
```bash
|
||||
# On CT 104
|
||||
mkdir -p /opt/stacks/ai/ingest
|
||||
cd /opt/stacks/ai/ingest
|
||||
|
||||
# Install deps (lightweight — no torch needed, calls services via HTTP)
|
||||
pip3 install qdrant-client requests tqdm
|
||||
|
||||
# Run
|
||||
python3 ingest.py --path /mnt/pve/unas/Notizen
|
||||
```
|
||||
|
||||
No container needed for the script itself — it runs on CT 104 bare Python and calls
|
||||
`http://docling:5001`, `http://nomic:80`, `http://qdrant:6333` via `ai-internal` (all
|
||||
on the same Docker network/host).
|
||||
|
||||
To run it from _outside_ CT 104 (e.g. the PVE host), wrap it in a container later.
|
||||
|
||||
## Open WebUI integration
|
||||
|
||||
OWUI's knowledge base already points at Qdrant (`VECTOR_DB=qdrant`). To surface
|
||||
documents from the `documents` collection in chat:
|
||||
|
||||
1. Admin → Knowledge → Create Knowledge Base
|
||||
2. Name: "Local Documents"
|
||||
3. The collection is populated by `ingest.py` — OWUI will search it on `#`-prefixed
|
||||
RAG queries or when the knowledge base is enabled in a chat.
|
||||
|
||||
> **Note**: OWUI creates its own internal collection names. To share the same `documents`
|
||||
> collection between `ingest.py` and OWUI, use the OWUI API to create a knowledge base
|
||||
> pointing to the pre-populated collection — or let OWUI manage its own collection and
|
||||
> have `ingest.py` add documents via the OWUI knowledge API (`POST /api/v1/knowledge/{id}/file/add`).
|
||||
> The OWUI API path is cleaner for OWUI search integration; the direct Qdrant path is
|
||||
> better for external tools (n8n, MCP).
|
||||
|
||||
## Ongoing ingestion (n8n)
|
||||
|
||||
After the one-shot crawl, wire n8n for continuous ingestion:
|
||||
|
||||
| Trigger | n8n nodes | Notes |
|
||||
|---------|-----------|-------|
|
||||
| Nextcloud webhook (file created/modified) | HTTP → SSH → `ingest.py --path <file>` | Nextcloud admin → Webhooks app |
|
||||
| Paperless post-consume webhook | HTTP → Docling → nomic → Qdrant upsert | Paperless has `POST_CONSUME_SCRIPT` hook |
|
||||
| Cron re-scan | Schedule → SSH → `ingest.py --path /mnt/pve/unas/Notizen` | weekly full re-sync |
|
||||
|
||||
The cron re-scan is safe because `ingest.py` skips unchanged files via state hash.
|
||||
|
||||
## Paths available on CT 104
|
||||
|
||||
| Path | Contents |
|
||||
|------|---------|
|
||||
| `/mnt/pve/unas/Notizen/` | Obsidian vault (Markdown) |
|
||||
| `/mnt/pve/unas/` | full UNAS NFS share |
|
||||
| `/opt/stacks/*/` | stack configs (already in Git, lower priority) |
|
||||
|
||||
Nextcloud files are on CT 105 (`192.168.1.41`). Access via:
|
||||
- WebDAV: `https://nc.nuclide.systems/remote.php/dav/files/fkrebs@nucli.de/`
|
||||
- Or mount the NC data volume — not currently mounted on CT 104
|
||||
|
||||
## Status
|
||||
|
||||
**Not yet implemented.** Design only. Next step: write `ingest.py`.
|
||||
Reference in New Issue
Block a user