Files
nexa/CLAUDE.md
T
Claude 17c7f5033f Stage embeddings: ship Path A (text) now, prepare Path C (text+visual) additively
Decision (Q15 resolved): start with text-only via TEI + bge-m3 in Phase 3.1,
prepare data shapes so Phase 3.2 (visual collection via infinity + jina-clip-v2)
is a pure additive operation — no rename, no schema migration, no n8n rewiring.

Concretely:
- Qdrant collection renamed nexa_knowledge → nexa_knowledge_text (1024-dim
  for bge-m3) with modality-aware payload (modality, source_type, media_uri,
  graph_iri, content_hash, context). Visual placeholder schema committed
  alongside (qdrant_schema_visual.json, 768-dim, jina-clip-v2).
- Image attachments captured in 3.1 are recorded in GraphDB as nexa:Note with
  nexa:modality "image" + nexa:pendingVisualIndex true; the 3.2 backfill
  workflow picks them up and embeds. No data lost between phases — the queue
  is the GraphDB itself.
- RDF schema (docs/08) gains nexa:modality, nexa:mediaUri,
  nexa:vectorCollection, nexa:pendingVisualIndex from day one.
- docs/02 roadmap split: 3.1 = text RAG (Path A), 3.2 = visual collection
  (Path C), 3.4 = Ontotext GraphDB.
- docs/09 grows a "Phase add-on: visual collection (Phase 3.2)" section with
  the TEI→infinity swap, second collection create, LiteLLM second model
  registration, and the SPARQL-driven backfill query.
- New open questions: Q16 (queue ergonomics + does SAIA already proxy an
  embed model?), Q17 (reuse Immich's CLIP for photo-library queries?).
- docs/03 + CLAUDE.md updated so future runs use the new collection names
  and don't re-decide the staging.
2026-05-04 21:33:59 +00:00

3.5 KiB

Instructions for Claude (and other agents)

This file tells future automated runs what they need to know about this repo.

Repo conventions

  • Documentation: all docs live in /docs/ and are numbered. Entry point is docs/index.md. When you add a doc, give it the next free NN- prefix and add a row to the index TOC.
  • Runtime artifacts: live in nexa-core/ (workflows, prompts, configs, scripts). Don't put .md documentation in there — link from /docs/ instead.
  • Source-of-truth: if a doc duplicates content from nexa-core/config/*.md, delete the duplicate. Single source of truth.

Real infrastructure (verified from screenshots, May 2026)

  • Proxmox host nuc at 192.168.1.20:8006 (PVE 9.1.9).
    • LXC 102 dns (AdGuard) — internal DNS, rewrites for *.nuclide.systems.
    • LXC 103 backrest — backup orchestration.
    • LXC 104 docker — main docker host at 192.168.1.40 (40 containers).
    • LXC 105 nextcloud — Nextcloud at nc.nuclide.systems.
    • LXC 106 octoprint — currently Exited; flagged in docs/11.
    • LXC 108 zoraxy — reverse proxy at 192.168.1.4:8000, TLS for *.nuclide.systems.
    • VM 100 haos — Home Assistant.
  • Already-running services on docker host (don't redeploy):
    • Memos :5230, n8n :5678, LiteLLM :4000 (UI LobeHub :3210), Qdrant (qdrant_scientific), ntfy :7998, Karakeep/Hoarder, Vaultwarden :11001, Pocket-ID :1411, Immich, Audiobookshelf, Paperless-ngx, Traccar, Prowlarr, plus MCP containers (crawl4ai-mcp, markitdown-mcp, papersearch-mcp).
  • Decided for Nexa (don't re-litigate without user input):
    • Vector store: reuse qdrant_scientific with collections suffixed by modality (nexa_knowledge_text, nexa_knowledge_visual).
    • Embeddings staged: Phase 3.1 TEI + BAAI/bge-m3 (text-only, 1024-dim). Phase 3.2 swap to infinity and add jinaai/jina-clip-v2 (768-dim, joint text+image space). All forward-compat fields (modality, media_uri, graph_iri, nexa:pendingVisualIndex) exist from 3.1 — adding the visual collection is additive.
    • Graph store: Ontotext GraphDB (SPARQL/RDF), Phase 3.4. RDF schema in docs/08 already includes nexa:modality / nexa:mediaUri / nexa:vectorCollection / nexa:pendingVisualIndex.
    • Chat model: SAIA via LiteLLM virtual key.

When working on Nexa

  1. Read docs/index.md first — it's the navigator.
  2. Open questions first. Before writing code or workflow JSON, scan docs/11-open-questions.md. If your task touches an unanswered Q, stop and ask rather than picking a default. Append new blockers to that doc as [ ] Q-NN.
  3. Optimization findings. When you spot infrastructure improvements, add them to docs/12-optimization-opportunities.md as a numbered bullet — don't just mention them in commit messages.
  4. Never inline secrets in workflow JSON or .env committed to git. Use n8n credentials, LiteLLM virtual keys, or (longer term) Vaultwarden.
  5. Keep deployment minimal. The default answer to "do we need a new container?" is no — the existing stack covers most needs.

Branch policy

  • This branch is claude/organize-docs-deployment-7N4v2. Push only here unless told otherwise.
  • New work for an unrelated feature → new branch under claude/<topic>.