Files
nexa/nexa-core/config/qdrant_schema.json
T
Claude 17c7f5033f Stage embeddings: ship Path A (text) now, prepare Path C (text+visual) additively
Decision (Q15 resolved): start with text-only via TEI + bge-m3 in Phase 3.1,
prepare data shapes so Phase 3.2 (visual collection via infinity + jina-clip-v2)
is a pure additive operation — no rename, no schema migration, no n8n rewiring.

Concretely:
- Qdrant collection renamed nexa_knowledge → nexa_knowledge_text (1024-dim
  for bge-m3) with modality-aware payload (modality, source_type, media_uri,
  graph_iri, content_hash, context). Visual placeholder schema committed
  alongside (qdrant_schema_visual.json, 768-dim, jina-clip-v2).
- Image attachments captured in 3.1 are recorded in GraphDB as nexa:Note with
  nexa:modality "image" + nexa:pendingVisualIndex true; the 3.2 backfill
  workflow picks them up and embeds. No data lost between phases — the queue
  is the GraphDB itself.
- RDF schema (docs/08) gains nexa:modality, nexa:mediaUri,
  nexa:vectorCollection, nexa:pendingVisualIndex from day one.
- docs/02 roadmap split: 3.1 = text RAG (Path A), 3.2 = visual collection
  (Path C), 3.4 = Ontotext GraphDB.
- docs/09 grows a "Phase add-on: visual collection (Phase 3.2)" section with
  the TEI→infinity swap, second collection create, LiteLLM second model
  registration, and the SPARQL-driven backfill query.
- New open questions: Q16 (queue ergonomics + does SAIA already proxy an
  embed model?), Q17 (reuse Immich's CLIP for photo-library queries?).
- docs/03 + CLAUDE.md updated so future runs use the new collection names
  and don't re-decide the staging.
2026-05-04 21:33:59 +00:00

25 lines
921 B
JSON

{
"collection_name": "nexa_knowledge_text",
"vector_config": {
"size": 1024,
"distance": "Cosine"
},
"payload_schema": {
"source": "keyword",
"source_type": "keyword",
"modality": "keyword",
"mime_type": "keyword",
"media_uri": "keyword",
"context": "keyword",
"tags": "keyword",
"graph_iri": "keyword",
"content_hash":"keyword",
"created_at": "datetime"
},
"_notes": {
"vector_size_rationale": "1024 = bge-m3 output. If Q15 resolves to nomic-embed-text-v1.5 instead, rebuild this collection with size: 768.",
"modality_field": "Always 'text' in this collection. Forward-compatibility for Phase 3.2 (see qdrant_schema_visual.json).",
"graph_iri_field": "IRI of the corresponding nexa:Note in GraphDB. The same value is stored on the GraphDB node as nexa:vectorId pointing back here. This is the cross-pillar bridge."
}
}