# Memory Service **Package:** `@nexusai/memory-service` **Location:** `packages/memory-service` **Deployed on:** Mini PC 1 (192.168.0.81) **Port:** 3002 ## Purpose Responsible for all reading and writing of long-term memory. Acts as the sole interface to both SQLite and Qdrant — no other service accesses these stores directly. On episode creation, automatically triggers entity and relationship extraction and embeds results into Qdrant. ## Dependencies - `express` — HTTP API - `better-sqlite3` — SQLite driver - `@qdrant/js-client-rest` — Qdrant vector store client - `dotenv` — environment variable loading - `@nexusai/shared` — shared utilities and constants ## Environment Variables | Variable | Required | Default | Description | |---|---|---|---| | PORT | No | 3002 | Port to listen on | | SQLITE_PATH | Yes | — | Path to SQLite database file | | QDRANT_URL | No | http://localhost:6333 | Qdrant instance URL | | EMBEDDING_SERVICE_URL | No | http://localhost:3003 | Embedding service URL | | INFERENCE_SERVICE_URL | No | http://localhost:3001 | Inference service URL — entity extraction routes through its `/utility/complete` endpoint | ## Internal Structure ``` src/ ├── db/ │ ├── index.js # SQLite connection + init + migrate() + one-time FTS backfill │ ├── migrations.js # Forward-only versioned migration runner (PRAGMA user_version) │ ├── schema.js # Complete current shape: tables, indexes, FTS5, triggers │ ├── projects.js # Project CRUD functions │ └── summaries.js # Summary CRUD functions ├── episodic/ │ └── index.js # Session + episode CRUD, FTS search, embedding write path ├── semantic/ │ └── index.js # Qdrant collection management, upsert, search, delete ├── entities/ │ ├── index.js # Entity + relationship CRUD (upsert, mention tracking) │ └── extraction.js # Automatic entity + relationship extraction via qwen2.5:3b ├── graph/ │ └── index.js # Knowledge graph traversal (neighborhood queries, recursive CTE) └── index.js # Express app + all route definitions ``` ## SQLite Schema Eight core tables: - **sessions** — top-level conversation containers. Fields: `external_id`, `name`, `project_id`, `metadata` - **episodes** — individual exchanges (user message + AI response) tied to a session - **entities** — named things the system learns about (people, places, concepts, etc.). Fields include `mention_count`, `confidence`, `source`, `last_seen_at` - **relationships** — directional labeled links between entities (`from_id`, `to_id`, `label`). Fields include `mention_count`, `notes` - **entity_episodes** — join table linking entities to the episodes where they were extracted. Used for provenance and orphan cleanup - **summaries** — condensed episode groups for efficient context retrieval - **projects** — named groupings of sessions with `name`, `description`, `colour`, `icon`, `isolated`, `notes`, `system_prompt` ### Schema & Migrations `schema.js` holds the **complete current shape** — every table, column, index, the FTS5 virtual table, and its triggers — as the single source of truth for a fresh database. It uses `CREATE TABLE IF NOT EXISTS`, so on a fresh DB it builds everything; on an existing DB it skips tables that already exist (and therefore does **not** reconcile columns on old tables — that's what migrations are for). `db/migrations.js` is a forward-only versioned runner keyed on `PRAGMA user_version`: ```js const migrations = [ (_db) => {}, // v0 → v1: consolidated baseline (historical ALTERs folded into schema.js) ]; const LATEST_VERSION = migrations.length; // derived, never hand-maintained ``` `migrate(db)` reads `user_version`, applies every entry newer than it (each in a transaction alongside its version bump), and stamps the result. A fresh DB is built whole by `schema.js` and simply stamped to `LATEST_VERSION`; the baseline entry is a no-op. **Adding a schema change:** append a new function to the `migrations` array (which bumps `LATEST_VERSION` automatically). Never edit an already-shipped entry, and never edit a table in `schema.js` expecting existing DBs to pick it up — they won't. This replaces the previous pattern of stacking silent `try/catch ALTER TABLE` statements in `db/index.js` on every boot. > **Consolidation note:** the historical ALTERs were folded into `schema.js` > rather than preserved as replayable migrations, so this assumes a fresh > database (which is the case post-wipe). An older, pre-consolidation database > would **not** auto-upgrade — `schema.js` skips its existing tables and the > baseline migration is a no-op. To support upgrading old DBs, the v1 baseline > would instead perform guarded (`ADD COLUMN if missing`) catch-up. ### FTS5 Full-Text Search An `episodes_fts` external-content virtual table enables keyword search across episodes. Three triggers (`episodes_fts_insert`, `episodes_fts_update`, `episodes_fts_delete`) keep the index in sync with the `episodes` table automatically during normal operation. A one-time backfill in `db/index.js` handles the case where the FTS table is created on a DB that already holds episodes (e.g. episodes predating FTS). It is gated on "did `episodes_fts` not exist before this boot," checked via `sqlite_master` **before** running the schema — not on a row-count comparison, because `COUNT(*)` on an external-content FTS5 table proxies the content table and cannot detect a desync. This replaced an unconditional full FTS rebuild that previously ran on every startup. ### SQLite Configuration - `journal_mode = WAL` — non-blocking reads during writes - `foreign_keys = ON` — enforces referential integrity and cascade deletes - PRAGMAs set via `db.pragma()`, not `db.exec()` ### Dynamic Updates Both `updateSession` and `updateProject` build their `SET` clause dynamically from only the fields passed — prevents partial updates from overwriting fields that weren't touched. `updateProject` allowlist: ```js const allowed = ['name', 'description', 'colour', 'icon', 'isolated', 'notes', 'system_prompt']; ``` ## Qdrant / Semantic Layer Three Qdrant collections are initialized on service startup via `semantic.initCollections()`: | Collection | Purpose | |---|---| | `episodes` | Embeddings for individual conversation exchanges | | `entities` | Embeddings for named entities | | `summaries` | Embeddings for condensed episode summaries | All collections use **768-dimension vectors** with **Cosine similarity**, matching `nomic-embed-text` via Ollama. Vector size and distance metric are defined in `@nexusai/shared` — not hardcoded here. `initCollections()` iterates `Object.values(COLLECTIONS)` and creates any collection that doesn't already exist at startup — all three collections are guaranteed to exist before any requests are handled. Each collection exposes upsert, search (with optional Qdrant filter), and delete operations. The `wait: true` flag is used on all writes. ## Embedding Write Path When a new episode is created: 1. Episode saved to SQLite synchronously — response returned immediately 2. User message + AI response combined: `User: ...\nAssistant: ...` 3. Text sent to embedding service (`POST /embed`) 4. Vector upserted into `episodes` Qdrant collection with payload `{ sessionId, createdAt }` This step is **fire-and-forget** — if embedding fails, the episode is still saved and searchable via FTS. The error is logged but not surfaced. > The Qdrant payload stores `sessionId` (the internal integer ID). See > `memory-isolation.md` for how project-level filtering works. ## Entity Layer Entities and relationships use upsert semantics with composite unique constraints to prevent duplicates: - `UNIQUE(name, type)` on entities — conflict increments `mention_count` and updates `last_seen_at` - `UNIQUE(from_id, to_id, label)` on relationships — conflict increments `mention_count` and preserves existing `notes` - `ON DELETE CASCADE` on relationship foreign keys After each episode is saved, `extraction.js` automatically extracts named entities **and relationships** from the conversation using `qwen2.5:3b` on Ollama — fire-and-forget. Each saved entity is also linked to the episode via the `entity_episodes` join table. > For full details on the extraction pipeline and JSON format, see `entity-extraction.md`. > For the knowledge graph traversal layer, see `knowledge-graph.md`. ## Knowledge Graph Layer `src/graph/index.js` provides SQLite-based graph traversal over the entities and relationships tables. Two functions are exposed via HTTP: - **`getNeighborhood(entityId, depth)`** — recursive CTE traversal, bidirectional, returns `{ nodes, edges }` - **`getEntityNeighbors(entityIds[])`** — bulk 1-hop traversal for orchestration context assembly > For design rationale, traversal queries, and integration with orchestration, see `knowledge-graph.md`. ## Summaries Layer Session summaries are generated by `orchestration-service/src/services/summarization.js` after each episode write and stored here via `POST /summaries`. The memory service is responsible only for CRUD — generation logic lives in orchestration. > For full details on trigger conditions, prompt format, cumulative updates, > and ChatML token stripping, see `summarization.md`. ## Delete Behaviour (SQLite + Qdrant consistency) SQLite cascades handle relational cleanup, but Qdrant is a separate store and must be cleaned explicitly. Each delete path that removes embedded rows also removes the corresponding vectors: | Delete | SQLite effect | Qdrant cleanup | |---|---|---| | `DELETE /episodes/:id` | Row removed | `semantic.deleteEpisode(id)` — vector by point ID | | `DELETE /sessions/by-external/:id` | Session + episodes cascade-deleted | `semantic.deleteEpisodesBySession(id)` — **payload-filter** delete on `sessionId` | | `DELETE /entities/:id` | Row removed, relationships cascade | `semantic.deleteEntity(id)` — vector by point ID | All three Qdrant deletes are **fire-and-forget** with error logging, matching the fire-and-forget write path — a Qdrant failure logs but does not fail the delete. The session path uses a **payload-filter** delete (matching on the `sessionId` field in the vector payload) rather than enumerating episode point IDs. This matters because the SQLite cascade has already removed the episode rows by the time cleanup runs, so there are no IDs left to enumerate — the filter deletes by payload regardless. It also cleans up any pre-existing orphans for that session as a side effect. > **Not cleaned on session delete:** entity vectors. Entities are shared across > sessions and projects (`UNIQUE(name, type)` is global), so deleting one > session must not remove entities that other sessions still reference. Entity > vector lifecycle is tied to explicit entity deletion and the (planned) memory > consolidation / orphan-cleanup pass. > **Historical orphans:** vectors orphaned by session deletes *before* this > cleanup existed are not removed retroactively. A one-time sweep (scroll the > `episodes` collection, delete points whose `sessionId` no longer exists in > SQLite) clears them. ## Project Delete Behaviour Deleting a project runs as a transaction — it first nulls out `project_id` on all assigned sessions, then deletes the project. This avoids a foreign key constraint failure since `sessions.project_id` has no `ON DELETE` rule: ```js const doDelete = db.transaction(() => { db.prepare(`UPDATE sessions SET project_id = NULL WHERE project_id = ?`).run(id); db.prepare(`DELETE FROM projects WHERE id = ?`).run(id); }); ``` For all HTTP endpoints, see `api-routes.md`.