Files
nexusAI/docs/services/memory-service.md
T
2026-08-23 23:27:53 -07:00

13 KiB

Memory Service

Package: @nexusai/memory-service
Location: packages/memory-service
Deployed on: Mini PC 1 (192.168.0.81)
Port: 3002

Purpose

Responsible for all reading and writing of long-term memory. Acts as the sole interface to both SQLite and Qdrant — no other service accesses these stores directly. On episode creation, automatically triggers entity and relationship extraction and embeds results into Qdrant.

Dependencies

  • express — HTTP API
  • better-sqlite3 — SQLite driver
  • @qdrant/js-client-rest — Qdrant vector store client
  • dotenv — environment variable loading
  • @nexusai/shared — shared utilities and constants

Environment Variables

Variable Required Default Description
PORT No 3002 Port to listen on
SQLITE_PATH Yes — Path to SQLite database file
QDRANT_URL No http://localhost:6333 Qdrant instance URL
EMBEDDING_SERVICE_URL No http://localhost:3003 Embedding service URL
INFERENCE_SERVICE_URL No http://localhost:3001 Inference service URL — entity extraction routes through its /utility/complete endpoint

Internal Structure

src/
├── db/
│   ├── index.js       # SQLite connection + init + migrate() + one-time FTS backfill
│   ├── migrations.js  # Forward-only versioned migration runner (PRAGMA user_version)
│   ├── schema.js      # Complete current shape: tables, indexes, FTS5, triggers
│   ├── projects.js    # Project CRUD functions
│   └── summaries.js   # Summary CRUD functions
├── episodic/
│   └── index.js       # Session + episode CRUD, FTS search, embedding write path
├── semantic/
│   └── index.js       # Qdrant collection management, upsert, search, delete
├── entities/
│   ├── index.js       # Entity + relationship CRUD (upsert, mention tracking)
│   └── extraction.js  # Automatic entity + relationship extraction via qwen2.5:3b
├── graph/
│   └── index.js       # Knowledge graph traversal (neighborhood queries, recursive CTE)
└── index.js           # Express app + all route definitions

SQLite Schema

Eight core tables:

  • sessions — top-level conversation containers. Fields: external_id, name, project_id, metadata
  • episodes — individual exchanges (user message + AI response) tied to a session
  • entities — named things the system learns about (people, places, concepts, etc.). Fields include mention_count, confidence, source, last_seen_at
  • relationships — directional labeled links between entities (from_id, to_id, label). Fields include mention_count, notes
  • entity_episodes — join table linking entities to the episodes where they were extracted. Used for provenance and orphan cleanup
  • summaries — condensed episode groups for efficient context retrieval
  • projects — named groupings of sessions with name, description, colour, icon, isolated, notes, system_prompt

Schema & Migrations

schema.js holds the complete current shape — every table, column, index, the FTS5 virtual table, and its triggers — as the single source of truth for a fresh database. It uses CREATE TABLE IF NOT EXISTS, so on a fresh DB it builds everything; on an existing DB it skips tables that already exist (and therefore does not reconcile columns on old tables — that's what migrations are for).

db/migrations.js is a forward-only versioned runner keyed on PRAGMA user_version:

const migrations = [
  (_db) => {},   // v0 → v1: consolidated baseline (historical ALTERs folded into schema.js)
  (db) => {      // v1 → v2: access tracking for consolidation lifecycle
    db.exec(`ALTER TABLE episodes ADD COLUMN last_accessed_at INTEGER`);
    db.exec(`ALTER TABLE episodes ADD COLUMN access_count INTEGER NOT NULL DEFAULT 0`);
    db.exec(`UPDATE episodes SET last_accessed_at = created_at`);  // backfill
  },
];
const LATEST_VERSION = migrations.length;  // derived, never hand-maintained

migrate(db) reads user_version, applies every entry newer than it (each in a transaction alongside its version bump), and stamps the result. A fresh DB is built whole by schema.js and simply stamped to LATEST_VERSION; the baseline entry is a no-op.

Adding a schema change: append a new function to the migrations array (which bumps LATEST_VERSION automatically). Never edit an already-shipped entry, and never edit a table in schema.js expecting existing DBs to pick it up — they won't. This replaces the previous pattern of stacking silent try/catch ALTER TABLE statements in db/index.js on every boot.

Consolidation note: the historical ALTERs were folded into schema.js rather than preserved as replayable migrations, so this assumes a fresh database (which is the case post-wipe). An older, pre-consolidation database would not auto-upgrade — schema.js skips its existing tables and the baseline migration is a no-op. To support upgrading old DBs, the v1 baseline would instead perform guarded (ADD COLUMN if missing) catch-up.

An episodes_fts external-content virtual table enables keyword search across episodes. Three triggers (episodes_fts_insert, episodes_fts_update, episodes_fts_delete) keep the index in sync with the episodes table automatically during normal operation.

A one-time backfill in db/index.js handles the case where the FTS table is created on a DB that already holds episodes (e.g. episodes predating FTS). It is gated on "did episodes_fts not exist before this boot," checked via sqlite_master before running the schema — not on a row-count comparison, because COUNT(*) on an external-content FTS5 table proxies the content table and cannot detect a desync. This replaced an unconditional full FTS rebuild that previously ran on every startup.

SQLite Configuration

  • journal_mode = WAL — non-blocking reads during writes
  • foreign_keys = ON — enforces referential integrity and cascade deletes
  • PRAGMAs set via db.pragma(), not db.exec()

Copying a live WAL database: cp on the .db file alone silently loses everything in the un-checkpointed -wal file (recent writes, even the migration version stamp). Always use sqlite3 nexusai.db "VACUUM INTO './copy.db'" (or .backup) — safe while the service is running, produces a complete single-file snapshot.

Dynamic Updates

Both updateSession and updateProject build their SET clause dynamically from only the fields passed — prevents partial updates from overwriting fields that weren't touched.

updateProject allowlist:

const allowed = ['name', 'description', 'colour', 'icon', 'isolated', 'notes', 'system_prompt'];

Qdrant / Semantic Layer

Three Qdrant collections are initialized on service startup via semantic.initCollections():

Collection Purpose
episodes Embeddings for individual conversation exchanges
entities Embeddings for named entities
summaries Embeddings for condensed episode summaries

All collections use 768-dimension vectors with Cosine similarity, matching nomic-embed-text via Ollama. Vector size and distance metric are defined in @nexusai/shared — not hardcoded here.

initCollections() iterates Object.values(COLLECTIONS) and creates any collection that doesn't already exist at startup — all three collections are guaranteed to exist before any requests are handled.

Each collection exposes upsert, search (with optional Qdrant filter), and delete operations. The wait: true flag is used on all writes.

Embedding Write Path

When a new episode is created:

  1. Episode saved to SQLite synchronously — response returned immediately
  2. User message + AI response combined: User: ...\nAssistant: ...
  3. Text sent to embedding service (POST /embed)
  4. Vector upserted into episodes Qdrant collection with payload { sessionId, createdAt }

This step is fire-and-forget — if embedding fails, the episode is still saved and searchable via FTS. The error is logged but not surfaced.

The Qdrant payload stores sessionId (the internal integer ID). See memory-isolation.md for how project-level filtering works.

Entity Layer

Entities and relationships use upsert semantics with composite unique constraints to prevent duplicates:

  • UNIQUE(name, type) on entities — conflict increments mention_count and updates last_seen_at
  • UNIQUE(from_id, to_id, label) on relationships — conflict increments mention_count and preserves existing notes
  • ON DELETE CASCADE on relationship foreign keys

After each episode is saved, extraction.js automatically extracts named entities and relationships from the conversation using qwen2.5:3b on Ollama — fire-and-forget. Each saved entity is also linked to the episode via the entity_episodes join table.

For full details on the extraction pipeline and JSON format, see entity-extraction.md.
For the knowledge graph traversal layer, see knowledge-graph.md.

Knowledge Graph Layer

src/graph/index.js provides SQLite-based graph traversal over the entities and relationships tables. Two functions are exposed via HTTP:

  • getNeighborhood(entityId, depth) — recursive CTE traversal, bidirectional, returns { nodes, edges }
  • getEntityNeighbors(entityIds[]) — bulk 1-hop traversal for orchestration context assembly

For design rationale, traversal queries, and integration with orchestration, see knowledge-graph.md.

Summaries Layer

Session summaries are generated by orchestration-service/src/services/summarization.js after each episode write and stored here via POST /summaries. The memory service is responsible only for CRUD — generation logic lives in orchestration.

For full details on trigger conditions, prompt format, cumulative updates, and ChatML token stripping, see summarization.md.

Access Tracking & Consolidation (dry-run)

Every episode selected into a chat context window (budget-selected, not the guaranteed-recency floor) gets an access bump via POST /episodes/touch — access_count incremented, last_accessed_at set to Date.now() (ms). Called fire-and-forget from orchestration; a failure loses one increment, nothing more.

GET /sessions/:id/consolidation-candidates scores episodes by access_count / (1 + days since last access) — never-accessed episodes fall back to created_at for the recency term and score exactly 0 (most eligible). Two floors apply: episodes younger than CONSOLIDATION.MIN_AGE_DAYS are excluded in SQL; sessions under CONSOLIDATION.MIN_SESSION_EPISODES return eligible: false before scoring runs. The endpoint is observe-only — the destructive pass (merge → summarize → Qdrant cleanup → orphan sweep) is not yet built.

Unit note: created_at is unix seconds (unixepoch()); last_accessed_at is unix milliseconds (Date.now()). The scoring query normalizes with created_at * 1000. Keep this in mind for any new queries touching both columns.

Delete Behaviour (SQLite + Qdrant consistency)

SQLite cascades handle relational cleanup, but Qdrant is a separate store and must be cleaned explicitly. Each delete path that removes embedded rows also removes the corresponding vectors:

Delete SQLite effect Qdrant cleanup
DELETE /episodes/:id Row removed semantic.deleteEpisode(id) — vector by point ID
DELETE /sessions/by-external/:id Session + episodes cascade-deleted semantic.deleteEpisodesBySession(id) — payload-filter delete on sessionId
DELETE /entities/:id Row removed, relationships cascade semantic.deleteEntity(id) — vector by point ID

All three Qdrant deletes are fire-and-forget with error logging, matching the fire-and-forget write path — a Qdrant failure logs but does not fail the delete.

The session path uses a payload-filter delete (matching on the sessionId field in the vector payload) rather than enumerating episode point IDs. This matters because the SQLite cascade has already removed the episode rows by the time cleanup runs, so there are no IDs left to enumerate — the filter deletes by payload regardless. It also cleans up any pre-existing orphans for that session as a side effect.

Not cleaned on session delete: entity vectors. Entities are shared across sessions and projects (UNIQUE(name, type) is global), so deleting one session must not remove entities that other sessions still reference. Entity vector lifecycle is tied to explicit entity deletion and the (planned) memory consolidation / orphan-cleanup pass.

Historical orphans: vectors orphaned by session deletes before this cleanup existed are not removed retroactively. A one-time sweep (scroll the episodes collection, delete points whose sessionId no longer exists in SQLite) clears them.

Project Delete Behaviour

Deleting a project runs as a transaction — it first nulls out project_id on all assigned sessions, then deletes the project. This avoids a foreign key constraint failure since sessions.project_id has no ON DELETE rule:

const doDelete = db.transaction(() => {
  db.prepare(`UPDATE sessions SET project_id = NULL WHERE project_id = ?`).run(id);
  db.prepare(`DELETE FROM projects WHERE id = ?`).run(id);
});

For all HTTP endpoints, see api-routes.md.