255 lines
12 KiB
Markdown
255 lines
12 KiB
Markdown
# Memory Service
|
|
|
|
**Package:** `@nexusai/memory-service`
|
|
**Location:** `packages/memory-service`
|
|
**Deployed on:** Mini PC 1 (192.168.0.81)
|
|
**Port:** 3002
|
|
|
|
## Purpose
|
|
|
|
Responsible for all reading and writing of long-term memory. Acts as the
|
|
sole interface to both SQLite and Qdrant — no other service accesses these
|
|
stores directly. On episode creation, automatically triggers entity and
|
|
relationship extraction and embeds results into Qdrant.
|
|
|
|
## Dependencies
|
|
|
|
- `express` — HTTP API
|
|
- `better-sqlite3` — SQLite driver
|
|
- `@qdrant/js-client-rest` — Qdrant vector store client
|
|
- `dotenv` — environment variable loading
|
|
- `@nexusai/shared` — shared utilities and constants
|
|
|
|
## Environment Variables
|
|
|
|
| Variable | Required | Default | Description |
|
|
|---|---|---|---|
|
|
| PORT | No | 3002 | Port to listen on |
|
|
| SQLITE_PATH | Yes | — | Path to SQLite database file |
|
|
| QDRANT_URL | No | http://localhost:6333 | Qdrant instance URL |
|
|
| EMBEDDING_SERVICE_URL | No | http://localhost:3003 | Embedding service URL |
|
|
| EXTRACTION_URL | No | http://localhost:11434 | Ollama URL for entity extraction |
|
|
| EXTRACTION_MODEL | No | qwen2.5:3b | Ollama model used for entity extraction |
|
|
|
|
## Internal Structure
|
|
|
|
```
|
|
src/
|
|
├── db/
|
|
│ ├── index.js # SQLite connection + init + migrate() + one-time FTS backfill
|
|
│ ├── migrations.js # Forward-only versioned migration runner (PRAGMA user_version)
|
|
│ ├── schema.js # Complete current shape: tables, indexes, FTS5, triggers
|
|
│ ├── projects.js # Project CRUD functions
|
|
│ └── summaries.js # Summary CRUD functions
|
|
├── episodic/
|
|
│ └── index.js # Session + episode CRUD, FTS search, embedding write path
|
|
├── semantic/
|
|
│ └── index.js # Qdrant collection management, upsert, search, delete
|
|
├── entities/
|
|
│ ├── index.js # Entity + relationship CRUD (upsert, mention tracking)
|
|
│ └── extraction.js # Automatic entity + relationship extraction via qwen2.5:3b
|
|
├── graph/
|
|
│ └── index.js # Knowledge graph traversal (neighborhood queries, recursive CTE)
|
|
└── index.js # Express app + all route definitions
|
|
```
|
|
|
|
## SQLite Schema
|
|
|
|
Eight core tables:
|
|
|
|
- **sessions** — top-level conversation containers. Fields: `external_id`, `name`, `project_id`, `metadata`
|
|
- **episodes** — individual exchanges (user message + AI response) tied to a session
|
|
- **entities** — named things the system learns about (people, places, concepts, etc.). Fields include `mention_count`, `confidence`, `source`, `last_seen_at`
|
|
- **relationships** — directional labeled links between entities (`from_id`, `to_id`, `label`). Fields include `mention_count`, `notes`
|
|
- **entity_episodes** — join table linking entities to the episodes where they were extracted. Used for provenance and orphan cleanup
|
|
- **summaries** — condensed episode groups for efficient context retrieval
|
|
- **projects** — named groupings of sessions with `name`, `description`, `colour`, `icon`, `isolated`, `notes`, `system_prompt`
|
|
|
|
### Schema & Migrations
|
|
|
|
`schema.js` holds the **complete current shape** — every table, column, index,
|
|
the FTS5 virtual table, and its triggers — as the single source of truth for a
|
|
fresh database. It uses `CREATE TABLE IF NOT EXISTS`, so on a fresh DB it builds
|
|
everything; on an existing DB it skips tables that already exist (and therefore
|
|
does **not** reconcile columns on old tables — that's what migrations are for).
|
|
|
|
`db/migrations.js` is a forward-only versioned runner keyed on
|
|
`PRAGMA user_version`:
|
|
|
|
```js
|
|
const migrations = [
|
|
(_db) => {}, // v0 → v1: consolidated baseline (historical ALTERs folded into schema.js)
|
|
];
|
|
const LATEST_VERSION = migrations.length; // derived, never hand-maintained
|
|
```
|
|
|
|
`migrate(db)` reads `user_version`, applies every entry newer than it (each in a
|
|
transaction alongside its version bump), and stamps the result. A fresh DB is
|
|
built whole by `schema.js` and simply stamped to `LATEST_VERSION`; the baseline
|
|
entry is a no-op.
|
|
|
|
**Adding a schema change:** append a new function to the `migrations` array
|
|
(which bumps `LATEST_VERSION` automatically). Never edit an already-shipped
|
|
entry, and never edit a table in `schema.js` expecting existing DBs to pick it
|
|
up — they won't. This replaces the previous pattern of stacking silent
|
|
`try/catch ALTER TABLE` statements in `db/index.js` on every boot.
|
|
|
|
> **Consolidation note:** the historical ALTERs were folded into `schema.js`
|
|
> rather than preserved as replayable migrations, so this assumes a fresh
|
|
> database (which is the case post-wipe). An older, pre-consolidation database
|
|
> would **not** auto-upgrade — `schema.js` skips its existing tables and the
|
|
> baseline migration is a no-op. To support upgrading old DBs, the v1 baseline
|
|
> would instead perform guarded (`ADD COLUMN if missing`) catch-up.
|
|
|
|
### FTS5 Full-Text Search
|
|
|
|
An `episodes_fts` external-content virtual table enables keyword search across
|
|
episodes. Three triggers (`episodes_fts_insert`, `episodes_fts_update`,
|
|
`episodes_fts_delete`) keep the index in sync with the `episodes` table
|
|
automatically during normal operation.
|
|
|
|
A one-time backfill in `db/index.js` handles the case where the FTS table is
|
|
created on a DB that already holds episodes (e.g. episodes predating FTS). It is
|
|
gated on "did `episodes_fts` not exist before this boot," checked via
|
|
`sqlite_master` **before** running the schema — not on a row-count comparison,
|
|
because `COUNT(*)` on an external-content FTS5 table proxies the content table
|
|
and cannot detect a desync. This replaced an unconditional full FTS rebuild that
|
|
previously ran on every startup.
|
|
|
|
### SQLite Configuration
|
|
|
|
- `journal_mode = WAL` — non-blocking reads during writes
|
|
- `foreign_keys = ON` — enforces referential integrity and cascade deletes
|
|
- PRAGMAs set via `db.pragma()`, not `db.exec()`
|
|
|
|
### Dynamic Updates
|
|
|
|
Both `updateSession` and `updateProject` build their `SET` clause dynamically
|
|
from only the fields passed — prevents partial updates from overwriting fields
|
|
that weren't touched.
|
|
|
|
`updateProject` allowlist:
|
|
```js
|
|
const allowed = ['name', 'description', 'colour', 'icon', 'isolated', 'notes', 'system_prompt'];
|
|
```
|
|
|
|
## Qdrant / Semantic Layer
|
|
|
|
Three Qdrant collections are initialized on service startup via `semantic.initCollections()`:
|
|
|
|
| Collection | Purpose |
|
|
|---|---|
|
|
| `episodes` | Embeddings for individual conversation exchanges |
|
|
| `entities` | Embeddings for named entities |
|
|
| `summaries` | Embeddings for condensed episode summaries |
|
|
|
|
All collections use **768-dimension vectors** with **Cosine similarity**,
|
|
matching `nomic-embed-text` via Ollama. Vector size and distance metric are
|
|
defined in `@nexusai/shared` — not hardcoded here.
|
|
|
|
`initCollections()` iterates `Object.values(COLLECTIONS)` and creates any
|
|
collection that doesn't already exist at startup — all three collections are
|
|
guaranteed to exist before any requests are handled.
|
|
|
|
Each collection exposes upsert, search (with optional Qdrant filter), and
|
|
delete operations. The `wait: true` flag is used on all writes.
|
|
|
|
## Embedding Write Path
|
|
|
|
When a new episode is created:
|
|
|
|
1. Episode saved to SQLite synchronously — response returned immediately
|
|
2. User message + AI response combined: `User: ...\nAssistant: ...`
|
|
3. Text sent to embedding service (`POST /embed`)
|
|
4. Vector upserted into `episodes` Qdrant collection with payload `{ sessionId, createdAt }`
|
|
|
|
This step is **fire-and-forget** — if embedding fails, the episode is still
|
|
saved and searchable via FTS. The error is logged but not surfaced.
|
|
|
|
> The Qdrant payload stores `sessionId` (the internal integer ID). See
|
|
> `memory-isolation.md` for how project-level filtering works.
|
|
|
|
## Entity Layer
|
|
|
|
Entities and relationships use upsert semantics with composite unique
|
|
constraints to prevent duplicates:
|
|
|
|
- `UNIQUE(name, type)` on entities — conflict increments `mention_count` and updates `last_seen_at`
|
|
- `UNIQUE(from_id, to_id, label)` on relationships — conflict increments `mention_count` and preserves existing `notes`
|
|
- `ON DELETE CASCADE` on relationship foreign keys
|
|
|
|
After each episode is saved, `extraction.js` automatically extracts named
|
|
entities **and relationships** from the conversation using `qwen2.5:3b` on
|
|
Ollama — fire-and-forget. Each saved entity is also linked to the episode
|
|
via the `entity_episodes` join table.
|
|
|
|
> For full details on the extraction pipeline and JSON format, see `entity-extraction.md`.
|
|
> For the knowledge graph traversal layer, see `knowledge-graph.md`.
|
|
|
|
## Knowledge Graph Layer
|
|
|
|
`src/graph/index.js` provides SQLite-based graph traversal over the entities
|
|
and relationships tables. Two functions are exposed via HTTP:
|
|
|
|
- **`getNeighborhood(entityId, depth)`** — recursive CTE traversal, bidirectional, returns `{ nodes, edges }`
|
|
- **`getEntityNeighbors(entityIds[])`** — bulk 1-hop traversal for orchestration context assembly
|
|
|
|
> For design rationale, traversal queries, and integration with orchestration, see `knowledge-graph.md`.
|
|
|
|
## Summaries Layer
|
|
|
|
Session summaries are generated by `orchestration-service/src/services/summarization.js`
|
|
after each episode write and stored here via `POST /summaries`. The memory
|
|
service is responsible only for CRUD — generation logic lives in orchestration.
|
|
|
|
> For full details on trigger conditions, prompt format, cumulative updates,
|
|
> and ChatML token stripping, see `summarization.md`.
|
|
|
|
## Delete Behaviour (SQLite + Qdrant consistency)
|
|
|
|
SQLite cascades handle relational cleanup, but Qdrant is a separate store and
|
|
must be cleaned explicitly. Each delete path that removes embedded rows also
|
|
removes the corresponding vectors:
|
|
|
|
| Delete | SQLite effect | Qdrant cleanup |
|
|
|---|---|---|
|
|
| `DELETE /episodes/:id` | Row removed | `semantic.deleteEpisode(id)` — vector by point ID |
|
|
| `DELETE /sessions/by-external/:id` | Session + episodes cascade-deleted | `semantic.deleteEpisodesBySession(id)` — **payload-filter** delete on `sessionId` |
|
|
| `DELETE /entities/:id` | Row removed, relationships cascade | `semantic.deleteEntity(id)` — vector by point ID |
|
|
|
|
All three Qdrant deletes are **fire-and-forget** with error logging, matching
|
|
the fire-and-forget write path — a Qdrant failure logs but does not fail the
|
|
delete.
|
|
|
|
The session path uses a **payload-filter** delete (matching on the `sessionId`
|
|
field in the vector payload) rather than enumerating episode point IDs. This
|
|
matters because the SQLite cascade has already removed the episode rows by the
|
|
time cleanup runs, so there are no IDs left to enumerate — the filter deletes
|
|
by payload regardless. It also cleans up any pre-existing orphans for that
|
|
session as a side effect.
|
|
|
|
> **Not cleaned on session delete:** entity vectors. Entities are shared across
|
|
> sessions and projects (`UNIQUE(name, type)` is global), so deleting one
|
|
> session must not remove entities that other sessions still reference. Entity
|
|
> vector lifecycle is tied to explicit entity deletion and the (planned) memory
|
|
> consolidation / orphan-cleanup pass.
|
|
|
|
> **Historical orphans:** vectors orphaned by session deletes *before* this
|
|
> cleanup existed are not removed retroactively. A one-time sweep (scroll the
|
|
> `episodes` collection, delete points whose `sessionId` no longer exists in
|
|
> SQLite) clears them.
|
|
|
|
## Project Delete Behaviour
|
|
|
|
Deleting a project runs as a transaction — it first nulls out `project_id`
|
|
on all assigned sessions, then deletes the project. This avoids a foreign
|
|
key constraint failure since `sessions.project_id` has no `ON DELETE` rule:
|
|
|
|
```js
|
|
const doDelete = db.transaction(() => {
|
|
db.prepare(`UPDATE sessions SET project_id = NULL WHERE project_id = ?`).run(id);
|
|
db.prepare(`DELETE FROM projects WHERE id = ?`).run(id);
|
|
});
|
|
```
|
|
|
|
For all HTTP endpoints, see `api-routes.md`. |