From 1e7c11bad8046e44b9e5f2cd42ed3124dcd8762b Mon Sep 17 00:00:00 2001 From: Storme-bit Date: Mon, 17 Aug 2026 05:31:19 -0700 Subject: [PATCH] documentation updates --- docs/reference/API-routes.md | 2 + docs/reference/testing.md | 46 +++++ docs/roadmap.md | 4 +- docs/services/entity-extraction.md | 14 ++ docs/services/memory-service.md | 67 ++++--- docs/services/retrieval-fusion.md | 43 ++++- docs/services/summarization.md | 16 +- nexusai-docs-update-2.patch | 287 +++++++++++++++++++++++++++++ 8 files changed, 451 insertions(+), 28 deletions(-) create mode 100644 docs/reference/testing.md create mode 100644 nexusai-docs-update-2.patch diff --git a/docs/reference/API-routes.md b/docs/reference/API-routes.md index 19e19a3..2d89fe2 100644 --- a/docs/reference/API-routes.md +++ b/docs/reference/API-routes.md @@ -277,6 +277,8 @@ Both fields are optional. Only provided fields are updated. | GET | /episodes/search?q=&limit= | FTS keyword search across all episodes | | GET | /episodes/:id | Get episode by ID | | GET | /sessions/:id/episodes?limit=&offset= | Paginated episodes for a session | +| GET | /sessions/:id/episode-stats | Aggregate: count, total tokens, max id (summarization threshold check) | +| GET | /sessions/:id/episodes/since/:afterId | Episodes newer than :afterId, chronological (un-summarized tail) | | DELETE | /episodes/:id | Delete episode (SQLite + Qdrant cleanup) | > Route ordering: `/episodes/search` must be defined before `/episodes/:id`. diff --git a/docs/reference/testing.md b/docs/reference/testing.md new file mode 100644 index 0000000..2ce8a07 --- /dev/null +++ b/docs/reference/testing.md @@ -0,0 +1,46 @@ +# Testing + +NexusAI has a lightweight regression suite built on Node's **built-in** test +runner (`node:test`) and assertions (`node:assert`) — no external test +dependencies. It runs offline on any node in the homelab. + +```bash +npm test # = node --test (discovers test/*.test.js at the repo root) +``` + +## What's covered + +| File | Subject | Imports the real… | +|---|---|---| +| `test/fusion.test.js` | RRF ranking math | `fuseEpisodeResults` (orchestration) | +| `test/fts-query.test.js` | Keyword tokenizer + FTS5 matching | `buildFtsQuery` (memory-service) | +| `test/llamacpp-stream.test.js` | SSE stream reassembly across chunks | `completeStream` (inference) | +| `test/summarization.test.js` | Summarization decision logic | `maybeSummarize` (orchestration) | +| `test/entity-extraction.test.js` | Greeting + regurgitation guards | `mentionedIn`, `isIgnoredName` (memory-service) | +| `test/schema.test.js` | Fresh-DB schema completeness | `schema.js` string (memory-service) | +| `test/migrations.test.js` | Migration version-stepping | `migrate` (memory-service) | + +Tests import the **real** functions rather than reimplementing logic — the +functions are exported for this purpose. External calls (Ollama, Qdrant, the +memory service) are mocked via `global.fetch`; SQLite-backed tests use the +built-in `node:sqlite` module against a throwaway in-memory database. + +## `node:sqlite` and skips + +`test/schema.test.js` and three cases in `test/fts-query.test.js` need +`node:sqlite`, which requires **Node ≥ 22.5**. On older Node they `skip` +themselves cleanly (guarded by `{ skip: !DatabaseSync }`) rather than failing, +so the suite stays green everywhere — but those checks only *verify* anything on +a node new enough to run them. A dev machine on current Node is the source of +truth for the schema and FTS matching tests. + +`node:sqlite` is still marked experimental, so runs print a one-line +`ExperimentalWarning`. It's harmless; `node --test --no-warnings` suppresses it. + +## Adding tests + +Keep the pattern: export the real function, import it, mock I/O at the boundary +(`global.fetch`) or use `node:sqlite` for DB behaviour. Prefer testing pure +logic (ranking, tokenizing, version-stepping, decision branches) over wiring. +Several of these tests were written *after* a bug slipped through — each new +class of mistake is worth a case so it can't recur silently. diff --git a/docs/roadmap.md b/docs/roadmap.md index 37735b5..097ae8c 100644 --- a/docs/roadmap.md +++ b/docs/roadmap.md @@ -68,7 +68,9 @@ The highest-leverage memory upgrade. Transforms NexusAI from "remembers conversa Multi-strategy retrieval merged into a single ranked result set. - [x] Reciprocal Rank Fusion (RRF) — merge semantic (Qdrant) + keyword (FTS5) results - [x] Configurable weights per retrieval strategy (`semanticWeight`, `keywordWeight` via `PATCH /settings`) -- [x] Score threshold retained per-strategy; FTS scoped to session/project sessions; `keywordWeight: 0` default (disabled until tuned) +- [x] Score threshold retained per-strategy; FTS scoped to session/project sessions +- [x] FTS query tokenization (`buildFtsQuery`) — tokenize + stopword-filter + quoted `OR` terms; replaced whole-phrase matching that made keyword recall nil +- [x] Fusion tuned and verified live (keyword path + project scoping confirmed via weight inversion); running `keywordWeight: 0.5` / `semanticWeight: 1.0` ### 3. Memory Consolidation Lifecycle Prevents long-term memory degradation and enables compression. diff --git a/docs/services/entity-extraction.md b/docs/services/entity-extraction.md index a9e5b7b..e16085d 100644 --- a/docs/services/entity-extraction.md +++ b/docs/services/entity-extraction.md @@ -131,6 +131,20 @@ check alone won't reject. > imperfectly — switch to `/[^\p{L}\p{N}\s]/gu` if the set gains non-ASCII > entries. +**Regurgitation guard (`mentionedIn`):** the extraction prompt feeds the model +a "known entities" hint block (the 20 most-recent entities) for spelling/type +consistency. The small model (qwen2.5:3b) will sometimes echo that list back as +if those entities appeared in the conversation — most visibly on contentless +turns (a greeting produced fake extractions of unrelated authors, game titles, +etc.). After parsing, each extracted name is checked against the actual +`userMessage + aiResponse` text (case- and whitespace-normalized substring); any +name not present is dropped before upsert. Since the prompt constrains names to +short proper nouns, a genuinely-discussed entity appears verbatim while a +regurgitated hint does not. Relationships referencing a dropped entity fall away +automatically (they resolve against the surviving `entityMap`). A prompt line +also tells the model the hint list is spelling-only — a backstop, with +`mentionedIn` as the deterministic guarantee. + ## Relationship Processing After all entities are saved, relationships are processed: diff --git a/docs/services/memory-service.md b/docs/services/memory-service.md index c103054..b2331be 100644 --- a/docs/services/memory-service.md +++ b/docs/services/memory-service.md @@ -36,8 +36,9 @@ relationship extraction and embeds results into Qdrant. ``` src/ ├── db/ -│ ├── index.js # SQLite connection + initialization + migrations -│ ├── schema.js # Table definitions, indexes, FTS5, triggers +│ ├── index.js # SQLite connection + init + migrate() + one-time FTS backfill +│ ├── migrations.js # Forward-only versioned migration runner (PRAGMA user_version) +│ ├── schema.js # Complete current shape: tables, indexes, FTS5, triggers │ ├── projects.js # Project CRUD functions │ └── summaries.js # Summary CRUD functions ├── episodic/ @@ -64,36 +65,56 @@ Eight core tables: - **summaries** — condensed episode groups for efficient context retrieval - **projects** — named groupings of sessions with `name`, `description`, `colour`, `icon`, `isolated`, `notes`, `system_prompt` -### Migrations +### Schema & Migrations -Schema changes that cannot use `CREATE TABLE IF NOT EXISTS` are applied as -idempotent migrations in `db/index.js` at startup: +`schema.js` holds the **complete current shape** — every table, column, index, +the FTS5 virtual table, and its triggers — as the single source of truth for a +fresh database. It uses `CREATE TABLE IF NOT EXISTS`, so on a fresh DB it builds +everything; on an existing DB it skips tables that already exist (and therefore +does **not** reconcile columns on old tables — that's what migrations are for). + +`db/migrations.js` is a forward-only versioned runner keyed on +`PRAGMA user_version`: ```js -try { db.exec(`ALTER TABLE sessions ADD COLUMN name TEXT`); } catch {} -try { db.exec(`ALTER TABLE sessions ADD COLUMN project_id INTEGER REFERENCES projects(id)`); } catch {} -try { db.exec(`CREATE INDEX IF NOT EXISTS idx_sessions_project ON sessions(project_id)`); } catch {} -try { db.exec(`ALTER TABLE projects ADD COLUMN isolated INTEGER NOT NULL DEFAULT 0`); } catch {} -try { db.exec(`ALTER TABLE projects ADD COLUMN notes TEXT`); } catch {} -try { db.exec(`ALTER TABLE projects ADD COLUMN system_prompt TEXT`); } catch {} -// Knowledge graph columns: -try { db.exec(`ALTER TABLE entities ADD COLUMN mention_count INTEGER NOT NULL DEFAULT 1`) } catch {} -try { db.exec(`ALTER TABLE entities ADD COLUMN confidence REAL NOT NULL DEFAULT 1.0`) } catch {} -try { db.exec(`ALTER TABLE entities ADD COLUMN source TEXT NOT NULL DEFAULT 'extraction'`) } catch {} -try { db.exec(`ALTER TABLE entities ADD COLUMN last_seen_at INTEGER`) } catch {} -try { db.exec(`ALTER TABLE relationships ADD COLUMN mention_count INTEGER NOT NULL DEFAULT 1`) } catch {} -try { db.exec(`ALTER TABLE relationships ADD COLUMN notes TEXT`) } catch {} +const migrations = [ + (_db) => {}, // v0 → v1: consolidated baseline (historical ALTERs folded into schema.js) +]; +const LATEST_VERSION = migrations.length; // derived, never hand-maintained ``` -`entity_episodes` is defined in `schema.js` itself (not a migration) since it is a new table. +`migrate(db)` reads `user_version`, applies every entry newer than it (each in a +transaction alongside its version bump), and stamps the result. A fresh DB is +built whole by `schema.js` and simply stamped to `LATEST_VERSION`; the baseline +entry is a no-op. -New migrations are always appended — never modify the schema file for existing tables since `ALTER TABLE` cannot use `IF NOT EXISTS`. +**Adding a schema change:** append a new function to the `migrations` array +(which bumps `LATEST_VERSION` automatically). Never edit an already-shipped +entry, and never edit a table in `schema.js` expecting existing DBs to pick it +up — they won't. This replaces the previous pattern of stacking silent +`try/catch ALTER TABLE` statements in `db/index.js` on every boot. + +> **Consolidation note:** the historical ALTERs were folded into `schema.js` +> rather than preserved as replayable migrations, so this assumes a fresh +> database (which is the case post-wipe). An older, pre-consolidation database +> would **not** auto-upgrade — `schema.js` skips its existing tables and the +> baseline migration is a no-op. To support upgrading old DBs, the v1 baseline +> would instead perform guarded (`ADD COLUMN if missing`) catch-up. ### FTS5 Full-Text Search -An `episodes_fts` virtual table enables keyword search across all episodes. -Three triggers (`episodes_fts_insert`, `episodes_fts_update`, `episodes_fts_delete`) -keep the FTS index automatically in sync with the episodes table. +An `episodes_fts` external-content virtual table enables keyword search across +episodes. Three triggers (`episodes_fts_insert`, `episodes_fts_update`, +`episodes_fts_delete`) keep the index in sync with the `episodes` table +automatically during normal operation. + +A one-time backfill in `db/index.js` handles the case where the FTS table is +created on a DB that already holds episodes (e.g. episodes predating FTS). It is +gated on "did `episodes_fts` not exist before this boot," checked via +`sqlite_master` **before** running the schema — not on a row-count comparison, +because `COUNT(*)` on an external-content FTS5 table proxies the content table +and cannot detect a desync. This replaced an unconditional full FTS rebuild that +previously ran on every startup. ### SQLite Configuration diff --git a/docs/services/retrieval-fusion.md b/docs/services/retrieval-fusion.md index 74971fb..35ae007 100644 --- a/docs/services/retrieval-fusion.md +++ b/docs/services/retrieval-fusion.md @@ -37,9 +37,10 @@ Fusion lives in orchestration — the service already coordinates multiple data sources, and fusion is a retrieval strategy, not a storage concern. ``` -getFusedEpisodes() -├── getSemanticEpisodes() — Qdrant embed+search → fetch full rows by ID -│ (existing path, unchanged) +getFusedEpisodes(…, queryVector, …) +├── getSemanticEpisodes(queryVector) — Qdrant search → fetch full rows by ID +│ (query embedded ONCE upstream in assembleContext and shared with entity +│ search — no longer embedded separately here) └── getFTSResults() — memory-service /episodes/search → full rows directly (skipped entirely if keywordWeight == 0) ↓ @@ -48,6 +49,10 @@ fuseEpisodeResults() — pure RRF, no I/O fusedEpisodes[] — top semanticLimit episodes by RRF score ``` +The query embedding is computed once per turn in `assembleContext` and passed +into both fused retrieval and entity search; if embedding fails, both receive +`null` and degrade to empty results rather than erroring. + ### Data Shape Consistency Both sides must enter fusion as `Episode[]` — full SQLite row objects with @@ -59,6 +64,38 @@ the same shape — and both must be filtered against `recentIds` first: FTS requests `semanticLimit * 2` results to provide headroom for the `recentIds` filter without under-serving the fusion. +## Query Tokenization + +Before FTS5 sees the query, `buildFtsQuery(query)` (in +`memory-service/src/episodic/index.js`) turns the raw message into a MATCH +expression: + +1. Lowercase and split on any non-letter/number (`/[^\p{L}\p{N}]+/u`, unicode-aware) +2. Drop stopwords (a small `FTS_STOPWORDS` set of common function words) and single-character tokens +3. Wrap each surviving token in double quotes and join with ` OR ` + +So `"How do I configure the Qdrant collection?"` becomes +`"configure" OR "qdrant" OR "collection"`. If nothing survives (an +all-stopword message like `"how do I do it?"`), it returns `null` and +`searchEpisodes` returns `[]` — keyword search sits out that turn and +semantic retrieval carries it. + +**Why this matters:** the earlier implementation quoted the *entire* message +as one FTS5 phrase, which required the whole string to appear verbatim in an +episode — so keyword recall was effectively nil for conversational queries. +Tokenizing into OR-joined terms is what makes `keywordWeight > 0` actually +contribute anything. + +**Injection safety:** quoting each token individually means any +FTS5-significant token inside the user's message (a literal `OR`, `*`, `"`, +etc.) is matched as a search term rather than parsed as an operator. This +replaces the safety the old whole-phrase quoting provided. + +The stopword set is deliberately conservative and tuned iteratively — high +frequency filler (`the`, `is`, `one`, `there`, …) is dropped, but borderline +words that can carry signal (`time`, `good`, `way`) are kept. Add to the set +when a common word is observed producing noisy matches. + ## FTS Session Scoping Without scoping, FTS5 searches across all episodes in the database. For diff --git a/docs/services/summarization.md b/docs/services/summarization.md index f8af7b4..bc51ace 100644 --- a/docs/services/summarization.md +++ b/docs/services/summarization.md @@ -12,7 +12,21 @@ the full context window with raw episodes. ## Trigger Conditions -`triggerSummary(session, allEpisodes)` calls `maybeSummarize` fire-and-forget. +`triggerSummary(session)` calls `maybeSummarize` fire-and-forget. It takes only +the `session` — it no longer receives the full episode list. Instead +`maybeSummarize` fetches exactly what it needs: + +1. `GET /sessions/:id/episode-stats` — a cheap aggregate (`COUNT`, `SUM(token_count)`, + `MAX(id)`) that gates the token threshold **without** pulling every episode row +2. Only if over threshold: `GET /sessions/:id/episodes/since/:afterId` — the + un-summarized tail (episodes newer than the last summary's range), fetched in + full for the summary prompt + +This replaced an earlier approach where `chat/index.js` fetched the entire +session (`getRecentEpisodes(session.id, 9999)`) on every message just to hand it +over — an O(session length) cost per turn. The stats query now runs on every +message; full episode text is fetched only when a summary actually fires. + `maybeSummarize` proceeds only when both conditions are met: 1. Total session token count exceeds `SUMMARIES.THRESHOLD_TOKENS` (default 200) diff --git a/nexusai-docs-update-2.patch b/nexusai-docs-update-2.patch new file mode 100644 index 0000000..c474fa8 --- /dev/null +++ b/nexusai-docs-update-2.patch @@ -0,0 +1,287 @@ +diff -ruN nexusai-baseline/docs/reference/API-routes.md nexusai/docs/reference/API-routes.md +--- nexusai-baseline/docs/reference/API-routes.md 2026-08-17 11:41:32.878153043 +0000 ++++ nexusai/docs/reference/API-routes.md 2026-08-17 12:28:03.473533125 +0000 +@@ -277,6 +277,8 @@ + | GET | /episodes/search?q=&limit= | FTS keyword search across all episodes | + | GET | /episodes/:id | Get episode by ID | + | GET | /sessions/:id/episodes?limit=&offset= | Paginated episodes for a session | ++| GET | /sessions/:id/episode-stats | Aggregate: count, total tokens, max id (summarization threshold check) | ++| GET | /sessions/:id/episodes/since/:afterId | Episodes newer than :afterId, chronological (un-summarized tail) | + | DELETE | /episodes/:id | Delete episode (SQLite + Qdrant cleanup) | + + > Route ordering: `/episodes/search` must be defined before `/episodes/:id`. +diff -ruN nexusai-baseline/docs/reference/testing.md nexusai/docs/reference/testing.md +--- nexusai-baseline/docs/reference/testing.md 1970-01-01 00:00:00.000000000 +0000 ++++ nexusai/docs/reference/testing.md 2026-08-17 12:28:47.370890122 +0000 +@@ -0,0 +1,46 @@ ++# Testing ++ ++NexusAI has a lightweight regression suite built on Node's **built-in** test ++runner (`node:test`) and assertions (`node:assert`) — no external test ++dependencies. It runs offline on any node in the homelab. ++ ++```bash ++npm test # = node --test (discovers test/*.test.js at the repo root) ++``` ++ ++## What's covered ++ ++| File | Subject | Imports the real… | ++|---|---|---| ++| `test/fusion.test.js` | RRF ranking math | `fuseEpisodeResults` (orchestration) | ++| `test/fts-query.test.js` | Keyword tokenizer + FTS5 matching | `buildFtsQuery` (memory-service) | ++| `test/llamacpp-stream.test.js` | SSE stream reassembly across chunks | `completeStream` (inference) | ++| `test/summarization.test.js` | Summarization decision logic | `maybeSummarize` (orchestration) | ++| `test/entity-extraction.test.js` | Greeting + regurgitation guards | `mentionedIn`, `isIgnoredName` (memory-service) | ++| `test/schema.test.js` | Fresh-DB schema completeness | `schema.js` string (memory-service) | ++| `test/migrations.test.js` | Migration version-stepping | `migrate` (memory-service) | ++ ++Tests import the **real** functions rather than reimplementing logic — the ++functions are exported for this purpose. External calls (Ollama, Qdrant, the ++memory service) are mocked via `global.fetch`; SQLite-backed tests use the ++built-in `node:sqlite` module against a throwaway in-memory database. ++ ++## `node:sqlite` and skips ++ ++`test/schema.test.js` and three cases in `test/fts-query.test.js` need ++`node:sqlite`, which requires **Node ≥ 22.5**. On older Node they `skip` ++themselves cleanly (guarded by `{ skip: !DatabaseSync }`) rather than failing, ++so the suite stays green everywhere — but those checks only *verify* anything on ++a node new enough to run them. A dev machine on current Node is the source of ++truth for the schema and FTS matching tests. ++ ++`node:sqlite` is still marked experimental, so runs print a one-line ++`ExperimentalWarning`. It's harmless; `node --test --no-warnings` suppresses it. ++ ++## Adding tests ++ ++Keep the pattern: export the real function, import it, mock I/O at the boundary ++(`global.fetch`) or use `node:sqlite` for DB behaviour. Prefer testing pure ++logic (ranking, tokenizing, version-stepping, decision branches) over wiring. ++Several of these tests were written *after* a bug slipped through — each new ++class of mistake is worth a case so it can't recur silently. +diff -ruN nexusai-baseline/docs/roadmap.md nexusai/docs/roadmap.md +--- nexusai-baseline/docs/roadmap.md 2026-08-17 11:41:32.878198690 +0000 ++++ nexusai/docs/roadmap.md 2026-08-17 12:28:30.788572199 +0000 +@@ -68,7 +68,9 @@ + Multi-strategy retrieval merged into a single ranked result set. + - [x] Reciprocal Rank Fusion (RRF) — merge semantic (Qdrant) + keyword (FTS5) results + - [x] Configurable weights per retrieval strategy (`semanticWeight`, `keywordWeight` via `PATCH /settings`) +-- [x] Score threshold retained per-strategy; FTS scoped to session/project sessions; `keywordWeight: 0` default (disabled until tuned) ++- [x] Score threshold retained per-strategy; FTS scoped to session/project sessions ++- [x] FTS query tokenization (`buildFtsQuery`) — tokenize + stopword-filter + quoted `OR` terms; replaced whole-phrase matching that made keyword recall nil ++- [x] Fusion tuned and verified live (keyword path + project scoping confirmed via weight inversion); running `keywordWeight: 0.5` / `semanticWeight: 1.0` + + ### 3. Memory Consolidation Lifecycle + Prevents long-term memory degradation and enables compression. +diff -ruN nexusai-baseline/docs/services/entity-extraction.md nexusai/docs/services/entity-extraction.md +--- nexusai-baseline/docs/services/entity-extraction.md 2026-08-17 11:41:32.878265351 +0000 ++++ nexusai/docs/services/entity-extraction.md 2026-08-17 12:27:44.172932395 +0000 +@@ -131,6 +131,20 @@ + > imperfectly — switch to `/[^\p{L}\p{N}\s]/gu` if the set gains non-ASCII + > entries. + ++**Regurgitation guard (`mentionedIn`):** the extraction prompt feeds the model ++a "known entities" hint block (the 20 most-recent entities) for spelling/type ++consistency. The small model (qwen2.5:3b) will sometimes echo that list back as ++if those entities appeared in the conversation — most visibly on contentless ++turns (a greeting produced fake extractions of unrelated authors, game titles, ++etc.). After parsing, each extracted name is checked against the actual ++`userMessage + aiResponse` text (case- and whitespace-normalized substring); any ++name not present is dropped before upsert. Since the prompt constrains names to ++short proper nouns, a genuinely-discussed entity appears verbatim while a ++regurgitated hint does not. Relationships referencing a dropped entity fall away ++automatically (they resolve against the surviving `entityMap`). A prompt line ++also tells the model the hint list is spelling-only — a backstop, with ++`mentionedIn` as the deterministic guarantee. ++ + ## Relationship Processing + + After all entities are saved, relationships are processed: +diff -ruN nexusai-baseline/docs/services/memory-service.md nexusai/docs/services/memory-service.md +--- nexusai-baseline/docs/services/memory-service.md 2026-08-17 11:41:32.878409053 +0000 ++++ nexusai/docs/services/memory-service.md 2026-08-17 12:27:35.716782770 +0000 +@@ -36,8 +36,9 @@ + ``` + src/ + ├── db/ +-│ ├── index.js # SQLite connection + initialization + migrations +-│ ├── schema.js # Table definitions, indexes, FTS5, triggers ++│ ├── index.js # SQLite connection + init + migrate() + one-time FTS backfill ++│ ├── migrations.js # Forward-only versioned migration runner (PRAGMA user_version) ++│ ├── schema.js # Complete current shape: tables, indexes, FTS5, triggers + │ ├── projects.js # Project CRUD functions + │ └── summaries.js # Summary CRUD functions + ├── episodic/ +@@ -64,36 +65,56 @@ + - **summaries** — condensed episode groups for efficient context retrieval + - **projects** — named groupings of sessions with `name`, `description`, `colour`, `icon`, `isolated`, `notes`, `system_prompt` + +-### Migrations ++### Schema & Migrations + +-Schema changes that cannot use `CREATE TABLE IF NOT EXISTS` are applied as +-idempotent migrations in `db/index.js` at startup: ++`schema.js` holds the **complete current shape** — every table, column, index, ++the FTS5 virtual table, and its triggers — as the single source of truth for a ++fresh database. It uses `CREATE TABLE IF NOT EXISTS`, so on a fresh DB it builds ++everything; on an existing DB it skips tables that already exist (and therefore ++does **not** reconcile columns on old tables — that's what migrations are for). ++ ++`db/migrations.js` is a forward-only versioned runner keyed on ++`PRAGMA user_version`: + + ```js +-try { db.exec(`ALTER TABLE sessions ADD COLUMN name TEXT`); } catch {} +-try { db.exec(`ALTER TABLE sessions ADD COLUMN project_id INTEGER REFERENCES projects(id)`); } catch {} +-try { db.exec(`CREATE INDEX IF NOT EXISTS idx_sessions_project ON sessions(project_id)`); } catch {} +-try { db.exec(`ALTER TABLE projects ADD COLUMN isolated INTEGER NOT NULL DEFAULT 0`); } catch {} +-try { db.exec(`ALTER TABLE projects ADD COLUMN notes TEXT`); } catch {} +-try { db.exec(`ALTER TABLE projects ADD COLUMN system_prompt TEXT`); } catch {} +-// Knowledge graph columns: +-try { db.exec(`ALTER TABLE entities ADD COLUMN mention_count INTEGER NOT NULL DEFAULT 1`) } catch {} +-try { db.exec(`ALTER TABLE entities ADD COLUMN confidence REAL NOT NULL DEFAULT 1.0`) } catch {} +-try { db.exec(`ALTER TABLE entities ADD COLUMN source TEXT NOT NULL DEFAULT 'extraction'`) } catch {} +-try { db.exec(`ALTER TABLE entities ADD COLUMN last_seen_at INTEGER`) } catch {} +-try { db.exec(`ALTER TABLE relationships ADD COLUMN mention_count INTEGER NOT NULL DEFAULT 1`) } catch {} +-try { db.exec(`ALTER TABLE relationships ADD COLUMN notes TEXT`) } catch {} ++const migrations = [ ++ (_db) => {}, // v0 → v1: consolidated baseline (historical ALTERs folded into schema.js) ++]; ++const LATEST_VERSION = migrations.length; // derived, never hand-maintained + ``` + +-`entity_episodes` is defined in `schema.js` itself (not a migration) since it is a new table. +- +-New migrations are always appended — never modify the schema file for existing tables since `ALTER TABLE` cannot use `IF NOT EXISTS`. ++`migrate(db)` reads `user_version`, applies every entry newer than it (each in a ++transaction alongside its version bump), and stamps the result. A fresh DB is ++built whole by `schema.js` and simply stamped to `LATEST_VERSION`; the baseline ++entry is a no-op. ++ ++**Adding a schema change:** append a new function to the `migrations` array ++(which bumps `LATEST_VERSION` automatically). Never edit an already-shipped ++entry, and never edit a table in `schema.js` expecting existing DBs to pick it ++up — they won't. This replaces the previous pattern of stacking silent ++`try/catch ALTER TABLE` statements in `db/index.js` on every boot. ++ ++> **Consolidation note:** the historical ALTERs were folded into `schema.js` ++> rather than preserved as replayable migrations, so this assumes a fresh ++> database (which is the case post-wipe). An older, pre-consolidation database ++> would **not** auto-upgrade — `schema.js` skips its existing tables and the ++> baseline migration is a no-op. To support upgrading old DBs, the v1 baseline ++> would instead perform guarded (`ADD COLUMN if missing`) catch-up. + + ### FTS5 Full-Text Search + +-An `episodes_fts` virtual table enables keyword search across all episodes. +-Three triggers (`episodes_fts_insert`, `episodes_fts_update`, `episodes_fts_delete`) +-keep the FTS index automatically in sync with the episodes table. ++An `episodes_fts` external-content virtual table enables keyword search across ++episodes. Three triggers (`episodes_fts_insert`, `episodes_fts_update`, ++`episodes_fts_delete`) keep the index in sync with the `episodes` table ++automatically during normal operation. ++ ++A one-time backfill in `db/index.js` handles the case where the FTS table is ++created on a DB that already holds episodes (e.g. episodes predating FTS). It is ++gated on "did `episodes_fts` not exist before this boot," checked via ++`sqlite_master` **before** running the schema — not on a row-count comparison, ++because `COUNT(*)` on an external-content FTS5 table proxies the content table ++and cannot detect a desync. This replaced an unconditional full FTS rebuild that ++previously ran on every startup. + + ### SQLite Configuration + +diff -ruN nexusai-baseline/docs/services/retrieval-fusion.md nexusai/docs/services/retrieval-fusion.md +--- nexusai-baseline/docs/services/retrieval-fusion.md 2026-08-17 11:41:32.878594989 +0000 ++++ nexusai/docs/services/retrieval-fusion.md 2026-08-17 12:27:07.267654899 +0000 +@@ -37,9 +37,10 @@ + sources, and fusion is a retrieval strategy, not a storage concern. + + ``` +-getFusedEpisodes() +-├── getSemanticEpisodes() — Qdrant embed+search → fetch full rows by ID +-│ (existing path, unchanged) ++getFusedEpisodes(…, queryVector, …) ++├── getSemanticEpisodes(queryVector) — Qdrant search → fetch full rows by ID ++│ (query embedded ONCE upstream in assembleContext and shared with entity ++│ search — no longer embedded separately here) + └── getFTSResults() — memory-service /episodes/search → full rows directly + (skipped entirely if keywordWeight == 0) + ↓ +@@ -48,6 +49,10 @@ + fusedEpisodes[] — top semanticLimit episodes by RRF score + ``` + ++The query embedding is computed once per turn in `assembleContext` and passed ++into both fused retrieval and entity search; if embedding fails, both receive ++`null` and degrade to empty results rather than erroring. ++ + ### Data Shape Consistency + + Both sides must enter fusion as `Episode[]` — full SQLite row objects with +@@ -59,6 +64,38 @@ + FTS requests `semanticLimit * 2` results to provide headroom for the + `recentIds` filter without under-serving the fusion. + ++## Query Tokenization ++ ++Before FTS5 sees the query, `buildFtsQuery(query)` (in ++`memory-service/src/episodic/index.js`) turns the raw message into a MATCH ++expression: ++ ++1. Lowercase and split on any non-letter/number (`/[^\p{L}\p{N}]+/u`, unicode-aware) ++2. Drop stopwords (a small `FTS_STOPWORDS` set of common function words) and single-character tokens ++3. Wrap each surviving token in double quotes and join with ` OR ` ++ ++So `"How do I configure the Qdrant collection?"` becomes ++`"configure" OR "qdrant" OR "collection"`. If nothing survives (an ++all-stopword message like `"how do I do it?"`), it returns `null` and ++`searchEpisodes` returns `[]` — keyword search sits out that turn and ++semantic retrieval carries it. ++ ++**Why this matters:** the earlier implementation quoted the *entire* message ++as one FTS5 phrase, which required the whole string to appear verbatim in an ++episode — so keyword recall was effectively nil for conversational queries. ++Tokenizing into OR-joined terms is what makes `keywordWeight > 0` actually ++contribute anything. ++ ++**Injection safety:** quoting each token individually means any ++FTS5-significant token inside the user's message (a literal `OR`, `*`, `"`, ++etc.) is matched as a search term rather than parsed as an operator. This ++replaces the safety the old whole-phrase quoting provided. ++ ++The stopword set is deliberately conservative and tuned iteratively — high ++frequency filler (`the`, `is`, `one`, `there`, …) is dropped, but borderline ++words that can carry signal (`time`, `good`, `way`) are kept. Add to the set ++when a common word is observed producing noisy matches. ++ + ## FTS Session Scoping + + Without scoping, FTS5 searches across all episodes in the database. For +diff -ruN nexusai-baseline/docs/services/summarization.md nexusai/docs/services/summarization.md +--- nexusai-baseline/docs/services/summarization.md 2026-08-17 11:41:32.878899850 +0000 ++++ nexusai/docs/services/summarization.md 2026-08-17 12:27:55.087746482 +0000 +@@ -12,7 +12,21 @@ + + ## Trigger Conditions + +-`triggerSummary(session, allEpisodes)` calls `maybeSummarize` fire-and-forget. ++`triggerSummary(session)` calls `maybeSummarize` fire-and-forget. It takes only ++the `session` — it no longer receives the full episode list. Instead ++`maybeSummarize` fetches exactly what it needs: ++ ++1. `GET /sessions/:id/episode-stats` — a cheap aggregate (`COUNT`, `SUM(token_count)`, ++ `MAX(id)`) that gates the token threshold **without** pulling every episode row ++2. Only if over threshold: `GET /sessions/:id/episodes/since/:afterId` — the ++ un-summarized tail (episodes newer than the last summary's range), fetched in ++ full for the summary prompt ++ ++This replaced an earlier approach where `chat/index.js` fetched the entire ++session (`getRecentEpisodes(session.id, 9999)`) on every message just to hand it ++over — an O(session length) cost per turn. The stats query now runs on every ++message; full episode text is fetched only when a summary actually fires. ++ + `maybeSummarize` proceeds only when both conditions are met: + + 1. Total session token count exceeds `SUMMARIES.THRESHOLD_TOKENS` (default 200)