documentation updates

This commit is contained in:
Storme-bit
2026-08-17 05:31:19 -07:00
parent 1df6acc427
commit 1e7c11bad8
8 changed files with 451 additions and 28 deletions
+2
View File
@@ -277,6 +277,8 @@ Both fields are optional. Only provided fields are updated.
| GET | /episodes/search?q=&limit= | FTS keyword search across all episodes |
| GET | /episodes/:id | Get episode by ID |
| GET | /sessions/:id/episodes?limit=&offset= | Paginated episodes for a session |
| GET | /sessions/:id/episode-stats | Aggregate: count, total tokens, max id (summarization threshold check) |
| GET | /sessions/:id/episodes/since/:afterId | Episodes newer than :afterId, chronological (un-summarized tail) |
| DELETE | /episodes/:id | Delete episode (SQLite + Qdrant cleanup) |
> Route ordering: `/episodes/search` must be defined before `/episodes/:id`.
+46
View File
@@ -0,0 +1,46 @@
# Testing
NexusAI has a lightweight regression suite built on Node's **built-in** test
runner (`node:test`) and assertions (`node:assert`) — no external test
dependencies. It runs offline on any node in the homelab.
```bash
npm test # = node --test (discovers test/*.test.js at the repo root)
```
## What's covered
| File | Subject | Imports the real… |
|---|---|---|
| `test/fusion.test.js` | RRF ranking math | `fuseEpisodeResults` (orchestration) |
| `test/fts-query.test.js` | Keyword tokenizer + FTS5 matching | `buildFtsQuery` (memory-service) |
| `test/llamacpp-stream.test.js` | SSE stream reassembly across chunks | `completeStream` (inference) |
| `test/summarization.test.js` | Summarization decision logic | `maybeSummarize` (orchestration) |
| `test/entity-extraction.test.js` | Greeting + regurgitation guards | `mentionedIn`, `isIgnoredName` (memory-service) |
| `test/schema.test.js` | Fresh-DB schema completeness | `schema.js` string (memory-service) |
| `test/migrations.test.js` | Migration version-stepping | `migrate` (memory-service) |
Tests import the **real** functions rather than reimplementing logic — the
functions are exported for this purpose. External calls (Ollama, Qdrant, the
memory service) are mocked via `global.fetch`; SQLite-backed tests use the
built-in `node:sqlite` module against a throwaway in-memory database.
## `node:sqlite` and skips
`test/schema.test.js` and three cases in `test/fts-query.test.js` need
`node:sqlite`, which requires **Node ≥ 22.5**. On older Node they `skip`
themselves cleanly (guarded by `{ skip: !DatabaseSync }`) rather than failing,
so the suite stays green everywhere — but those checks only *verify* anything on
a node new enough to run them. A dev machine on current Node is the source of
truth for the schema and FTS matching tests.
`node:sqlite` is still marked experimental, so runs print a one-line
`ExperimentalWarning`. It's harmless; `node --test --no-warnings` suppresses it.
## Adding tests
Keep the pattern: export the real function, import it, mock I/O at the boundary
(`global.fetch`) or use `node:sqlite` for DB behaviour. Prefer testing pure
logic (ranking, tokenizing, version-stepping, decision branches) over wiring.
Several of these tests were written *after* a bug slipped through — each new
class of mistake is worth a case so it can't recur silently.
+3 -1
View File
@@ -68,7 +68,9 @@ The highest-leverage memory upgrade. Transforms NexusAI from "remembers conversa
Multi-strategy retrieval merged into a single ranked result set.
- [x] Reciprocal Rank Fusion (RRF) — merge semantic (Qdrant) + keyword (FTS5) results
- [x] Configurable weights per retrieval strategy (`semanticWeight`, `keywordWeight` via `PATCH /settings`)
- [x] Score threshold retained per-strategy; FTS scoped to session/project sessions; `keywordWeight: 0` default (disabled until tuned)
- [x] Score threshold retained per-strategy; FTS scoped to session/project sessions
- [x] FTS query tokenization (`buildFtsQuery`) — tokenize + stopword-filter + quoted `OR` terms; replaced whole-phrase matching that made keyword recall nil
- [x] Fusion tuned and verified live (keyword path + project scoping confirmed via weight inversion); running `keywordWeight: 0.5` / `semanticWeight: 1.0`
### 3. Memory Consolidation Lifecycle
Prevents long-term memory degradation and enables compression.
+14
View File
@@ -131,6 +131,20 @@ check alone won't reject.
> imperfectly — switch to `/[^\p{L}\p{N}\s]/gu` if the set gains non-ASCII
> entries.
**Regurgitation guard (`mentionedIn`):** the extraction prompt feeds the model
a "known entities" hint block (the 20 most-recent entities) for spelling/type
consistency. The small model (qwen2.5:3b) will sometimes echo that list back as
if those entities appeared in the conversation — most visibly on contentless
turns (a greeting produced fake extractions of unrelated authors, game titles,
etc.). After parsing, each extracted name is checked against the actual
`userMessage + aiResponse` text (case- and whitespace-normalized substring); any
name not present is dropped before upsert. Since the prompt constrains names to
short proper nouns, a genuinely-discussed entity appears verbatim while a
regurgitated hint does not. Relationships referencing a dropped entity fall away
automatically (they resolve against the surviving `entityMap`). A prompt line
also tells the model the hint list is spelling-only — a backstop, with
`mentionedIn` as the deterministic guarantee.
## Relationship Processing
After all entities are saved, relationships are processed:
+44 -23
View File
@@ -36,8 +36,9 @@ relationship extraction and embeds results into Qdrant.
```
src/
├── db/
│ ├── index.js # SQLite connection + initialization + migrations
│ ├── schema.js # Table definitions, indexes, FTS5, triggers
│ ├── index.js # SQLite connection + init + migrate() + one-time FTS backfill
│ ├── migrations.js # Forward-only versioned migration runner (PRAGMA user_version)
│ ├── schema.js # Complete current shape: tables, indexes, FTS5, triggers
│ ├── projects.js # Project CRUD functions
│ └── summaries.js # Summary CRUD functions
├── episodic/
@@ -64,36 +65,56 @@ Eight core tables:
- **summaries** — condensed episode groups for efficient context retrieval
- **projects** — named groupings of sessions with `name`, `description`, `colour`, `icon`, `isolated`, `notes`, `system_prompt`
### Migrations
### Schema & Migrations
Schema changes that cannot use `CREATE TABLE IF NOT EXISTS` are applied as
idempotent migrations in `db/index.js` at startup:
`schema.js` holds the **complete current shape** — every table, column, index,
the FTS5 virtual table, and its triggers — as the single source of truth for a
fresh database. It uses `CREATE TABLE IF NOT EXISTS`, so on a fresh DB it builds
everything; on an existing DB it skips tables that already exist (and therefore
does **not** reconcile columns on old tables — that's what migrations are for).
`db/migrations.js` is a forward-only versioned runner keyed on
`PRAGMA user_version`:
```js
try { db.exec(`ALTER TABLE sessions ADD COLUMN name TEXT`); } catch {}
try { db.exec(`ALTER TABLE sessions ADD COLUMN project_id INTEGER REFERENCES projects(id)`); } catch {}
try { db.exec(`CREATE INDEX IF NOT EXISTS idx_sessions_project ON sessions(project_id)`); } catch {}
try { db.exec(`ALTER TABLE projects ADD COLUMN isolated INTEGER NOT NULL DEFAULT 0`); } catch {}
try { db.exec(`ALTER TABLE projects ADD COLUMN notes TEXT`); } catch {}
try { db.exec(`ALTER TABLE projects ADD COLUMN system_prompt TEXT`); } catch {}
// Knowledge graph columns:
try { db.exec(`ALTER TABLE entities ADD COLUMN mention_count INTEGER NOT NULL DEFAULT 1`) } catch {}
try { db.exec(`ALTER TABLE entities ADD COLUMN confidence REAL NOT NULL DEFAULT 1.0`) } catch {}
try { db.exec(`ALTER TABLE entities ADD COLUMN source TEXT NOT NULL DEFAULT 'extraction'`) } catch {}
try { db.exec(`ALTER TABLE entities ADD COLUMN last_seen_at INTEGER`) } catch {}
try { db.exec(`ALTER TABLE relationships ADD COLUMN mention_count INTEGER NOT NULL DEFAULT 1`) } catch {}
try { db.exec(`ALTER TABLE relationships ADD COLUMN notes TEXT`) } catch {}
const migrations = [
(_db) => {}, // v0 → v1: consolidated baseline (historical ALTERs folded into schema.js)
];
const LATEST_VERSION = migrations.length; // derived, never hand-maintained
```
`entity_episodes` is defined in `schema.js` itself (not a migration) since it is a new table.
`migrate(db)` reads `user_version`, applies every entry newer than it (each in a
transaction alongside its version bump), and stamps the result. A fresh DB is
built whole by `schema.js` and simply stamped to `LATEST_VERSION`; the baseline
entry is a no-op.
New migrations are always appended — never modify the schema file for existing tables since `ALTER TABLE` cannot use `IF NOT EXISTS`.
**Adding a schema change:** append a new function to the `migrations` array
(which bumps `LATEST_VERSION` automatically). Never edit an already-shipped
entry, and never edit a table in `schema.js` expecting existing DBs to pick it
up — they won't. This replaces the previous pattern of stacking silent
`try/catch ALTER TABLE` statements in `db/index.js` on every boot.
> **Consolidation note:** the historical ALTERs were folded into `schema.js`
> rather than preserved as replayable migrations, so this assumes a fresh
> database (which is the case post-wipe). An older, pre-consolidation database
> would **not** auto-upgrade — `schema.js` skips its existing tables and the
> baseline migration is a no-op. To support upgrading old DBs, the v1 baseline
> would instead perform guarded (`ADD COLUMN if missing`) catch-up.
### FTS5 Full-Text Search
An `episodes_fts` virtual table enables keyword search across all episodes.
Three triggers (`episodes_fts_insert`, `episodes_fts_update`, `episodes_fts_delete`)
keep the FTS index automatically in sync with the episodes table.
An `episodes_fts` external-content virtual table enables keyword search across
episodes. Three triggers (`episodes_fts_insert`, `episodes_fts_update`,
`episodes_fts_delete`) keep the index in sync with the `episodes` table
automatically during normal operation.
A one-time backfill in `db/index.js` handles the case where the FTS table is
created on a DB that already holds episodes (e.g. episodes predating FTS). It is
gated on "did `episodes_fts` not exist before this boot," checked via
`sqlite_master` **before** running the schema — not on a row-count comparison,
because `COUNT(*)` on an external-content FTS5 table proxies the content table
and cannot detect a desync. This replaced an unconditional full FTS rebuild that
previously ran on every startup.
### SQLite Configuration
+40 -3
View File
@@ -37,9 +37,10 @@ Fusion lives in orchestration — the service already coordinates multiple data
sources, and fusion is a retrieval strategy, not a storage concern.
```
getFusedEpisodes()
├── getSemanticEpisodes() — Qdrant embed+search → fetch full rows by ID
│ (existing path, unchanged)
getFusedEpisodes(…, queryVector, …)
├── getSemanticEpisodes(queryVector) — Qdrant search → fetch full rows by ID
│ (query embedded ONCE upstream in assembleContext and shared with entity
│ search — no longer embedded separately here)
└── getFTSResults() — memory-service /episodes/search → full rows directly
(skipped entirely if keywordWeight == 0)
↓
@@ -48,6 +49,10 @@ fuseEpisodeResults() — pure RRF, no I/O
fusedEpisodes[] — top semanticLimit episodes by RRF score
```
The query embedding is computed once per turn in `assembleContext` and passed
into both fused retrieval and entity search; if embedding fails, both receive
`null` and degrade to empty results rather than erroring.
### Data Shape Consistency
Both sides must enter fusion as `Episode[]` — full SQLite row objects with
@@ -59,6 +64,38 @@ the same shape — and both must be filtered against `recentIds` first:
FTS requests `semanticLimit * 2` results to provide headroom for the
`recentIds` filter without under-serving the fusion.
## Query Tokenization
Before FTS5 sees the query, `buildFtsQuery(query)` (in
`memory-service/src/episodic/index.js`) turns the raw message into a MATCH
expression:
1. Lowercase and split on any non-letter/number (`/[^\p{L}\p{N}]+/u`, unicode-aware)
2. Drop stopwords (a small `FTS_STOPWORDS` set of common function words) and single-character tokens
3. Wrap each surviving token in double quotes and join with ` OR `
So `"How do I configure the Qdrant collection?"` becomes
`"configure" OR "qdrant" OR "collection"`. If nothing survives (an
all-stopword message like `"how do I do it?"`), it returns `null` and
`searchEpisodes` returns `[]` — keyword search sits out that turn and
semantic retrieval carries it.
**Why this matters:** the earlier implementation quoted the *entire* message
as one FTS5 phrase, which required the whole string to appear verbatim in an
episode — so keyword recall was effectively nil for conversational queries.
Tokenizing into OR-joined terms is what makes `keywordWeight > 0` actually
contribute anything.
**Injection safety:** quoting each token individually means any
FTS5-significant token inside the user's message (a literal `OR`, `*`, `"`,
etc.) is matched as a search term rather than parsed as an operator. This
replaces the safety the old whole-phrase quoting provided.
The stopword set is deliberately conservative and tuned iteratively — high
frequency filler (`the`, `is`, `one`, `there`, …) is dropped, but borderline
words that can carry signal (`time`, `good`, `way`) are kept. Add to the set
when a common word is observed producing noisy matches.
## FTS Session Scoping
Without scoping, FTS5 searches across all episodes in the database. For
+15 -1
View File
@@ -12,7 +12,21 @@ the full context window with raw episodes.
## Trigger Conditions
`triggerSummary(session, allEpisodes)` calls `maybeSummarize` fire-and-forget.
`triggerSummary(session)` calls `maybeSummarize` fire-and-forget. It takes only
the `session` — it no longer receives the full episode list. Instead
`maybeSummarize` fetches exactly what it needs:
1. `GET /sessions/:id/episode-stats` — a cheap aggregate (`COUNT`, `SUM(token_count)`,
`MAX(id)`) that gates the token threshold **without** pulling every episode row
2. Only if over threshold: `GET /sessions/:id/episodes/since/:afterId` — the
un-summarized tail (episodes newer than the last summary's range), fetched in
full for the summary prompt
This replaced an earlier approach where `chat/index.js` fetched the entire
session (`getRecentEpisodes(session.id, 9999)`) on every message just to hand it
over — an O(session length) cost per turn. The stats query now runs on every
message; full episode text is fetched only when a summary actually fires.
`maybeSummarize` proceeds only when both conditions are met:
1. Total session token count exceeds `SUMMARIES.THRESHOLD_TOKENS` (default 200)
+287
View File
@@ -0,0 +1,287 @@
diff -ruN nexusai-baseline/docs/reference/API-routes.md nexusai/docs/reference/API-routes.md
--- nexusai-baseline/docs/reference/API-routes.md 2026-08-17 11:41:32.878153043 +0000
+++ nexusai/docs/reference/API-routes.md 2026-08-17 12:28:03.473533125 +0000
@@ -277,6 +277,8 @@
| GET | /episodes/search?q=&limit= | FTS keyword search across all episodes |
| GET | /episodes/:id | Get episode by ID |
| GET | /sessions/:id/episodes?limit=&offset= | Paginated episodes for a session |
+| GET | /sessions/:id/episode-stats | Aggregate: count, total tokens, max id (summarization threshold check) |
+| GET | /sessions/:id/episodes/since/:afterId | Episodes newer than :afterId, chronological (un-summarized tail) |
| DELETE | /episodes/:id | Delete episode (SQLite + Qdrant cleanup) |
> Route ordering: `/episodes/search` must be defined before `/episodes/:id`.
diff -ruN nexusai-baseline/docs/reference/testing.md nexusai/docs/reference/testing.md
--- nexusai-baseline/docs/reference/testing.md 1970-01-01 00:00:00.000000000 +0000
+++ nexusai/docs/reference/testing.md 2026-08-17 12:28:47.370890122 +0000
@@ -0,0 +1,46 @@
+# Testing
+
+NexusAI has a lightweight regression suite built on Node's **built-in** test
+runner (`node:test`) and assertions (`node:assert`) — no external test
+dependencies. It runs offline on any node in the homelab.
+
+```bash
+npm test # = node --test (discovers test/*.test.js at the repo root)
+```
+
+## What's covered
+
+| File | Subject | Imports the real… |
+|---|---|---|
+| `test/fusion.test.js` | RRF ranking math | `fuseEpisodeResults` (orchestration) |
+| `test/fts-query.test.js` | Keyword tokenizer + FTS5 matching | `buildFtsQuery` (memory-service) |
+| `test/llamacpp-stream.test.js` | SSE stream reassembly across chunks | `completeStream` (inference) |
+| `test/summarization.test.js` | Summarization decision logic | `maybeSummarize` (orchestration) |
+| `test/entity-extraction.test.js` | Greeting + regurgitation guards | `mentionedIn`, `isIgnoredName` (memory-service) |
+| `test/schema.test.js` | Fresh-DB schema completeness | `schema.js` string (memory-service) |
+| `test/migrations.test.js` | Migration version-stepping | `migrate` (memory-service) |
+
+Tests import the **real** functions rather than reimplementing logic — the
+functions are exported for this purpose. External calls (Ollama, Qdrant, the
+memory service) are mocked via `global.fetch`; SQLite-backed tests use the
+built-in `node:sqlite` module against a throwaway in-memory database.
+
+## `node:sqlite` and skips
+
+`test/schema.test.js` and three cases in `test/fts-query.test.js` need
+`node:sqlite`, which requires **Node ≥ 22.5**. On older Node they `skip`
+themselves cleanly (guarded by `{ skip: !DatabaseSync }`) rather than failing,
+so the suite stays green everywhere — but those checks only *verify* anything on
+a node new enough to run them. A dev machine on current Node is the source of
+truth for the schema and FTS matching tests.
+
+`node:sqlite` is still marked experimental, so runs print a one-line
+`ExperimentalWarning`. It's harmless; `node --test --no-warnings` suppresses it.
+
+## Adding tests
+
+Keep the pattern: export the real function, import it, mock I/O at the boundary
+(`global.fetch`) or use `node:sqlite` for DB behaviour. Prefer testing pure
+logic (ranking, tokenizing, version-stepping, decision branches) over wiring.
+Several of these tests were written *after* a bug slipped through — each new
+class of mistake is worth a case so it can't recur silently.
diff -ruN nexusai-baseline/docs/roadmap.md nexusai/docs/roadmap.md
--- nexusai-baseline/docs/roadmap.md 2026-08-17 11:41:32.878198690 +0000
+++ nexusai/docs/roadmap.md 2026-08-17 12:28:30.788572199 +0000
@@ -68,7 +68,9 @@
Multi-strategy retrieval merged into a single ranked result set.
- [x] Reciprocal Rank Fusion (RRF) — merge semantic (Qdrant) + keyword (FTS5) results
- [x] Configurable weights per retrieval strategy (`semanticWeight`, `keywordWeight` via `PATCH /settings`)
-- [x] Score threshold retained per-strategy; FTS scoped to session/project sessions; `keywordWeight: 0` default (disabled until tuned)
+- [x] Score threshold retained per-strategy; FTS scoped to session/project sessions
+- [x] FTS query tokenization (`buildFtsQuery`) — tokenize + stopword-filter + quoted `OR` terms; replaced whole-phrase matching that made keyword recall nil
+- [x] Fusion tuned and verified live (keyword path + project scoping confirmed via weight inversion); running `keywordWeight: 0.5` / `semanticWeight: 1.0`
### 3. Memory Consolidation Lifecycle
Prevents long-term memory degradation and enables compression.
diff -ruN nexusai-baseline/docs/services/entity-extraction.md nexusai/docs/services/entity-extraction.md
--- nexusai-baseline/docs/services/entity-extraction.md 2026-08-17 11:41:32.878265351 +0000
+++ nexusai/docs/services/entity-extraction.md 2026-08-17 12:27:44.172932395 +0000
@@ -131,6 +131,20 @@
> imperfectly — switch to `/[^\p{L}\p{N}\s]/gu` if the set gains non-ASCII
> entries.
+**Regurgitation guard (`mentionedIn`):** the extraction prompt feeds the model
+a "known entities" hint block (the 20 most-recent entities) for spelling/type
+consistency. The small model (qwen2.5:3b) will sometimes echo that list back as
+if those entities appeared in the conversation — most visibly on contentless
+turns (a greeting produced fake extractions of unrelated authors, game titles,
+etc.). After parsing, each extracted name is checked against the actual
+`userMessage + aiResponse` text (case- and whitespace-normalized substring); any
+name not present is dropped before upsert. Since the prompt constrains names to
+short proper nouns, a genuinely-discussed entity appears verbatim while a
+regurgitated hint does not. Relationships referencing a dropped entity fall away
+automatically (they resolve against the surviving `entityMap`). A prompt line
+also tells the model the hint list is spelling-only — a backstop, with
+`mentionedIn` as the deterministic guarantee.
+
## Relationship Processing
After all entities are saved, relationships are processed:
diff -ruN nexusai-baseline/docs/services/memory-service.md nexusai/docs/services/memory-service.md
--- nexusai-baseline/docs/services/memory-service.md 2026-08-17 11:41:32.878409053 +0000
+++ nexusai/docs/services/memory-service.md 2026-08-17 12:27:35.716782770 +0000
@@ -36,8 +36,9 @@
```
src/
├── db/
-│ ├── index.js # SQLite connection + initialization + migrations
-│ ├── schema.js # Table definitions, indexes, FTS5, triggers
+│ ├── index.js # SQLite connection + init + migrate() + one-time FTS backfill
+│ ├── migrations.js # Forward-only versioned migration runner (PRAGMA user_version)
+│ ├── schema.js # Complete current shape: tables, indexes, FTS5, triggers
│ ├── projects.js # Project CRUD functions
│ └── summaries.js # Summary CRUD functions
├── episodic/
@@ -64,36 +65,56 @@
- **summaries** — condensed episode groups for efficient context retrieval
- **projects** — named groupings of sessions with `name`, `description`, `colour`, `icon`, `isolated`, `notes`, `system_prompt`
-### Migrations
+### Schema & Migrations
-Schema changes that cannot use `CREATE TABLE IF NOT EXISTS` are applied as
-idempotent migrations in `db/index.js` at startup:
+`schema.js` holds the **complete current shape** — every table, column, index,
+the FTS5 virtual table, and its triggers — as the single source of truth for a
+fresh database. It uses `CREATE TABLE IF NOT EXISTS`, so on a fresh DB it builds
+everything; on an existing DB it skips tables that already exist (and therefore
+does **not** reconcile columns on old tables — that's what migrations are for).
+
+`db/migrations.js` is a forward-only versioned runner keyed on
+`PRAGMA user_version`:
```js
-try { db.exec(`ALTER TABLE sessions ADD COLUMN name TEXT`); } catch {}
-try { db.exec(`ALTER TABLE sessions ADD COLUMN project_id INTEGER REFERENCES projects(id)`); } catch {}
-try { db.exec(`CREATE INDEX IF NOT EXISTS idx_sessions_project ON sessions(project_id)`); } catch {}
-try { db.exec(`ALTER TABLE projects ADD COLUMN isolated INTEGER NOT NULL DEFAULT 0`); } catch {}
-try { db.exec(`ALTER TABLE projects ADD COLUMN notes TEXT`); } catch {}
-try { db.exec(`ALTER TABLE projects ADD COLUMN system_prompt TEXT`); } catch {}
-// Knowledge graph columns:
-try { db.exec(`ALTER TABLE entities ADD COLUMN mention_count INTEGER NOT NULL DEFAULT 1`) } catch {}
-try { db.exec(`ALTER TABLE entities ADD COLUMN confidence REAL NOT NULL DEFAULT 1.0`) } catch {}
-try { db.exec(`ALTER TABLE entities ADD COLUMN source TEXT NOT NULL DEFAULT 'extraction'`) } catch {}
-try { db.exec(`ALTER TABLE entities ADD COLUMN last_seen_at INTEGER`) } catch {}
-try { db.exec(`ALTER TABLE relationships ADD COLUMN mention_count INTEGER NOT NULL DEFAULT 1`) } catch {}
-try { db.exec(`ALTER TABLE relationships ADD COLUMN notes TEXT`) } catch {}
+const migrations = [
+ (_db) => {}, // v0 → v1: consolidated baseline (historical ALTERs folded into schema.js)
+];
+const LATEST_VERSION = migrations.length; // derived, never hand-maintained
```
-`entity_episodes` is defined in `schema.js` itself (not a migration) since it is a new table.
-
-New migrations are always appended — never modify the schema file for existing tables since `ALTER TABLE` cannot use `IF NOT EXISTS`.
+`migrate(db)` reads `user_version`, applies every entry newer than it (each in a
+transaction alongside its version bump), and stamps the result. A fresh DB is
+built whole by `schema.js` and simply stamped to `LATEST_VERSION`; the baseline
+entry is a no-op.
+
+**Adding a schema change:** append a new function to the `migrations` array
+(which bumps `LATEST_VERSION` automatically). Never edit an already-shipped
+entry, and never edit a table in `schema.js` expecting existing DBs to pick it
+up — they won't. This replaces the previous pattern of stacking silent
+`try/catch ALTER TABLE` statements in `db/index.js` on every boot.
+
+> **Consolidation note:** the historical ALTERs were folded into `schema.js`
+> rather than preserved as replayable migrations, so this assumes a fresh
+> database (which is the case post-wipe). An older, pre-consolidation database
+> would **not** auto-upgrade — `schema.js` skips its existing tables and the
+> baseline migration is a no-op. To support upgrading old DBs, the v1 baseline
+> would instead perform guarded (`ADD COLUMN if missing`) catch-up.
### FTS5 Full-Text Search
-An `episodes_fts` virtual table enables keyword search across all episodes.
-Three triggers (`episodes_fts_insert`, `episodes_fts_update`, `episodes_fts_delete`)
-keep the FTS index automatically in sync with the episodes table.
+An `episodes_fts` external-content virtual table enables keyword search across
+episodes. Three triggers (`episodes_fts_insert`, `episodes_fts_update`,
+`episodes_fts_delete`) keep the index in sync with the `episodes` table
+automatically during normal operation.
+
+A one-time backfill in `db/index.js` handles the case where the FTS table is
+created on a DB that already holds episodes (e.g. episodes predating FTS). It is
+gated on "did `episodes_fts` not exist before this boot," checked via
+`sqlite_master` **before** running the schema — not on a row-count comparison,
+because `COUNT(*)` on an external-content FTS5 table proxies the content table
+and cannot detect a desync. This replaced an unconditional full FTS rebuild that
+previously ran on every startup.
### SQLite Configuration
diff -ruN nexusai-baseline/docs/services/retrieval-fusion.md nexusai/docs/services/retrieval-fusion.md
--- nexusai-baseline/docs/services/retrieval-fusion.md 2026-08-17 11:41:32.878594989 +0000
+++ nexusai/docs/services/retrieval-fusion.md 2026-08-17 12:27:07.267654899 +0000
@@ -37,9 +37,10 @@
sources, and fusion is a retrieval strategy, not a storage concern.
```
-getFusedEpisodes()
-├── getSemanticEpisodes() — Qdrant embed+search → fetch full rows by ID
-│ (existing path, unchanged)
+getFusedEpisodes(…, queryVector, …)
+├── getSemanticEpisodes(queryVector) — Qdrant search → fetch full rows by ID
+│ (query embedded ONCE upstream in assembleContext and shared with entity
+│ search — no longer embedded separately here)
└── getFTSResults() — memory-service /episodes/search → full rows directly
(skipped entirely if keywordWeight == 0)
↓
@@ -48,6 +49,10 @@
fusedEpisodes[] — top semanticLimit episodes by RRF score
```
+The query embedding is computed once per turn in `assembleContext` and passed
+into both fused retrieval and entity search; if embedding fails, both receive
+`null` and degrade to empty results rather than erroring.
+
### Data Shape Consistency
Both sides must enter fusion as `Episode[]` — full SQLite row objects with
@@ -59,6 +64,38 @@
FTS requests `semanticLimit * 2` results to provide headroom for the
`recentIds` filter without under-serving the fusion.
+## Query Tokenization
+
+Before FTS5 sees the query, `buildFtsQuery(query)` (in
+`memory-service/src/episodic/index.js`) turns the raw message into a MATCH
+expression:
+
+1. Lowercase and split on any non-letter/number (`/[^\p{L}\p{N}]+/u`, unicode-aware)
+2. Drop stopwords (a small `FTS_STOPWORDS` set of common function words) and single-character tokens
+3. Wrap each surviving token in double quotes and join with ` OR `
+
+So `"How do I configure the Qdrant collection?"` becomes
+`"configure" OR "qdrant" OR "collection"`. If nothing survives (an
+all-stopword message like `"how do I do it?"`), it returns `null` and
+`searchEpisodes` returns `[]` — keyword search sits out that turn and
+semantic retrieval carries it.
+
+**Why this matters:** the earlier implementation quoted the *entire* message
+as one FTS5 phrase, which required the whole string to appear verbatim in an
+episode — so keyword recall was effectively nil for conversational queries.
+Tokenizing into OR-joined terms is what makes `keywordWeight > 0` actually
+contribute anything.
+
+**Injection safety:** quoting each token individually means any
+FTS5-significant token inside the user's message (a literal `OR`, `*`, `"`,
+etc.) is matched as a search term rather than parsed as an operator. This
+replaces the safety the old whole-phrase quoting provided.
+
+The stopword set is deliberately conservative and tuned iteratively — high
+frequency filler (`the`, `is`, `one`, `there`, …) is dropped, but borderline
+words that can carry signal (`time`, `good`, `way`) are kept. Add to the set
+when a common word is observed producing noisy matches.
+
## FTS Session Scoping
Without scoping, FTS5 searches across all episodes in the database. For
diff -ruN nexusai-baseline/docs/services/summarization.md nexusai/docs/services/summarization.md
--- nexusai-baseline/docs/services/summarization.md 2026-08-17 11:41:32.878899850 +0000
+++ nexusai/docs/services/summarization.md 2026-08-17 12:27:55.087746482 +0000
@@ -12,7 +12,21 @@
## Trigger Conditions
-`triggerSummary(session, allEpisodes)` calls `maybeSummarize` fire-and-forget.
+`triggerSummary(session)` calls `maybeSummarize` fire-and-forget. It takes only
+the `session` — it no longer receives the full episode list. Instead
+`maybeSummarize` fetches exactly what it needs:
+
+1. `GET /sessions/:id/episode-stats` — a cheap aggregate (`COUNT`, `SUM(token_count)`,
+ `MAX(id)`) that gates the token threshold **without** pulling every episode row
+2. Only if over threshold: `GET /sessions/:id/episodes/since/:afterId` — the
+ un-summarized tail (episodes newer than the last summary's range), fetched in
+ full for the summary prompt
+
+This replaced an earlier approach where `chat/index.js` fetched the entire
+session (`getRecentEpisodes(session.id, 9999)`) on every message just to hand it
+over — an O(session length) cost per turn. The stats query now runs on every
+message; full episode text is fetched only when a summary actually fires.
+
`maybeSummarize` proceeds only when both conditions are met:
1. Total session token count exceeds `SUMMARIES.THRESHOLD_TOKENS` (default 200)