diff --git a/nexusai-docs-update-2.patch b/nexusai-docs-update-2.patch deleted file mode 100644 index c474fa8..0000000 --- a/nexusai-docs-update-2.patch +++ /dev/null @@ -1,287 +0,0 @@ -diff -ruN nexusai-baseline/docs/reference/API-routes.md nexusai/docs/reference/API-routes.md ---- nexusai-baseline/docs/reference/API-routes.md 2026-08-17 11:41:32.878153043 +0000 -+++ nexusai/docs/reference/API-routes.md 2026-08-17 12:28:03.473533125 +0000 -@@ -277,6 +277,8 @@ - | GET | /episodes/search?q=&limit= | FTS keyword search across all episodes | - | GET | /episodes/:id | Get episode by ID | - | GET | /sessions/:id/episodes?limit=&offset= | Paginated episodes for a session | -+| GET | /sessions/:id/episode-stats | Aggregate: count, total tokens, max id (summarization threshold check) | -+| GET | /sessions/:id/episodes/since/:afterId | Episodes newer than :afterId, chronological (un-summarized tail) | - | DELETE | /episodes/:id | Delete episode (SQLite + Qdrant cleanup) | - - > Route ordering: `/episodes/search` must be defined before `/episodes/:id`. -diff -ruN nexusai-baseline/docs/reference/testing.md nexusai/docs/reference/testing.md ---- nexusai-baseline/docs/reference/testing.md 1970-01-01 00:00:00.000000000 +0000 -+++ nexusai/docs/reference/testing.md 2026-08-17 12:28:47.370890122 +0000 -@@ -0,0 +1,46 @@ -+# Testing -+ -+NexusAI has a lightweight regression suite built on Node's **built-in** test -+runner (`node:test`) and assertions (`node:assert`) — no external test -+dependencies. It runs offline on any node in the homelab. -+ -+```bash -+npm test # = node --test (discovers test/*.test.js at the repo root) -+``` -+ -+## What's covered -+ -+| File | Subject | Imports the real… | -+|---|---|---| -+| `test/fusion.test.js` | RRF ranking math | `fuseEpisodeResults` (orchestration) | -+| `test/fts-query.test.js` | Keyword tokenizer + FTS5 matching | `buildFtsQuery` (memory-service) | -+| `test/llamacpp-stream.test.js` | SSE stream reassembly across chunks | `completeStream` (inference) | -+| `test/summarization.test.js` | Summarization decision logic | `maybeSummarize` (orchestration) | -+| `test/entity-extraction.test.js` | Greeting + regurgitation guards | `mentionedIn`, `isIgnoredName` (memory-service) | -+| `test/schema.test.js` | Fresh-DB schema completeness | `schema.js` string (memory-service) | -+| `test/migrations.test.js` | Migration version-stepping | `migrate` (memory-service) | -+ -+Tests import the **real** functions rather than reimplementing logic — the -+functions are exported for this purpose. External calls (Ollama, Qdrant, the -+memory service) are mocked via `global.fetch`; SQLite-backed tests use the -+built-in `node:sqlite` module against a throwaway in-memory database. -+ -+## `node:sqlite` and skips -+ -+`test/schema.test.js` and three cases in `test/fts-query.test.js` need -+`node:sqlite`, which requires **Node ≥ 22.5**. On older Node they `skip` -+themselves cleanly (guarded by `{ skip: !DatabaseSync }`) rather than failing, -+so the suite stays green everywhere — but those checks only *verify* anything on -+a node new enough to run them. A dev machine on current Node is the source of -+truth for the schema and FTS matching tests. -+ -+`node:sqlite` is still marked experimental, so runs print a one-line -+`ExperimentalWarning`. It's harmless; `node --test --no-warnings` suppresses it. -+ -+## Adding tests -+ -+Keep the pattern: export the real function, import it, mock I/O at the boundary -+(`global.fetch`) or use `node:sqlite` for DB behaviour. Prefer testing pure -+logic (ranking, tokenizing, version-stepping, decision branches) over wiring. -+Several of these tests were written *after* a bug slipped through — each new -+class of mistake is worth a case so it can't recur silently. -diff -ruN nexusai-baseline/docs/roadmap.md nexusai/docs/roadmap.md ---- nexusai-baseline/docs/roadmap.md 2026-08-17 11:41:32.878198690 +0000 -+++ nexusai/docs/roadmap.md 2026-08-17 12:28:30.788572199 +0000 -@@ -68,7 +68,9 @@ - Multi-strategy retrieval merged into a single ranked result set. - - [x] Reciprocal Rank Fusion (RRF) — merge semantic (Qdrant) + keyword (FTS5) results - - [x] Configurable weights per retrieval strategy (`semanticWeight`, `keywordWeight` via `PATCH /settings`) --- [x] Score threshold retained per-strategy; FTS scoped to session/project sessions; `keywordWeight: 0` default (disabled until tuned) -+- [x] Score threshold retained per-strategy; FTS scoped to session/project sessions -+- [x] FTS query tokenization (`buildFtsQuery`) — tokenize + stopword-filter + quoted `OR` terms; replaced whole-phrase matching that made keyword recall nil -+- [x] Fusion tuned and verified live (keyword path + project scoping confirmed via weight inversion); running `keywordWeight: 0.5` / `semanticWeight: 1.0` - - ### 3. Memory Consolidation Lifecycle - Prevents long-term memory degradation and enables compression. -diff -ruN nexusai-baseline/docs/services/entity-extraction.md nexusai/docs/services/entity-extraction.md ---- nexusai-baseline/docs/services/entity-extraction.md 2026-08-17 11:41:32.878265351 +0000 -+++ nexusai/docs/services/entity-extraction.md 2026-08-17 12:27:44.172932395 +0000 -@@ -131,6 +131,20 @@ - > imperfectly — switch to `/[^\p{L}\p{N}\s]/gu` if the set gains non-ASCII - > entries. - -+**Regurgitation guard (`mentionedIn`):** the extraction prompt feeds the model -+a "known entities" hint block (the 20 most-recent entities) for spelling/type -+consistency. The small model (qwen2.5:3b) will sometimes echo that list back as -+if those entities appeared in the conversation — most visibly on contentless -+turns (a greeting produced fake extractions of unrelated authors, game titles, -+etc.). After parsing, each extracted name is checked against the actual -+`userMessage + aiResponse` text (case- and whitespace-normalized substring); any -+name not present is dropped before upsert. Since the prompt constrains names to -+short proper nouns, a genuinely-discussed entity appears verbatim while a -+regurgitated hint does not. Relationships referencing a dropped entity fall away -+automatically (they resolve against the surviving `entityMap`). A prompt line -+also tells the model the hint list is spelling-only — a backstop, with -+`mentionedIn` as the deterministic guarantee. -+ - ## Relationship Processing - - After all entities are saved, relationships are processed: -diff -ruN nexusai-baseline/docs/services/memory-service.md nexusai/docs/services/memory-service.md ---- nexusai-baseline/docs/services/memory-service.md 2026-08-17 11:41:32.878409053 +0000 -+++ nexusai/docs/services/memory-service.md 2026-08-17 12:27:35.716782770 +0000 -@@ -36,8 +36,9 @@ - ``` - src/ - ├── db/ --│ ├── index.js # SQLite connection + initialization + migrations --│ ├── schema.js # Table definitions, indexes, FTS5, triggers -+│ ├── index.js # SQLite connection + init + migrate() + one-time FTS backfill -+│ ├── migrations.js # Forward-only versioned migration runner (PRAGMA user_version) -+│ ├── schema.js # Complete current shape: tables, indexes, FTS5, triggers - │ ├── projects.js # Project CRUD functions - │ └── summaries.js # Summary CRUD functions - ├── episodic/ -@@ -64,36 +65,56 @@ - - **summaries** — condensed episode groups for efficient context retrieval - - **projects** — named groupings of sessions with `name`, `description`, `colour`, `icon`, `isolated`, `notes`, `system_prompt` - --### Migrations -+### Schema & Migrations - --Schema changes that cannot use `CREATE TABLE IF NOT EXISTS` are applied as --idempotent migrations in `db/index.js` at startup: -+`schema.js` holds the **complete current shape** — every table, column, index, -+the FTS5 virtual table, and its triggers — as the single source of truth for a -+fresh database. It uses `CREATE TABLE IF NOT EXISTS`, so on a fresh DB it builds -+everything; on an existing DB it skips tables that already exist (and therefore -+does **not** reconcile columns on old tables — that's what migrations are for). -+ -+`db/migrations.js` is a forward-only versioned runner keyed on -+`PRAGMA user_version`: - - ```js --try { db.exec(`ALTER TABLE sessions ADD COLUMN name TEXT`); } catch {} --try { db.exec(`ALTER TABLE sessions ADD COLUMN project_id INTEGER REFERENCES projects(id)`); } catch {} --try { db.exec(`CREATE INDEX IF NOT EXISTS idx_sessions_project ON sessions(project_id)`); } catch {} --try { db.exec(`ALTER TABLE projects ADD COLUMN isolated INTEGER NOT NULL DEFAULT 0`); } catch {} --try { db.exec(`ALTER TABLE projects ADD COLUMN notes TEXT`); } catch {} --try { db.exec(`ALTER TABLE projects ADD COLUMN system_prompt TEXT`); } catch {} --// Knowledge graph columns: --try { db.exec(`ALTER TABLE entities ADD COLUMN mention_count INTEGER NOT NULL DEFAULT 1`) } catch {} --try { db.exec(`ALTER TABLE entities ADD COLUMN confidence REAL NOT NULL DEFAULT 1.0`) } catch {} --try { db.exec(`ALTER TABLE entities ADD COLUMN source TEXT NOT NULL DEFAULT 'extraction'`) } catch {} --try { db.exec(`ALTER TABLE entities ADD COLUMN last_seen_at INTEGER`) } catch {} --try { db.exec(`ALTER TABLE relationships ADD COLUMN mention_count INTEGER NOT NULL DEFAULT 1`) } catch {} --try { db.exec(`ALTER TABLE relationships ADD COLUMN notes TEXT`) } catch {} -+const migrations = [ -+ (_db) => {}, // v0 → v1: consolidated baseline (historical ALTERs folded into schema.js) -+]; -+const LATEST_VERSION = migrations.length; // derived, never hand-maintained - ``` - --`entity_episodes` is defined in `schema.js` itself (not a migration) since it is a new table. -- --New migrations are always appended — never modify the schema file for existing tables since `ALTER TABLE` cannot use `IF NOT EXISTS`. -+`migrate(db)` reads `user_version`, applies every entry newer than it (each in a -+transaction alongside its version bump), and stamps the result. A fresh DB is -+built whole by `schema.js` and simply stamped to `LATEST_VERSION`; the baseline -+entry is a no-op. -+ -+**Adding a schema change:** append a new function to the `migrations` array -+(which bumps `LATEST_VERSION` automatically). Never edit an already-shipped -+entry, and never edit a table in `schema.js` expecting existing DBs to pick it -+up — they won't. This replaces the previous pattern of stacking silent -+`try/catch ALTER TABLE` statements in `db/index.js` on every boot. -+ -+> **Consolidation note:** the historical ALTERs were folded into `schema.js` -+> rather than preserved as replayable migrations, so this assumes a fresh -+> database (which is the case post-wipe). An older, pre-consolidation database -+> would **not** auto-upgrade — `schema.js` skips its existing tables and the -+> baseline migration is a no-op. To support upgrading old DBs, the v1 baseline -+> would instead perform guarded (`ADD COLUMN if missing`) catch-up. - - ### FTS5 Full-Text Search - --An `episodes_fts` virtual table enables keyword search across all episodes. --Three triggers (`episodes_fts_insert`, `episodes_fts_update`, `episodes_fts_delete`) --keep the FTS index automatically in sync with the episodes table. -+An `episodes_fts` external-content virtual table enables keyword search across -+episodes. Three triggers (`episodes_fts_insert`, `episodes_fts_update`, -+`episodes_fts_delete`) keep the index in sync with the `episodes` table -+automatically during normal operation. -+ -+A one-time backfill in `db/index.js` handles the case where the FTS table is -+created on a DB that already holds episodes (e.g. episodes predating FTS). It is -+gated on "did `episodes_fts` not exist before this boot," checked via -+`sqlite_master` **before** running the schema — not on a row-count comparison, -+because `COUNT(*)` on an external-content FTS5 table proxies the content table -+and cannot detect a desync. This replaced an unconditional full FTS rebuild that -+previously ran on every startup. - - ### SQLite Configuration - -diff -ruN nexusai-baseline/docs/services/retrieval-fusion.md nexusai/docs/services/retrieval-fusion.md ---- nexusai-baseline/docs/services/retrieval-fusion.md 2026-08-17 11:41:32.878594989 +0000 -+++ nexusai/docs/services/retrieval-fusion.md 2026-08-17 12:27:07.267654899 +0000 -@@ -37,9 +37,10 @@ - sources, and fusion is a retrieval strategy, not a storage concern. - - ``` --getFusedEpisodes() --├── getSemanticEpisodes() — Qdrant embed+search → fetch full rows by ID --│ (existing path, unchanged) -+getFusedEpisodes(…, queryVector, …) -+├── getSemanticEpisodes(queryVector) — Qdrant search → fetch full rows by ID -+│ (query embedded ONCE upstream in assembleContext and shared with entity -+│ search — no longer embedded separately here) - └── getFTSResults() — memory-service /episodes/search → full rows directly - (skipped entirely if keywordWeight == 0) - ↓ -@@ -48,6 +49,10 @@ - fusedEpisodes[] — top semanticLimit episodes by RRF score - ``` - -+The query embedding is computed once per turn in `assembleContext` and passed -+into both fused retrieval and entity search; if embedding fails, both receive -+`null` and degrade to empty results rather than erroring. -+ - ### Data Shape Consistency - - Both sides must enter fusion as `Episode[]` — full SQLite row objects with -@@ -59,6 +64,38 @@ - FTS requests `semanticLimit * 2` results to provide headroom for the - `recentIds` filter without under-serving the fusion. - -+## Query Tokenization -+ -+Before FTS5 sees the query, `buildFtsQuery(query)` (in -+`memory-service/src/episodic/index.js`) turns the raw message into a MATCH -+expression: -+ -+1. Lowercase and split on any non-letter/number (`/[^\p{L}\p{N}]+/u`, unicode-aware) -+2. Drop stopwords (a small `FTS_STOPWORDS` set of common function words) and single-character tokens -+3. Wrap each surviving token in double quotes and join with ` OR ` -+ -+So `"How do I configure the Qdrant collection?"` becomes -+`"configure" OR "qdrant" OR "collection"`. If nothing survives (an -+all-stopword message like `"how do I do it?"`), it returns `null` and -+`searchEpisodes` returns `[]` — keyword search sits out that turn and -+semantic retrieval carries it. -+ -+**Why this matters:** the earlier implementation quoted the *entire* message -+as one FTS5 phrase, which required the whole string to appear verbatim in an -+episode — so keyword recall was effectively nil for conversational queries. -+Tokenizing into OR-joined terms is what makes `keywordWeight > 0` actually -+contribute anything. -+ -+**Injection safety:** quoting each token individually means any -+FTS5-significant token inside the user's message (a literal `OR`, `*`, `"`, -+etc.) is matched as a search term rather than parsed as an operator. This -+replaces the safety the old whole-phrase quoting provided. -+ -+The stopword set is deliberately conservative and tuned iteratively — high -+frequency filler (`the`, `is`, `one`, `there`, …) is dropped, but borderline -+words that can carry signal (`time`, `good`, `way`) are kept. Add to the set -+when a common word is observed producing noisy matches. -+ - ## FTS Session Scoping - - Without scoping, FTS5 searches across all episodes in the database. For -diff -ruN nexusai-baseline/docs/services/summarization.md nexusai/docs/services/summarization.md ---- nexusai-baseline/docs/services/summarization.md 2026-08-17 11:41:32.878899850 +0000 -+++ nexusai/docs/services/summarization.md 2026-08-17 12:27:55.087746482 +0000 -@@ -12,7 +12,21 @@ - - ## Trigger Conditions - --`triggerSummary(session, allEpisodes)` calls `maybeSummarize` fire-and-forget. -+`triggerSummary(session)` calls `maybeSummarize` fire-and-forget. It takes only -+the `session` — it no longer receives the full episode list. Instead -+`maybeSummarize` fetches exactly what it needs: -+ -+1. `GET /sessions/:id/episode-stats` — a cheap aggregate (`COUNT`, `SUM(token_count)`, -+ `MAX(id)`) that gates the token threshold **without** pulling every episode row -+2. Only if over threshold: `GET /sessions/:id/episodes/since/:afterId` — the -+ un-summarized tail (episodes newer than the last summary's range), fetched in -+ full for the summary prompt -+ -+This replaced an earlier approach where `chat/index.js` fetched the entire -+session (`getRecentEpisodes(session.id, 9999)`) on every message just to hand it -+over — an O(session length) cost per turn. The stats query now runs on every -+message; full episode text is fetched only when a summary actually fires. -+ - `maybeSummarize` proceeds only when both conditions are met: - - 1. Total session token count exceeds `SUMMARIES.THRESHOLD_TOKENS` (default 200)