Compare commits

...
2 Commits
Author SHA1 Message Date
Storme-bit 3d7372d04f documentation update 2026-08-23 23:27:53 -07:00
Storme-bit c31fb786f4 documentation update 2026-08-23 23:26:51 -07:00
7 changed files with 111 additions and 58 deletions
+21
View File
@@ -205,6 +205,9 @@ Returns `503` if llama-server is unreachable.
| `scoreThreshold` | float | 0–1 | Minimum similarity score for Qdrant results | | `scoreThreshold` | float | 0–1 | Minimum similarity score for Qdrant results |
| `semanticWeight` | float | 0–5 | RRF weight for Qdrant semantic results | | `semanticWeight` | float | 0–5 | RRF weight for Qdrant semantic results |
| `keywordWeight` | float | 0–5 | RRF weight for FTS5 keyword results (`0` = disabled) | | `keywordWeight` | float | 0–5 | RRF weight for FTS5 keyword results (`0` = disabled) |
| `contextBudget` | integer | — | Token budget for context assembly (char/4 estimation) |
| `entityWeight` | float | — | Scoring bonus for entity-linked episodes in the context pool |
| `minRecentEpisodes` | integer | — | Guaranteed floor of recent episodes always included in context |
| `modelsFolderPath` | string | — | Path to folder containing .gguf files | | `modelsFolderPath` | string | — | Path to folder containing .gguf files |
| `temperature` | float | 0–2 | Inference randomness | | `temperature` | float | 0–2 | Inference randomness |
| `repeatPenalty` | float | 1–2 | Repeat token penalty | | `repeatPenalty` | float | 1–2 | Repeat token penalty |
@@ -253,6 +256,7 @@ orchestration.
| GET | /sessions/by-external/:externalId | Get session by external ID | | GET | /sessions/by-external/:externalId | Get session by external ID |
| PATCH | /sessions/by-external/:externalId | Update session fields | | PATCH | /sessions/by-external/:externalId | Update session fields |
| DELETE | /sessions/by-external/:externalId | Delete session (cascades to episodes) | | DELETE | /sessions/by-external/:externalId | Delete session (cascades to episodes) |
| GET | /sessions/:id/entity-ids | Entity IDs linked to this session's episodes (non-project memory isolation scoping) |
> Route ordering: `by-external/:externalId` must be defined before `/:id` > Route ordering: `by-external/:externalId` must be defined before `/:id`
> to prevent `by-external` being captured as an ID param. > to prevent `by-external` being captured as an ID param.
@@ -279,6 +283,8 @@ Both fields are optional. Only provided fields are updated.
| GET | /sessions/:id/episodes?limit=&offset= | Paginated episodes for a session | | GET | /sessions/:id/episodes?limit=&offset= | Paginated episodes for a session |
| GET | /sessions/:id/episode-stats | Aggregate: count, total tokens, max id (summarization threshold check) | | GET | /sessions/:id/episode-stats | Aggregate: count, total tokens, max id (summarization threshold check) |
| GET | /sessions/:id/episodes/since/:afterId | Episodes newer than :afterId, chronological (un-summarized tail) | | GET | /sessions/:id/episodes/since/:afterId | Episodes newer than :afterId, chronological (un-summarized tail) |
| POST | /episodes/touch | Batch access-tracking bump — increments `access_count`, sets `last_accessed_at` |
| GET | /sessions/:id/consolidation-candidates | Dry-run consolidation scoring (aging score, floors applied) |
| DELETE | /episodes/:id | Delete episode (SQLite + Qdrant cleanup) | | DELETE | /episodes/:id | Delete episode (SQLite + Qdrant cleanup) |
> Route ordering: `/episodes/search` must be defined before `/episodes/:id`. > Route ordering: `/episodes/search` must be defined before `/episodes/:id`.
@@ -293,6 +299,21 @@ Both fields are optional. Only provided fields are updated.
} }
``` ```
**POST /episodes/touch — body:**
```json
{ "ids": [54, 55, 56] }
```
Returns `{ "touched": 3 }`. Called fire-and-forget by orchestration after
budget selection. Nonexistent IDs are silently skipped; `touched` echoes the
request count, not rows matched.
**GET /sessions/:id/consolidation-candidates** — returns
`{ eligible, reason?, candidates: [...] }` where each candidate is
`{ id, access_count, created_at, last_accessed_at, preview, aging_score }`,
sorted by `aging_score` ascending (most eligible first). Sessions under
`CONSOLIDATION.MIN_SESSION_EPISODES` return `eligible: false` with a `reason`.
Observe-only — nothing is modified.
### Projects ### Projects
| Method | Path | Description | | Method | Path | Description |
+8
View File
@@ -18,6 +18,7 @@ npm test # = node --test (discovers test/*.test.js at the repo root)
| `test/summarization.test.js` | Summarization decision logic | `maybeSummarize` (orchestration) | | `test/summarization.test.js` | Summarization decision logic | `maybeSummarize` (orchestration) |
| `test/entity-extraction.test.js` | Greeting + regurgitation guards | `mentionedIn`, `isIgnoredName` (memory-service) | | `test/entity-extraction.test.js` | Greeting + regurgitation guards | `mentionedIn`, `isIgnoredName` (memory-service) |
| `test/schema.test.js` | Fresh-DB schema completeness | `schema.js` string (memory-service) | | `test/schema.test.js` | Fresh-DB schema completeness | `schema.js` string (memory-service) |
| `test/trivial-turn.test.js` | Greeting/trivial-turn detection | `isTrivialTurn` (shared) |
| `test/migrations.test.js` | Migration version-stepping | `migrate` (memory-service) | | `test/migrations.test.js` | Migration version-stepping | `migrate` (memory-service) |
Tests import the **real** functions rather than reimplementing logic — the Tests import the **real** functions rather than reimplementing logic — the
@@ -44,3 +45,10 @@ Keep the pattern: export the real function, import it, mock I/O at the boundary
logic (ranking, tokenizing, version-stepping, decision branches) over wiring. logic (ranking, tokenizing, version-stepping, decision branches) over wiring.
Several of these tests were written *after* a bug slipped through — each new Several of these tests were written *after* a bug slipped through — each new
class of mistake is worth a case so it can't recur silently. class of mistake is worth a case so it can't recur silently.
When a mocked service call changes shape (e.g. the `utilityInference` refactor
moving summarization from Ollama's `/api/generate` to the inference service's
`/utility/complete`), the mock URL router must move with it — an "unexpected
fetch" throw in these tests usually means the code under test evolved, not
broke. Runner-contract tests (migrations) inject stub no-op migration arrays
rather than letting the real SQL-bearing migrations hit the minimal fake db.
+8 -7
View File
@@ -74,8 +74,9 @@ Multi-strategy retrieval merged into a single ranked result set.
### 3. Memory Consolidation Lifecycle ### 3. Memory Consolidation Lifecycle
Prevents long-term memory degradation and enables compression. Prevents long-term memory degradation and enables compression.
- [ ] Episode aging — score/weight episodes by recency and access frequency - [x] Episode aging — `access_count` + `last_accessed_at` columns (v2 migration), batch touch on retrieval selection (`POST /episodes/touch`), aging score `access_count / (1 + days since last access)`
- [ ] Consolidation pass — merge related low-weight episodes into summary nodes - [x] Dry-run candidates endpoint — `GET /sessions/:id/consolidation-candidates` with age + session-size floors (observe-only phase before destructive pass)
- [ ] Consolidation pass — merge related low-weight episodes into summary nodes (incl. Qdrant vector + entity link cleanup)
- [ ] Orphan cleanup — remove entities no longer referenced by active episodes - [ ] Orphan cleanup — remove entities no longer referenced by active episodes
### 4. User Preference Model ### 4. User Preference Model
@@ -90,11 +91,11 @@ Short-circuit simple requests before they reach the LLM.
- [ ] Confidence bands — FAST PATH (memory lookup only) vs FULL (LLM + context) - [ ] Confidence bands — FAST PATH (memory lookup only) vs FULL (LLM + context)
- [ ] Fast-path handlers — direct memory queries, session lookups, factual recalls - [ ] Fast-path handlers — direct memory queries, session lookups, factual recalls
### 6. Smarter Context Assembly *(inspired by acid2lake)* ### 6. Smarter Context Assembly *(inspired by acid2lake)* ✅
Budget-aware context selection instead of dumping all relevant memory into the prompt. Budget-aware context selection instead of dumping all relevant memory into the prompt.
- [ ] Token budget manager in orchestration - [x] Token budget manager in orchestration (`selectWithinBudget`, char/4 estimation on stored text)
- [ ] Priority scoring — recency × relevance × entity weight - [x] Priority scoring — RRF fusion + recency + entity boost (`buildScoredPool`)
- [ ] Configurable context budget via env var - [x] Configurable via settings (`contextBudget`, `entityWeight`, `minRecentEpisodes`) — live, no restart
### 7. Procedural Memory Store *(inspired by acid2lake)* ### 7. Procedural Memory Store *(inspired by acid2lake)*
Learns "how NexusAI has successfully handled this type of request before." Learns "how NexusAI has successfully handled this type of request before."
@@ -227,4 +228,4 @@ The JARVIS moment — NexusAI reasons, plans, and acts across multiple steps.
--- ---
*Last updated: April 2026* *Last updated: August 2026*
+33
View File
@@ -78,6 +78,11 @@ does **not** reconcile columns on old tables — that's what migrations are for)
```js ```js
const migrations = [ const migrations = [
(_db) => {}, // v0 → v1: consolidated baseline (historical ALTERs folded into schema.js) (_db) => {}, // v0 → v1: consolidated baseline (historical ALTERs folded into schema.js)
(db) => { // v1 → v2: access tracking for consolidation lifecycle
db.exec(`ALTER TABLE episodes ADD COLUMN last_accessed_at INTEGER`);
db.exec(`ALTER TABLE episodes ADD COLUMN access_count INTEGER NOT NULL DEFAULT 0`);
db.exec(`UPDATE episodes SET last_accessed_at = created_at`); // backfill
},
]; ];
const LATEST_VERSION = migrations.length; // derived, never hand-maintained const LATEST_VERSION = migrations.length; // derived, never hand-maintained
``` ```
@@ -121,6 +126,12 @@ previously ran on every startup.
- `foreign_keys = ON` — enforces referential integrity and cascade deletes - `foreign_keys = ON` — enforces referential integrity and cascade deletes
- PRAGMAs set via `db.pragma()`, not `db.exec()` - PRAGMAs set via `db.pragma()`, not `db.exec()`
> **Copying a live WAL database:** `cp` on the `.db` file alone silently loses
> everything in the un-checkpointed `-wal` file (recent writes, even the
> migration version stamp). Always use
> `sqlite3 nexusai.db "VACUUM INTO './copy.db'"` (or `.backup`) — safe while
> the service is running, produces a complete single-file snapshot.
### Dynamic Updates ### Dynamic Updates
Both `updateSession` and `updateProject` build their `SET` clause dynamically Both `updateSession` and `updateProject` build their `SET` clause dynamically
@@ -204,6 +215,28 @@ service is responsible only for CRUD — generation logic lives in orchestration
> For full details on trigger conditions, prompt format, cumulative updates, > For full details on trigger conditions, prompt format, cumulative updates,
> and ChatML token stripping, see `summarization.md`. > and ChatML token stripping, see `summarization.md`.
## Access Tracking & Consolidation (dry-run)
Every episode selected into a chat context window (budget-selected, not the
guaranteed-recency floor) gets an access bump via `POST /episodes/touch` —
`access_count` incremented, `last_accessed_at` set to `Date.now()` (ms).
Called fire-and-forget from orchestration; a failure loses one increment,
nothing more.
`GET /sessions/:id/consolidation-candidates` scores episodes by
`access_count / (1 + days since last access)` — never-accessed episodes fall
back to `created_at` for the recency term and score exactly 0 (most eligible).
Two floors apply: episodes younger than `CONSOLIDATION.MIN_AGE_DAYS` are
excluded in SQL; sessions under `CONSOLIDATION.MIN_SESSION_EPISODES` return
`eligible: false` before scoring runs. The endpoint is observe-only — the
destructive pass (merge → summarize → Qdrant cleanup → orphan sweep) is not
yet built.
> **Unit note:** `created_at` is unix **seconds** (`unixepoch()`);
> `last_accessed_at` is unix **milliseconds** (`Date.now()`). The scoring
> query normalizes with `created_at * 1000`. Keep this in mind for any new
> queries touching both columns.
## Delete Behaviour (SQLite + Qdrant consistency) ## Delete Behaviour (SQLite + Qdrant consistency)
SQLite cascades handle relational cleanup, but Qdrant is a separate store and SQLite cascades handle relational cleanup, but Qdrant is a separate store and
+31 -20
View File
@@ -73,6 +73,9 @@ via `appSettings.load()` — changes apply immediately without a service restart
| `scoreThreshold` | 0.5 | Minimum similarity score for Qdrant semantic results | | `scoreThreshold` | 0.5 | Minimum similarity score for Qdrant semantic results |
| `semanticWeight` | 1.0 | RRF weight for Qdrant semantic results | | `semanticWeight` | 1.0 | RRF weight for Qdrant semantic results |
| `keywordWeight` | 0 | RRF weight for FTS5 keyword results (`0` = disabled) | | `keywordWeight` | 0 | RRF weight for FTS5 keyword results (`0` = disabled) |
| `contextBudget` | — | Token budget for context assembly (char/4 estimation on stored text) |
| `entityWeight` | — | Scoring bonus for entity-linked episodes in the context pool |
| `minRecentEpisodes` | — | Guaranteed floor of recent episodes always included in context |
| `modelsFolderPath` | `/mnt/nexus-models` | Path to folder containing .gguf files | | `modelsFolderPath` | `/mnt/nexus-models` | Path to folder containing .gguf files |
| `temperature` | 0.7 | Inference temperature | | `temperature` | 0.7 | Inference temperature |
| `repeatPenalty` | 1.1 | Repeat token penalty | | `repeatPenalty` | 1.1 | Repeat token penalty |
@@ -101,35 +104,43 @@ difference is how the inference response is delivered to the client.
4. **Recent episode retrieval** — fetch most recent episodes (`recentEpisodeLimit`). 4. **Recent episode retrieval** — fetch most recent episodes (`recentEpisodeLimit`).
5. **Fused episode retrieval** — runs semantic (Qdrant) and keyword (FTS5) 5. **Trivial-turn gate** — greetings/pleasantries (`isTrivialTurn`) skip all
retrieval (semantic, keyword, entity); recent history alone is the context.
Breaks the greeting → marginal-retrieval → confabulation loop.
6. **Fused episode retrieval** — runs semantic (Qdrant) and keyword (FTS5)
search in parallel, then merges results via Reciprocal Rank Fusion (RRF). search in parallel, then merges results via Reciprocal Rank Fusion (RRF).
Both paths are filtered against `recentIds` before fusion. FTS is scoped The query is embedded once and shared with entity search. Both paths are
to the current session or all project sessions. If `keywordWeight` is `0`, filtered against `recentIds` before fusion. FTS is scoped to the current
the FTS call is skipped entirely. Non-critical — failures fall back to session or all project sessions. If `keywordWeight` is `0`, the FTS call
whichever strategy succeeded. is skipped entirely. Non-critical — failures fall back to whichever
strategy succeeded.
6. **Entity search** — query `entities` Qdrant collection filtered by 7. **Entity search + graph expansion** — query `entities` Qdrant collection
`projectId`. Returns entity IDs alongside Qdrant payload data (the Qdrant (project-scoped, or session-scoped via `/sessions/:id/entity-ids` for
point ID equals the SQLite entity ID). Non-critical. non-project chats). Entity IDs are expanded into a 1-hop subgraph via
`POST /graph/neighbors`; on failure, falls back to flat entity list.
Non-critical.
7. **Graph neighborhood expansion** — call `POST /graph/neighbors` on 8. **Scored pool + budget selection** — `buildScoredPool` combines RRF scores,
memory-service with the entity IDs from step 6. Returns a 1-hop subgraph recency, and entity-linkage bonus; `selectWithinBudget` fills `contextBudget`
`{ nodes, edges }` — entity objects plus the relationships connecting them. (char/4 token estimation on stored text) above a guaranteed floor of
If no entities were found or the graph call fails, falls back to flat entity `minRecentEpisodes` recent episodes. Selected episode IDs are then reported
list (no edges). Non-critical. to `POST /episodes/touch` fire-and-forget (access tracking for the
consolidation lifecycle).
8. **Prompt assembly** — combine system prompt, graph context, fused episodes, 9. **Prompt assembly** — combine system prompt, graph context, selected
recent episodes, and user message. episodes, guaranteed recent episodes, and user message.
9. **Inference** — send to inference service. `/chat` awaits full response; 10. **Inference** — send to inference service. `/chat` awaits full response;
`/chat/stream` pipes SSE chunks to the client. `/chat/stream` pipes SSE chunks to the client.
10. **Episode write** — write exchange back to memory with `projectId`. 11. **Episode write** — write exchange back to memory with `projectId`.
11. **Summarisation trigger** — `triggerSummary(session, allEpisodes)` called 12. **Summarisation trigger** — `triggerSummary(session)` called
fire-and-forget. See `summarization.md` for full details. fire-and-forget. See `summarization.md` for full details.
12. **Auto-naming** — on first message with no session name, fires a secondary 13. **Auto-naming** — on first message with no session name, fires a secondary
inference call (max 20 tokens, temperature 0.3) to generate a session name. inference call (max 20 tokens, temperature 0.3) to generate a session name.
### Prompt Structure ### Prompt Structure
+10
View File
@@ -203,6 +203,16 @@ SUMMARY_MAX_TOKENS=800
SUMMARY_MIN_EPISODES=5 SUMMARY_MIN_EPISODES=5
``` ```
#### `CONSOLIDATION`
Controls the memory consolidation lifecycle (currently dry-run only).
| Key | Value | Description |
|---|---|---|
| `MIN_AGE_DAYS` | `7` | Episodes younger than this are never consolidation candidates |
| `MIN_SESSION_EPISODES` | `20` | Sessions with fewer episodes are skipped entirely |
| `CANDIDATE_LIMIT` | `50` | Max candidates returned per scoring query |
#### `SQLITE` #### `SQLITE`
| Key | Value | Description | | Key | Value | Description |
-31
View File
@@ -1,31 +0,0 @@
// scratch.js — run with: SQLITE_PATH=./packages/memory-service/data/test-copy.db node scratch.js
const { getDB } = require('./packages/memory-service/src/db');
const episodic = require('./packages/memory-service/src/episodic');
const db = getDB();
// Backdate half the episodes to 10-40 days old, give some fake access history
const eps = db.prepare('SELECT id FROM episodes').all();
const now = Math.floor(Date.now() / 1000);
for (const { id } of eps) {
if (Math.random() < 0.5) continue; // leave half young, so the age floor stays testable
const daysOld = 10 + Math.floor(Math.random() * 30);
const createdAt = now - daysOld * 86400;
const touched = Math.random() < 0.6;
db.prepare(`
UPDATE episodes
SET created_at = ?,
access_count = ?,
last_accessed_at = ?
WHERE id = ?
`).run(
createdAt,
touched ? 1 + Math.floor(Math.random() * 5) : 0,
touched ? (createdAt + Math.floor(Math.random() * daysOld) * 86400) * 1000 : null, // accessed sometime after creation, in ms
id
);
}
// Now run the actual candidate function against the doctored data
console.table(episodic.getConsolidationCandidates(19)); // whatever session id has rows