documentation update
This commit is contained in:
@@ -78,6 +78,11 @@ does **not** reconcile columns on old tables — that's what migrations are for)
|
||||
```js
|
||||
const migrations = [
|
||||
(_db) => {}, // v0 → v1: consolidated baseline (historical ALTERs folded into schema.js)
|
||||
(db) => { // v1 → v2: access tracking for consolidation lifecycle
|
||||
db.exec(`ALTER TABLE episodes ADD COLUMN last_accessed_at INTEGER`);
|
||||
db.exec(`ALTER TABLE episodes ADD COLUMN access_count INTEGER NOT NULL DEFAULT 0`);
|
||||
db.exec(`UPDATE episodes SET last_accessed_at = created_at`); // backfill
|
||||
},
|
||||
];
|
||||
const LATEST_VERSION = migrations.length; // derived, never hand-maintained
|
||||
```
|
||||
@@ -121,6 +126,12 @@ previously ran on every startup.
|
||||
- `foreign_keys = ON` — enforces referential integrity and cascade deletes
|
||||
- PRAGMAs set via `db.pragma()`, not `db.exec()`
|
||||
|
||||
> **Copying a live WAL database:** `cp` on the `.db` file alone silently loses
|
||||
> everything in the un-checkpointed `-wal` file (recent writes, even the
|
||||
> migration version stamp). Always use
|
||||
> `sqlite3 nexusai.db "VACUUM INTO './copy.db'"` (or `.backup`) — safe while
|
||||
> the service is running, produces a complete single-file snapshot.
|
||||
|
||||
### Dynamic Updates
|
||||
|
||||
Both `updateSession` and `updateProject` build their `SET` clause dynamically
|
||||
@@ -204,6 +215,28 @@ service is responsible only for CRUD — generation logic lives in orchestration
|
||||
> For full details on trigger conditions, prompt format, cumulative updates,
|
||||
> and ChatML token stripping, see `summarization.md`.
|
||||
|
||||
## Access Tracking & Consolidation (dry-run)
|
||||
|
||||
Every episode selected into a chat context window (budget-selected, not the
|
||||
guaranteed-recency floor) gets an access bump via `POST /episodes/touch` —
|
||||
`access_count` incremented, `last_accessed_at` set to `Date.now()` (ms).
|
||||
Called fire-and-forget from orchestration; a failure loses one increment,
|
||||
nothing more.
|
||||
|
||||
`GET /sessions/:id/consolidation-candidates` scores episodes by
|
||||
`access_count / (1 + days since last access)` — never-accessed episodes fall
|
||||
back to `created_at` for the recency term and score exactly 0 (most eligible).
|
||||
Two floors apply: episodes younger than `CONSOLIDATION.MIN_AGE_DAYS` are
|
||||
excluded in SQL; sessions under `CONSOLIDATION.MIN_SESSION_EPISODES` return
|
||||
`eligible: false` before scoring runs. The endpoint is observe-only — the
|
||||
destructive pass (merge → summarize → Qdrant cleanup → orphan sweep) is not
|
||||
yet built.
|
||||
|
||||
> **Unit note:** `created_at` is unix **seconds** (`unixepoch()`);
|
||||
> `last_accessed_at` is unix **milliseconds** (`Date.now()`). The scoring
|
||||
> query normalizes with `created_at * 1000`. Keep this in mind for any new
|
||||
> queries touching both columns.
|
||||
|
||||
## Delete Behaviour (SQLite + Qdrant consistency)
|
||||
|
||||
SQLite cascades handle relational cleanup, but Qdrant is a separate store and
|
||||
|
||||
@@ -73,6 +73,9 @@ via `appSettings.load()` — changes apply immediately without a service restart
|
||||
| `scoreThreshold` | 0.5 | Minimum similarity score for Qdrant semantic results |
|
||||
| `semanticWeight` | 1.0 | RRF weight for Qdrant semantic results |
|
||||
| `keywordWeight` | 0 | RRF weight for FTS5 keyword results (`0` = disabled) |
|
||||
| `contextBudget` | — | Token budget for context assembly (char/4 estimation on stored text) |
|
||||
| `entityWeight` | — | Scoring bonus for entity-linked episodes in the context pool |
|
||||
| `minRecentEpisodes` | — | Guaranteed floor of recent episodes always included in context |
|
||||
| `modelsFolderPath` | `/mnt/nexus-models` | Path to folder containing .gguf files |
|
||||
| `temperature` | 0.7 | Inference temperature |
|
||||
| `repeatPenalty` | 1.1 | Repeat token penalty |
|
||||
@@ -101,35 +104,43 @@ difference is how the inference response is delivered to the client.
|
||||
|
||||
4. **Recent episode retrieval** — fetch most recent episodes (`recentEpisodeLimit`).
|
||||
|
||||
5. **Fused episode retrieval** — runs semantic (Qdrant) and keyword (FTS5)
|
||||
5. **Trivial-turn gate** — greetings/pleasantries (`isTrivialTurn`) skip all
|
||||
retrieval (semantic, keyword, entity); recent history alone is the context.
|
||||
Breaks the greeting → marginal-retrieval → confabulation loop.
|
||||
|
||||
6. **Fused episode retrieval** — runs semantic (Qdrant) and keyword (FTS5)
|
||||
search in parallel, then merges results via Reciprocal Rank Fusion (RRF).
|
||||
Both paths are filtered against `recentIds` before fusion. FTS is scoped
|
||||
to the current session or all project sessions. If `keywordWeight` is `0`,
|
||||
the FTS call is skipped entirely. Non-critical — failures fall back to
|
||||
whichever strategy succeeded.
|
||||
The query is embedded once and shared with entity search. Both paths are
|
||||
filtered against `recentIds` before fusion. FTS is scoped to the current
|
||||
session or all project sessions. If `keywordWeight` is `0`, the FTS call
|
||||
is skipped entirely. Non-critical — failures fall back to whichever
|
||||
strategy succeeded.
|
||||
|
||||
6. **Entity search** — query `entities` Qdrant collection filtered by
|
||||
`projectId`. Returns entity IDs alongside Qdrant payload data (the Qdrant
|
||||
point ID equals the SQLite entity ID). Non-critical.
|
||||
7. **Entity search + graph expansion** — query `entities` Qdrant collection
|
||||
(project-scoped, or session-scoped via `/sessions/:id/entity-ids` for
|
||||
non-project chats). Entity IDs are expanded into a 1-hop subgraph via
|
||||
`POST /graph/neighbors`; on failure, falls back to flat entity list.
|
||||
Non-critical.
|
||||
|
||||
7. **Graph neighborhood expansion** — call `POST /graph/neighbors` on
|
||||
memory-service with the entity IDs from step 6. Returns a 1-hop subgraph
|
||||
`{ nodes, edges }` — entity objects plus the relationships connecting them.
|
||||
If no entities were found or the graph call fails, falls back to flat entity
|
||||
list (no edges). Non-critical.
|
||||
8. **Scored pool + budget selection** — `buildScoredPool` combines RRF scores,
|
||||
recency, and entity-linkage bonus; `selectWithinBudget` fills `contextBudget`
|
||||
(char/4 token estimation on stored text) above a guaranteed floor of
|
||||
`minRecentEpisodes` recent episodes. Selected episode IDs are then reported
|
||||
to `POST /episodes/touch` fire-and-forget (access tracking for the
|
||||
consolidation lifecycle).
|
||||
|
||||
8. **Prompt assembly** — combine system prompt, graph context, fused episodes,
|
||||
recent episodes, and user message.
|
||||
9. **Prompt assembly** — combine system prompt, graph context, selected
|
||||
episodes, guaranteed recent episodes, and user message.
|
||||
|
||||
9. **Inference** — send to inference service. `/chat` awaits full response;
|
||||
`/chat/stream` pipes SSE chunks to the client.
|
||||
10. **Inference** — send to inference service. `/chat` awaits full response;
|
||||
`/chat/stream` pipes SSE chunks to the client.
|
||||
|
||||
10. **Episode write** — write exchange back to memory with `projectId`.
|
||||
11. **Episode write** — write exchange back to memory with `projectId`.
|
||||
|
||||
11. **Summarisation trigger** — `triggerSummary(session, allEpisodes)` called
|
||||
12. **Summarisation trigger** — `triggerSummary(session)` called
|
||||
fire-and-forget. See `summarization.md` for full details.
|
||||
|
||||
12. **Auto-naming** — on first message with no session name, fires a secondary
|
||||
13. **Auto-naming** — on first message with no session name, fires a secondary
|
||||
inference call (max 20 tokens, temperature 0.3) to generate a session name.
|
||||
|
||||
### Prompt Structure
|
||||
|
||||
@@ -203,6 +203,16 @@ SUMMARY_MAX_TOKENS=800
|
||||
SUMMARY_MIN_EPISODES=5
|
||||
```
|
||||
|
||||
#### `CONSOLIDATION`
|
||||
|
||||
Controls the memory consolidation lifecycle (currently dry-run only).
|
||||
|
||||
| Key | Value | Description |
|
||||
|---|---|---|
|
||||
| `MIN_AGE_DAYS` | `7` | Episodes younger than this are never consolidation candidates |
|
||||
| `MIN_SESSION_EPISODES` | `20` | Sessions with fewer episodes are skipped entirely |
|
||||
| `CANDIDATE_LIMIT` | `50` | Max candidates returned per scoring query |
|
||||
|
||||
#### `SQLITE`
|
||||
|
||||
| Key | Value | Description |
|
||||
|
||||
Reference in New Issue
Block a user