204 lines
7.1 KiB
Markdown
204 lines
7.1 KiB
Markdown
# Summarization
|
|
|
|
Session summarization generates rolling plain-text summaries of conversation
|
|
history, giving the model a condensed view of past context without consuming
|
|
the full context window with raw episodes.
|
|
|
|
**Location:** `packages/orchestration-service/src/services/summarization.js`
|
|
**Triggered by:** `chat/index.js` after every episode write (fire-and-forget)
|
|
**Model:** the utility model served by the inference service (`/utility/complete`), set via `UTILITY_MODEL` (backed by Ollama on Mini PC 1, 192.168.0.81)
|
|
|
|
---
|
|
|
|
## Trigger Conditions
|
|
|
|
`triggerSummary(session)` calls `maybeSummarize` fire-and-forget. It takes only
|
|
the `session` — it no longer receives the full episode list. Instead
|
|
`maybeSummarize` fetches exactly what it needs:
|
|
|
|
1. `GET /sessions/:id/episode-stats` — a cheap aggregate (`COUNT`, `SUM(token_count)`,
|
|
`MAX(id)`) that gates the token threshold **without** pulling every episode row
|
|
2. Only if over threshold: `GET /sessions/:id/episodes/since/:afterId` — the
|
|
un-summarized tail (episodes newer than the last summary's range), fetched in
|
|
full for the summary prompt
|
|
|
|
This replaced an earlier approach where `chat/index.js` fetched the entire
|
|
session (`getRecentEpisodes(session.id, 9999)`) on every message just to hand it
|
|
over — an O(session length) cost per turn. The stats query now runs on every
|
|
message; full episode text is fetched only when a summary actually fires.
|
|
|
|
`maybeSummarize` proceeds only when both conditions are met:
|
|
|
|
1. Total session token count exceeds `SUMMARIES.THRESHOLD_TOKENS` (default 200)
|
|
2. At least `SUMMARIES.MIN_EPISODES_SINCE` (default 5) new episodes have
|
|
accumulated since the last summary
|
|
|
|
The token threshold is intentionally low — it ensures summaries start
|
|
generating early in a session's life rather than only after very long
|
|
conversations.
|
|
|
|
---
|
|
|
|
## Summary Rows and Cumulative Updates
|
|
|
|
Each session can have multiple summary rows in the `summaries` table.
|
|
The update strategy depends on the size of the most recent summary:
|
|
|
|
| Condition | Action |
|
|
|---|---|
|
|
| No existing summary | Generate fresh summary from all episodes |
|
|
| Latest summary under `MAX_SUMMARY_TOKENS` | Update: summarise new episodes with existing summary as context |
|
|
| Latest summary over `MAX_SUMMARY_TOKENS` | Create new row: treat as fresh summarisation |
|
|
|
|
This produces a chain of summary rows over time. Each row's `episode_range`
|
|
covers only the episodes summarised in that specific pass (e.g. `259-263`),
|
|
not all episodes in the session.
|
|
|
|
---
|
|
|
|
## Utility Inference Request
|
|
|
|
Summaries are generated through the shared `utilityInference()` helper, which
|
|
POSTs to the inference service's `/utility/complete` endpoint. `buildSummaryPrompt`
|
|
returns a plain instruction string (no template tags) passed as the `user` message:
|
|
|
|
```js
|
|
const content = await utilityInference({
|
|
user: buildSummaryPrompt(episodesToSummarize, existingSummary),
|
|
temperature: SUMMARIES.TEMPERATURE, // 0.2
|
|
maxTokens: SUMMARIES.SESSION_GEN_MAX_TOKENS, // 500
|
|
});
|
|
```
|
|
|
|
`TEMPERATURE` (0.2) is slightly higher than extraction (0.1) — summaries benefit
|
|
from some fluency. `SESSION_GEN_MAX_TOKENS` (500) gives room for ~5 thorough
|
|
sentences without runoff. Both live in `@nexusai/shared` `SUMMARIES` constants.
|
|
|
|
There is no `json: true` here — summaries are free-text, unlike entity extraction.
|
|
|
|
---
|
|
|
|
## Prompt Format
|
|
|
|
The prompt is plain text describing the task; the model's own prompt template
|
|
(ChatML for qwen, etc.) is applied **server-side** by the inference service via
|
|
Ollama's `/api/chat`. No `<|im_start|>` tags belong in this codebase, and the
|
|
utility model can be swapped (via `UTILITY_MODEL` on the inference service) with
|
|
no prompt changes here.
|
|
|
|
Fresh summary instruction:
|
|
|
|
```
|
|
Summarize the conversation below in 3-5 sentences.
|
|
Write in third person. Do not quote directly — paraphrase only.
|
|
Do not include greetings, sign-offs, or filler. Output only the summary text.
|
|
|
|
Conversation:
|
|
{context}
|
|
```
|
|
|
|
Cumulative update instruction:
|
|
|
|
```
|
|
Update the summary below to incorporate the new exchanges.
|
|
Write 3-5 sentences in third person. Do not quote directly — paraphrase only.
|
|
Do not include greetings, sign-offs, or filler. Output only the updated summary text.
|
|
|
|
Previous summary:
|
|
{existingSummary}
|
|
|
|
New exchanges:
|
|
{context}
|
|
```
|
|
|
|
### Input truncation
|
|
|
|
Episode context is truncated to `MAX_CHARS = 3000` characters, keeping the
|
|
most recent exchanges (sliced from the end). This keeps the model focused and
|
|
prevents the prompt from exceeding its effective context window.
|
|
|
|
---
|
|
|
|
## Output Handling
|
|
|
|
Because `/api/chat` applies and removes the prompt template server-side, the
|
|
returned text is already clean — the previous ChatML token-stripping step (and
|
|
the class of bug where leaked tokens got stored and re-injected into the next
|
|
summarisation prompt) no longer applies.
|
|
|
|
---
|
|
|
|
## Episode Range Tracking
|
|
|
|
Each summary row stores `episode_range` as `"firstId-lastId"` covering only
|
|
the episodes summarised in that pass:
|
|
|
|
```js
|
|
const summarizedIds = episodesToSummarize.map(ep => ep.id).sort((a,b) => a - b);
|
|
const episodeRange = `${summarizedIds.at(0)}-${summarizedIds.at(-1)}`;
|
|
```
|
|
|
|
This makes SummaryView cards meaningful — "Episodes 259-263" tells you
|
|
exactly which exchanges that summary covers, rather than always showing
|
|
the full session range.
|
|
|
|
---
|
|
|
|
## Summary Storage
|
|
|
|
Summaries are written directly to the memory service from orchestration:
|
|
|
|
```js
|
|
// Create new row
|
|
await fetch(`${MEMORY_URL}/summaries`, {
|
|
method: 'POST',
|
|
body: JSON.stringify({ sessionId: session.id, content, tokenCount, episodeRange }),
|
|
});
|
|
|
|
// Update existing row
|
|
await fetch(`${MEMORY_URL}/summaries/${latest.id}`, {
|
|
method: 'PATCH',
|
|
body: JSON.stringify({ content, tokenCount, episodeRange }),
|
|
});
|
|
```
|
|
|
|
`session.id` here is the internal SQLite integer ID — not the external UUID.
|
|
It is available directly on the `session` object passed from `chat/index.js`.
|
|
|
|
---
|
|
|
|
## Client-Side Indicator
|
|
|
|
The chat client shows a "Summarising…" spinner in the `ChatWindow` header
|
|
and on the InfoPanel's Session Memory button while summarisation may be
|
|
in progress.
|
|
|
|
Since summarisation is fire-and-forget with no completion signal back to
|
|
the client, the indicator is timer-based: it activates when the stream
|
|
finishes and clears after 8 seconds.
|
|
|
|
```js
|
|
// In App.jsx, watching the streaming state from useChat:
|
|
useEffect(() => {
|
|
if (prevStreaming.current && !streaming) {
|
|
setSummarising(true);
|
|
const t = setTimeout(() => setSummarising(false), 8000);
|
|
return () => clearTimeout(t);
|
|
}
|
|
prevStreaming.current = streaming;
|
|
}, [streaming]);
|
|
```
|
|
|
|
---
|
|
|
|
## Environment Variables
|
|
|
|
Set in `packages/orchestration-service/src/.env`:
|
|
|
|
| Variable | Default | Description |
|
|
|---|---|---|
|
|
| `INFERENCE_SERVICE_URL` | `http://localhost:3001` | Inference service — summaries route through its `/utility/complete` endpoint (model set via `UTILITY_MODEL` there) |
|
|
| `MEMORY_SERVICE_URL` | `http://localhost:3002` | Memory service URL |
|
|
| `SUMMARY_THRESHOLD_TOKENS` | `200` | Token threshold before summarisation triggers |
|
|
| `SUMMARY_MAX_TOKENS` | `800` | Max summary length before a new row is created |
|
|
| `SUMMARY_MIN_EPISODES` | `5` | Min new episodes since last summary before re-summarising |s |