7.1 KiB
Summarization
Session summarization generates rolling plain-text summaries of conversation history, giving the model a condensed view of past context without consuming the full context window with raw episodes.
Location: packages/orchestration-service/src/services/summarization.js
Triggered by: chat/index.js after every episode write (fire-and-forget)
Model: the utility model served by the inference service (/utility/complete), set via UTILITY_MODEL (backed by Ollama on Mini PC 1, 192.168.0.81)
Trigger Conditions
triggerSummary(session) calls maybeSummarize fire-and-forget. It takes only
the session — it no longer receives the full episode list. Instead
maybeSummarize fetches exactly what it needs:
GET /sessions/:id/episode-stats— a cheap aggregate (COUNT,SUM(token_count),MAX(id)) that gates the token threshold without pulling every episode row- Only if over threshold:
GET /sessions/:id/episodes/since/:afterId— the un-summarized tail (episodes newer than the last summary's range), fetched in full for the summary prompt
This replaced an earlier approach where chat/index.js fetched the entire
session (getRecentEpisodes(session.id, 9999)) on every message just to hand it
over — an O(session length) cost per turn. The stats query now runs on every
message; full episode text is fetched only when a summary actually fires.
maybeSummarize proceeds only when both conditions are met:
- Total session token count exceeds
SUMMARIES.THRESHOLD_TOKENS(default 200) - At least
SUMMARIES.MIN_EPISODES_SINCE(default 5) new episodes have accumulated since the last summary
The token threshold is intentionally low — it ensures summaries start generating early in a session's life rather than only after very long conversations.
Summary Rows and Cumulative Updates
Each session can have multiple summary rows in the summaries table.
The update strategy depends on the size of the most recent summary:
| Condition | Action |
|---|---|
| No existing summary | Generate fresh summary from all episodes |
Latest summary under MAX_SUMMARY_TOKENS |
Update: summarise new episodes with existing summary as context |
Latest summary over MAX_SUMMARY_TOKENS |
Create new row: treat as fresh summarisation |
This produces a chain of summary rows over time. Each row's episode_range
covers only the episodes summarised in that specific pass (e.g. 259-263),
not all episodes in the session.
Utility Inference Request
Summaries are generated through the shared utilityInference() helper, which
POSTs to the inference service's /utility/complete endpoint. buildSummaryPrompt
returns a plain instruction string (no template tags) passed as the user message:
const content = await utilityInference({
user: buildSummaryPrompt(episodesToSummarize, existingSummary),
temperature: SUMMARIES.TEMPERATURE, // 0.2
maxTokens: SUMMARIES.SESSION_GEN_MAX_TOKENS, // 500
});
TEMPERATURE (0.2) is slightly higher than extraction (0.1) — summaries benefit
from some fluency. SESSION_GEN_MAX_TOKENS (500) gives room for ~5 thorough
sentences without runoff. Both live in @nexusai/shared SUMMARIES constants.
There is no json: true here — summaries are free-text, unlike entity extraction.
Prompt Format
The prompt is plain text describing the task; the model's own prompt template
(ChatML for qwen, etc.) is applied server-side by the inference service via
Ollama's /api/chat. No <|im_start|> tags belong in this codebase, and the
utility model can be swapped (via UTILITY_MODEL on the inference service) with
no prompt changes here.
Fresh summary instruction:
Summarize the conversation below in 3-5 sentences.
Write in third person. Do not quote directly — paraphrase only.
Do not include greetings, sign-offs, or filler. Output only the summary text.
Conversation:
{context}
Cumulative update instruction:
Update the summary below to incorporate the new exchanges.
Write 3-5 sentences in third person. Do not quote directly — paraphrase only.
Do not include greetings, sign-offs, or filler. Output only the updated summary text.
Previous summary:
{existingSummary}
New exchanges:
{context}
Input truncation
Episode context is truncated to MAX_CHARS = 3000 characters, keeping the
most recent exchanges (sliced from the end). This keeps the model focused and
prevents the prompt from exceeding its effective context window.
Output Handling
Because /api/chat applies and removes the prompt template server-side, the
returned text is already clean — the previous ChatML token-stripping step (and
the class of bug where leaked tokens got stored and re-injected into the next
summarisation prompt) no longer applies.
Episode Range Tracking
Each summary row stores episode_range as "firstId-lastId" covering only
the episodes summarised in that pass:
const summarizedIds = episodesToSummarize.map(ep => ep.id).sort((a,b) => a - b);
const episodeRange = `${summarizedIds.at(0)}-${summarizedIds.at(-1)}`;
This makes SummaryView cards meaningful — "Episodes 259-263" tells you exactly which exchanges that summary covers, rather than always showing the full session range.
Summary Storage
Summaries are written directly to the memory service from orchestration:
// Create new row
await fetch(`${MEMORY_URL}/summaries`, {
method: 'POST',
body: JSON.stringify({ sessionId: session.id, content, tokenCount, episodeRange }),
});
// Update existing row
await fetch(`${MEMORY_URL}/summaries/${latest.id}`, {
method: 'PATCH',
body: JSON.stringify({ content, tokenCount, episodeRange }),
});
session.id here is the internal SQLite integer ID — not the external UUID.
It is available directly on the session object passed from chat/index.js.
Client-Side Indicator
The chat client shows a "Summarising…" spinner in the ChatWindow header
and on the InfoPanel's Session Memory button while summarisation may be
in progress.
Since summarisation is fire-and-forget with no completion signal back to the client, the indicator is timer-based: it activates when the stream finishes and clears after 8 seconds.
// In App.jsx, watching the streaming state from useChat:
useEffect(() => {
if (prevStreaming.current && !streaming) {
setSummarising(true);
const t = setTimeout(() => setSummarising(false), 8000);
return () => clearTimeout(t);
}
prevStreaming.current = streaming;
}, [streaming]);
Environment Variables
Set in packages/orchestration-service/src/.env:
| Variable | Default | Description |
|---|---|---|
INFERENCE_SERVICE_URL |
http://localhost:3001 |
Inference service — summaries route through its /utility/complete endpoint (model set via UTILITY_MODEL there) |
MEMORY_SERVICE_URL |
http://localhost:3002 |
Memory service URL |
SUMMARY_THRESHOLD_TOKENS |
200 |
Token threshold before summarisation triggers |
SUMMARY_MAX_TOKENS |
800 |
Max summary length before a new row is created |
SUMMARY_MIN_EPISODES |
5 |
Min new episodes since last summary before re-summarising |