Files
nexusAI/docs/services/summarization.md
T
2026-08-17 23:48:15 -07:00

7.1 KiB

Summarization

Session summarization generates rolling plain-text summaries of conversation history, giving the model a condensed view of past context without consuming the full context window with raw episodes.

Location: packages/orchestration-service/src/services/summarization.js
Triggered by: chat/index.js after every episode write (fire-and-forget)
Model: the utility model served by the inference service (/utility/complete), set via UTILITY_MODEL (backed by Ollama on Mini PC 1, 192.168.0.81)


Trigger Conditions

triggerSummary(session) calls maybeSummarize fire-and-forget. It takes only the session — it no longer receives the full episode list. Instead maybeSummarize fetches exactly what it needs:

  1. GET /sessions/:id/episode-stats — a cheap aggregate (COUNT, SUM(token_count), MAX(id)) that gates the token threshold without pulling every episode row
  2. Only if over threshold: GET /sessions/:id/episodes/since/:afterId — the un-summarized tail (episodes newer than the last summary's range), fetched in full for the summary prompt

This replaced an earlier approach where chat/index.js fetched the entire session (getRecentEpisodes(session.id, 9999)) on every message just to hand it over — an O(session length) cost per turn. The stats query now runs on every message; full episode text is fetched only when a summary actually fires.

maybeSummarize proceeds only when both conditions are met:

  1. Total session token count exceeds SUMMARIES.THRESHOLD_TOKENS (default 200)
  2. At least SUMMARIES.MIN_EPISODES_SINCE (default 5) new episodes have accumulated since the last summary

The token threshold is intentionally low — it ensures summaries start generating early in a session's life rather than only after very long conversations.


Summary Rows and Cumulative Updates

Each session can have multiple summary rows in the summaries table. The update strategy depends on the size of the most recent summary:

Condition Action
No existing summary Generate fresh summary from all episodes
Latest summary under MAX_SUMMARY_TOKENS Update: summarise new episodes with existing summary as context
Latest summary over MAX_SUMMARY_TOKENS Create new row: treat as fresh summarisation

This produces a chain of summary rows over time. Each row's episode_range covers only the episodes summarised in that specific pass (e.g. 259-263), not all episodes in the session.


Utility Inference Request

Summaries are generated through the shared utilityInference() helper, which POSTs to the inference service's /utility/complete endpoint. buildSummaryPrompt returns a plain instruction string (no template tags) passed as the user message:

const content = await utilityInference({
    user: buildSummaryPrompt(episodesToSummarize, existingSummary),
    temperature: SUMMARIES.TEMPERATURE,            // 0.2
    maxTokens:   SUMMARIES.SESSION_GEN_MAX_TOKENS,  // 500
});

TEMPERATURE (0.2) is slightly higher than extraction (0.1) — summaries benefit from some fluency. SESSION_GEN_MAX_TOKENS (500) gives room for ~5 thorough sentences without runoff. Both live in @nexusai/shared SUMMARIES constants.

There is no json: true here — summaries are free-text, unlike entity extraction.


Prompt Format

The prompt is plain text describing the task; the model's own prompt template (ChatML for qwen, etc.) is applied server-side by the inference service via Ollama's /api/chat. No <|im_start|> tags belong in this codebase, and the utility model can be swapped (via UTILITY_MODEL on the inference service) with no prompt changes here.

Fresh summary instruction:

Summarize the conversation below in 3-5 sentences.
Write in third person. Do not quote directly — paraphrase only.
Do not include greetings, sign-offs, or filler. Output only the summary text.

Conversation:
{context}

Cumulative update instruction:

Update the summary below to incorporate the new exchanges.
Write 3-5 sentences in third person. Do not quote directly — paraphrase only.
Do not include greetings, sign-offs, or filler. Output only the updated summary text.

Previous summary:
{existingSummary}

New exchanges:
{context}

Input truncation

Episode context is truncated to MAX_CHARS = 3000 characters, keeping the most recent exchanges (sliced from the end). This keeps the model focused and prevents the prompt from exceeding its effective context window.


Output Handling

Because /api/chat applies and removes the prompt template server-side, the returned text is already clean — the previous ChatML token-stripping step (and the class of bug where leaked tokens got stored and re-injected into the next summarisation prompt) no longer applies.


Episode Range Tracking

Each summary row stores episode_range as "firstId-lastId" covering only the episodes summarised in that pass:

const summarizedIds = episodesToSummarize.map(ep => ep.id).sort((a,b) => a - b);
const episodeRange = `${summarizedIds.at(0)}-${summarizedIds.at(-1)}`;

This makes SummaryView cards meaningful — "Episodes 259-263" tells you exactly which exchanges that summary covers, rather than always showing the full session range.


Summary Storage

Summaries are written directly to the memory service from orchestration:

// Create new row
await fetch(`${MEMORY_URL}/summaries`, {
    method: 'POST',
    body: JSON.stringify({ sessionId: session.id, content, tokenCount, episodeRange }),
});

// Update existing row
await fetch(`${MEMORY_URL}/summaries/${latest.id}`, {
    method: 'PATCH',
    body: JSON.stringify({ content, tokenCount, episodeRange }),
});

session.id here is the internal SQLite integer ID — not the external UUID. It is available directly on the session object passed from chat/index.js.


Client-Side Indicator

The chat client shows a "Summarising…" spinner in the ChatWindow header and on the InfoPanel's Session Memory button while summarisation may be in progress.

Since summarisation is fire-and-forget with no completion signal back to the client, the indicator is timer-based: it activates when the stream finishes and clears after 8 seconds.

// In App.jsx, watching the streaming state from useChat:
useEffect(() => {
    if (prevStreaming.current && !streaming) {
        setSummarising(true);
        const t = setTimeout(() => setSummarising(false), 8000);
        return () => clearTimeout(t);
    }
    prevStreaming.current = streaming;
}, [streaming]);

Environment Variables

Set in packages/orchestration-service/src/.env:

Variable Default Description
INFERENCE_SERVICE_URL http://localhost:3001 Inference service — summaries route through its /utility/complete endpoint (model set via UTILITY_MODEL there)
MEMORY_SERVICE_URL http://localhost:3002 Memory service URL
SUMMARY_THRESHOLD_TOKENS 200 Token threshold before summarisation triggers
SUMMARY_MAX_TOKENS 800 Max summary length before a new row is created
SUMMARY_MIN_EPISODES 5 Min new episodes since last summary before re-summarising