documentation update
This commit is contained in:
@@ -147,4 +147,36 @@ data: [DONE]
|
||||
`model` and `tokenCount` are captured from the llama.cpp `finish_reason: stop`
|
||||
chunk and emitted on the done event.
|
||||
|
||||
### SSE parsing (buffered)
|
||||
|
||||
`llama-server` emits SSE events as `data:` lines, but a single network chunk
|
||||
from `res.body` is **not** guaranteed to contain whole lines — an event can be
|
||||
split across two chunks (mid-`data:` prefix or even mid-JSON). The provider
|
||||
therefore accumulates a string `buffer`, splits on `\n`, and retains the final
|
||||
(possibly incomplete) element for the next chunk rather than parsing each raw
|
||||
chunk in isolation:
|
||||
|
||||
```js
|
||||
let buffer = '';
|
||||
for await (const chunk of res.body) {
|
||||
buffer += Buffer.from(chunk).toString('utf-8');
|
||||
const lines = buffer.split('\n');
|
||||
buffer = lines.pop() ?? ''; // keep incomplete tail for next chunk
|
||||
for (const line of lines) yield* processLine(line.trim());
|
||||
}
|
||||
if (buffer.trim()) yield* processLine(buffer.trim()); // flush remainder
|
||||
```
|
||||
|
||||
`processLine()` handles exactly one line: it skips non-`data:` lines and
|
||||
`[DONE]`, guards `JSON.parse` in a try/catch (a malformed line logs a warning
|
||||
and is skipped rather than throwing and killing the stream), extracts
|
||||
`delta.content`, and captures `model` / token usage from the terminal chunks.
|
||||
Content is emitted only when a delta is present (`if (delta) yield { response,
|
||||
done: false }`).
|
||||
|
||||
> This mirrors the same buffering the orchestration service uses when reading
|
||||
> the inference stream. Without it, long responses intermittently die
|
||||
> mid-stream on a chunk boundary, and a run where every content delta lands on
|
||||
> a split line surfaces as an empty response with a `done` event.
|
||||
|
||||
For all HTTP endpoints, see `api-routes.md`.
|
||||
Reference in New Issue
Block a user