Chat Pipeline & Streaming
Everything a chat turn touches, from POST /api/chat to streamed tokens.
The flow
POST /api/chat (src/routes/api/chat.ts):
requireUser(request)— any of the three auth modes.- Validate body with zod:
{conversationId: uuid, message: string 1..8000}→400on failure. ensureConversation(conversationId, user.id, message)— theconversationIdis client-supplied:- exists, owned, not deleted → reuse;
- exists but foreign or deleted →
NotFoundError→404; - absent → created in one transaction: the
conversationsrow (model = INFERENCE_MODEL,title = message.slice(0,80)) and a seedmessagesrow (role: 'system', the Clippy persona), so a conversation is never observable without its system row. A concurrent create on the same id is caught via unique violation → the winner is returned.
- Persist the user turn —
appendMessage(id, 'user', message)(also bumpsupdatedAt). - Build the prompt —
foldSystem(trimHistory(loadHistory(id), CONTEXT_BUDGET))whereCONTEXT_BUDGET = 8192 − 1024 = 7168tokens (reserving 1024 for the completion). - Stream — open a
ReadableStream; for each event fromchatStream(history, signal)emit SSE, accumulate the assistant text and token usage; on completion persist the assistant row and emitdone.
Chat-template compatibility normalization
The production model is @vllm2/unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_XL, served through the AIRS
gateway. Clippy retains two conservative transforms in
src/lib/chat/history.ts from its earlier Mistral deployment so persisted conversations remain
portable across strict chat templates:
foldSystem— strict templates may reject thesystemrole, so all system messages are concatenated and prepended to the first user turn's content; standalone system entries are dropped.trimHistory— keeps the system message(s) + newest turns that fitmaxTokens(approxceil(len/4)tokens per message), dropping the oldest prefix. Then it enforces a user-first guarantee: leadingassistantturns are shifted off so the first non-system turn is always auserturn (strict alternating-role templates reject a payload that starts withassistant).
:::tip Swapping models
INFERENCE_MODEL sets the model. If you point Clippy at a model without strict alternation rules,
the folding/trimming is still safe — it just becomes a no-op-ish normalization.
:::
The persona
CLIPPY_SYSTEM_PROMPT (src/lib/chat/persona.ts):
You are Clippy 📎, a friendly, upbeat AI assistant … Light paperclip humor is welcome; never let it get in the way of a good answer.
The inference call
chatStream (src/lib/chat/inference.ts) does POST ${INFERENCE_BASE_URL}/v1/chat/completions with:
{
"model": "<INFERENCE_MODEL>",
"messages": [ ... ],
"stream": true,
"stream_options": { "include_usage": true }
}
- Adds
Authorization: Bearer ${INFERENCE_API_KEY}only ifINFERENCE_API_KEYis set (cluster vLLM runs with--api-key; a dev vLLM may not). - Parses the upstream SSE stream, yielding
{type:'delta', content}and{type:'usage', promptTokens, completionTokens}. A non-2xx (or missing body) throwsinference error: <status> <text>. The reader is always cancelled infinally.
The SSE contract (app → client)
Wire format is event: <name>\ndata: <json>\n\n:
| Event | Data | Meaning |
|---|---|---|
delta | {"content": "<token chunk>"} | one chunk of the reply — concatenate |
done | {"messageId","promptTokens?","completionTokens?"} | end marker, after the assistant row is persisted |
error | {"message":"inference failed, try again"} | a real inference error (not a client disconnect) |
Response headers: content-type: text/event-stream, cache-control: no-cache,
connection: keep-alive.
Interruption handling
- Client disconnects (
request.signal.aborted) with partial content → persist the partial assistant row withinterrupted: true, emit nothing. - Inference error → persist the partial row with
interrupted: true, emiterror, log. safeEnqueueswallows enqueue-after-close races.
The client consumer streamChat (src/lib/api.ts) POSTs, redirects to /login on 401, parses
the SSE frames, and invokes on.delta / on.done / on.error. The red-team adapter reimplements
this same parse in Python — see The Adapter.