Architecture
Clippy Chat is a single deployable Node app. One TanStack Start server renders the React SSR UI and serves the file-based JSON/SSE API routes. It talks to Postgres (persistence), a vLLM server (inference), an OIDC provider, and a separate authenticated MCP tool server.
Request flow
red-team runner ─┐ Bearer JWT (scope: clippy-api)
browser ─────────┤ session cookie / OIDC
▼
┌──────────────────────────────────────────────┐
│ Ingress — TLS termination + rate-limit │
└──────────────────────────────────────────────┘
▼
┌──────────────────────────────────────────────┐
│ clippy-chat (TanStack Start, Node 22) │
│ src/routes/** (React pages + /api routes) │
│ resolveUser → requireUser / requireAdmin │
└───────────────┬───────────────┬──────────────┘
┌────┴─────────┬───────────────┐
▼ ▼ ▼
┌──────────────┐ ┌──────────────┐ ┌──────────────────┐
│ vLLM │ │ clippy-mcp │ │ Postgres 17 │
│ /v1/chat/... │ │ bearer auth │ │ conversations, │
│ stream:true │ │ tools │ │ messages, users │
└──────────────┘ └──────────────┘ └──────────────────┘
- Browser → UI. React pages (
index.tsx,login.tsx,admin.tsx,c.$conversationId.tsx) fetch data via React Query calling the client API layersrc/lib/api.ts(plainfetchto/api/*). - Browser / machine → API. Handlers under
src/routes/api/**resolve identity withresolveUser/requireUser/requireAdmin(src/lib/auth/middleware.ts), then hit the DB or vLLM. - API → vLLM. Only
POST /api/chatstreams from vLLM (src/lib/chat/inference.ts) via the OpenAI-compatible/v1/chat/completionsendpoint withstream: true, and re-emits the tokens to the browser as Server-Sent Events. - API → OIDC provider.
openid-clienthandles auth-code + PKCE for web login (src/lib/auth/oidc.ts);joseverifies machine JWTs against the provider's remote JWKS (src/lib/auth/bearer.ts). - API → MCP. Server-side code mints/caches a
clippy-mcp-clienttoken and sends it asAuthorization: Bearerto the ClusterIP-onlyclippy-mcpservice. The MCP server independently validates issuer, audience, client, scope, workspace, signature, and lifetime.
External callers use AIRS instead of the ClusterIP. See the AI Gateway and MCP architecture for both two-stage header modes, identity forwarding, trust boundaries, and production diagrams.
The HTML shell
src/routes/__root.tsx is the document shell. It sets the page title, loads src/styles.css,
and wraps the app in a fresh React Query QueryClient per mount — a new client per SSR
request, so one user's cached data never leaks into another's server-rendered response.
Persistence layer
src/db/client.ts exposes a lazily-created pg.Pool (max 10 connections, 5s connect
timeout). Schema and queries use Drizzle ORM. See
Database Schema for the tables.
Multi-replica awareness
In production the app runs as multiple replicas behind the ingress. That topology shows up in the code as deliberate race handling:
- Admin bootstrap seeds the local admin exactly once across replicas, catching the unique
violation and returning the winner (
src/lib/auth/bootstrap.ts). - Conversation create tolerates two replicas creating the same client-supplied
conversationIdconcurrently (src/lib/chat/service.ts). - OIDC discovery / JWKS are cached per process and refetched on demand.
Dev vs. production server
- Dev:
npm run dev→ Vite dev server; TanStack Start serves SSR itself. - Production:
npm run buildbundles the server with Nitro to.output/server/index.mjs;npm startruns it. Nitro is build-only — it is not used by the dev server.
See Deployment for Docker and Kubernetes.