Skip to main content

Architecture

Clippy Chat is a single deployable Node app. One TanStack Start server renders the React SSR UI and serves the file-based JSON/SSE API routes. It talks to Postgres (persistence), a vLLM server (inference), an OIDC provider, and a separate authenticated MCP tool server.

Request flow​

red-team runner ─┐ Bearer JWT (scope: clippy-api)
browser ─────────┤ session cookie / OIDC
▼
┌──────────────────────────────────────────────┐
│ Ingress — TLS termination + rate-limit │
└──────────────────────────────────────────────┘
▼
┌──────────────────────────────────────────────┐
│ clippy-chat (TanStack Start, Node 22) │
│ src/routes/** (React pages + /api routes) │
│ resolveUser → requireUser / requireAdmin │
└───────────────┬───────────────┬──────────────┘
┌────┴─────────┬───────────────┐
▼ ▼ ▼
┌──────────────┐ ┌──────────────┐ ┌──────────────────┐
│ vLLM │ │ clippy-mcp │ │ Postgres 17 │
│ /v1/chat/... │ │ bearer auth │ │ conversations, │
│ stream:true │ │ tools │ │ messages, users │
└──────────────┘ └──────────────┘ └──────────────────┘
  • Browser → UI. React pages (index.tsx, login.tsx, admin.tsx, c.$conversationId.tsx) fetch data via React Query calling the client API layer src/lib/api.ts (plain fetch to /api/*).
  • Browser / machine → API. Handlers under src/routes/api/** resolve identity with resolveUser / requireUser / requireAdmin (src/lib/auth/middleware.ts), then hit the DB or vLLM.
  • API → vLLM. Only POST /api/chat streams from vLLM (src/lib/chat/inference.ts) via the OpenAI-compatible /v1/chat/completions endpoint with stream: true, and re-emits the tokens to the browser as Server-Sent Events.
  • API → OIDC provider. openid-client handles auth-code + PKCE for web login (src/lib/auth/oidc.ts); jose verifies machine JWTs against the provider's remote JWKS (src/lib/auth/bearer.ts).
  • API → MCP. Server-side code mints/caches a clippy-mcp-client token and sends it as Authorization: Bearer to the ClusterIP-only clippy-mcp service. The MCP server independently validates issuer, audience, client, scope, workspace, signature, and lifetime.

External callers use AIRS instead of the ClusterIP. See the AI Gateway and MCP architecture for both two-stage header modes, identity forwarding, trust boundaries, and production diagrams.

The HTML shell​

src/routes/__root.tsx is the document shell. It sets the page title, loads src/styles.css, and wraps the app in a fresh React Query QueryClient per mount — a new client per SSR request, so one user's cached data never leaks into another's server-rendered response.

Persistence layer​

src/db/client.ts exposes a lazily-created pg.Pool (max 10 connections, 5s connect timeout). Schema and queries use Drizzle ORM. See Database Schema for the tables.

Multi-replica awareness​

In production the app runs as multiple replicas behind the ingress. That topology shows up in the code as deliberate race handling:

  • Admin bootstrap seeds the local admin exactly once across replicas, catching the unique violation and returning the winner (src/lib/auth/bootstrap.ts).
  • Conversation create tolerates two replicas creating the same client-supplied conversationId concurrently (src/lib/chat/service.ts).
  • OIDC discovery / JWKS are cached per process and refetched on demand.

Dev vs. production server​

  • Dev: npm run dev → Vite dev server; TanStack Start serves SSR itself.
  • Production: npm run build bundles the server with Nitro to .output/server/index.mjs; npm start runs it. Nitro is build-only — it is not used by the dev server.

See Deployment for Docker and Kubernetes.