Zero dependencies · Node ≥ 18 · Apache-2.0

llm-session-proxy

A configurable local reverse proxy for LLM APIs. It fills in the session header your client never sends, injects arbitrary custom parameters, rewrites model aliases and request paths, and can convert between the OpenAI, Anthropic and Responses protocols — then forwards the request, streaming SSE responses included.

npx llm-session-proxy

The problem it solves

OpenCode Go wants a session identifier on every request, and most clients cannot set custom headers.

The upstream expects a header like this:

x-opencode-session : stable per conversation (used for prompt caching and routing)

If the client does not send it, the upstream refuses the call outright:

400 Request is missing x-opencode-session and cannot be routed efficiently

This tool moves that job onto your machine: point the client's Base URL at the local proxy and the proxy handles the rest. Nothing in it is tied to one provider — swap upstream and inject.headers and it fronts any OpenAI- or Anthropic-compatible endpoint.

What it does

Every item is adjustable through a config file, an environment variable, or a CLI flag.

session

Automatic session IDs

Three-tier strategy: explicit client identifier → content fingerprint (system + first user message) → one-off random. Stable within a conversation, so prompt caching actually works.

inject

Arbitrary parameter injection

Inject request headers and request body fields. Values support templates such as {{session.id}}, {{uuid}} and {{env.HOME}}.

model

Model alias rewriting

Prefix stripping (proxy-glm → glm-5.3) and exact mapping, and the two compose. The OpenCode Go alias table ships built in, so the proxy- prefix the client needs actually resolves — and an alias that strips down to nothing is reported by name instead of being forwarded silently.

router

Rule-based routing into buckets

Four buckets — default / background / think / longContext — each with its own model override and body transforms. Rules match on path prefix, model prefix, a body field or request size; they run top-down with first match winning, and the conditions inside one rule are ANDed. Off by default, so upgrading changes nothing.

transform

Composable body transforms

Five named, in-tree transforms (drop-fields, drop-empty-fields, rename-fields, clamp-max-tokens, noop), applied globally or per bucket in a fixed order. The registry is a hard-coded list on purpose: the proxy never loads code from a path.

protocol

Protocol conversion

OpenAI Chat ↔ Anthropic Messages ↔ OpenAI Responses, for request bodies, JSON responses and event-by-event SSE transcoding. A client that speaks one protocol reaches models that speak another. Routed by model prefix, off by default.

path

Request path rewriting

When the client can only be pointed at /v1, one regex moves it to the path the upstream actually wants.

stream

Unbuffered SSE pass-through

Chunks are forwarded as they arrive instead of being buffered, so streaming stays streaming. Upstream 4xx bodies come back verbatim.

ops

Logs on by default

~/.lsp/logs/llm-session-proxy.log with size or date rotation and 30-day archival — a vanished process still leaves clues behind.

check

Preflight checks

--dry-run prints the effective routing, the rendered injection table and the model-resolution chain with no network I/O. --doctor adds DNS/TCP/TLS reachability and listen-port checks, and exits non-zero on problems.

deps

Zero dependencies

Built-in Node modules only, no third-party packages, and nothing to install before npx starts.

Quick start

Requires Node.js ≥ 18.

# The default config already targets OpenCode Go — just start it
npx llm-session-proxy

# Client Base URL:  http://127.0.0.1:9355/zen/go/v1
# Model ID:         proxy- + the real model name, e.g. proxy-glm-5.3-flash
# API Key:          your OpenCode API key
The proxy- prefix is not optional. Clients like Trae decide whether to use a custom channel based on the model ID. Use a built-in preset name such as glm-5.3-flash and the traffic gets picked up by the client's own cloud channel, bypassing this proxy entirely — and still returning 400. Adding the prefix forces the request through the proxy, which strips the prefix before forwarding upstream.
npx llm-session-proxy \
  --upstream https://api.deepseek.com \
  --port 8788 \
  --inject "x-session-id={{session.id}}" \
  --inject "x-trace-id={{uuid}}" \
  --model-map fast=deepseek-chat \
  --model-map smart=deepseek-reasoner \
  --model-prefix ""

# Client Base URL: http://127.0.0.1:8788/v1
# Model names can be short aliases like fast / smart
# Generate an annotated sample config
npx llm-session-proxy --init

# Edit it, then start
npx llm-session-proxy -c llm-session-proxy.config.json
{
  // supports // and /* */ comments, and trailing commas
  "listen": { "host": "127.0.0.1", "port": 9355 },
  "upstream": { "host": "opencode.ai" },
  "inject": {
    "headers": {
      "x-opencode-session": "{{session.id}}",
      "x-opencode-request": "{{session.requestId}}"
    }
  },
  "model": { "stripPrefixes": ["proxy-"] }
}

What to put in the client's Base URL

If the client can only be givenStart the proxy with
http://127.0.0.1:9355/zen/go/v1no extra flags (passed through as-is)
http://127.0.0.1:9355/v1--path-rewrite "^/v1/=>/zen/go/v1/"
http://127.0.0.1:9355 (no path)--base-path /zen/go/v1

OpenCode Go model cheat sheet

The upstream splits models across three endpoints — whether a model is usable depends on the protocol your client speaks, which matters more than the model name. Every client model ID below resolves out of the box: the aliases are built into the proxy. Curious what any alias turns into? npx llm-session-proxy --dry-run --model proxy-glm.

① /zen/go/v1/chat/completions · OpenAI-compatible, what most clients use

ModelClient model IDReal upstream ID
GLM-5.3proxy-glmglm-5.3
GLM-5.3-Flashproxy-glm-flashglm-5.3-flash
GLM-5.2 / 5.1proxy-glm-5.2glm-5.2
Kimi K3proxy-kimikimi-k3
Kimi K2.7 Codeproxy-kimi-codekimi-k2.7-code
Kimi K2.6proxy-kimi-k2.6kimi-k2.6
DeepSeek V4.1 Flashproxy-deepseekdeepseek-flash
DeepSeek V4 Proproxy-deepseek-prodeepseek-v4-pro
DeepSeek V4 Flashproxy-deepseek-v4-flashdeepseek-v4-flash
DeepSeek V4 Flash Visionproxy-deepseek-visiondeepseek-v4-flash-vision-exp
LongCat-2.0proxy-longcatlongcat-2.0
MiMo-V2.5 / Proproxy-mimomimo-v2.5
Hy3 / Hy4 previewproxy-hy3hy3

② /zen/go/v1/responses · OpenAI Responses API, client must support that protocol

ModelClient model IDReal upstream ID
Grok 4.6proxy-grokgrok-4.6
GPT 5.6 Lunaproxy-gpt-lunagpt-5.6-luna
Muse Spark 1.3 Contributorproxy-muse-1.3muse-spark-1.3-contributor

③ /zen/go/v1/messages · Anthropic protocol, client must be able to speak it

ModelClient model IDReal upstream ID
MiniMax M3 / M2.7 / M2.5proxy-minimaxminimax-m3
Qwen3.8 Maxproxy-qwen-maxqwen3.8-max
Qwen3.8 Flashproxy-qwen-flashqwen3.8-flash
Qwen3.7 / 3.6 Plusproxy-qwen-plusqwen3.6-plus
This proxy does no protocol conversion. When the client sends an OpenAI-shaped body, the models in ② and ③ are out of reach — the proxy only adds headers, renames models and swaps paths; it will not translate a Chat Completions body into a Messages body. Writing the real upstream ID directly (no alias, no prefix) also works: the proxy passes it through untouched. The model list changes over time — treat the upstream docs as the source of truth.

Configuration reference

Precedence: defaults < config file < environment variables < CLI flags.

upstream — where requests are forwarded

FieldDefaultMeaning
protocolhttpshttps or http
hostopencode.aiUpstream hostname; a full URL also works
portnullnull means the protocol default
basePath""Prefix added to every forwarded path, e.g. /zen/go/v1
rewriteHosttrueWhether the Host header is rewritten to the upstream host

session — session IDs

FieldDefaultMeaning
enabledtrueWhether session ID injection is on
headerNamesx-opencode-session, …Read a client-supplied session ID from these headers, in order
bodyFieldssession_id, …Read it from these body fields, in order (dot paths supported)
contentHash.fields["system", …]Fields that feed the content fingerprint
includeFirstUserMessagetrueWhether the first user message feeds the fingerprint
idPrefixses_Prefix for generated IDs
idFormathex26hex26 / hex / uuid / base36 / short
requestIdFormatmsg_{{session.count}}Template for the per-request ID
maxSessions512Session table cap; the oldest entries are evicted first
ttlSeconds0Session expiry in seconds; 0 means never

inject — parameter injection

FieldDefaultMeaning
headersfour x-opencode-*Headers to inject; values are templates. Set to null to skip
body{}Fields merged into the request body; dot paths and templates supported
removeBodyFields[]Body field paths to delete
overwritetruefalse keeps a client-supplied header or field of the same name

model — model name rewriting

FieldDefaultMeaning
stripPrefixes["proxy-"]Prefixes to strip; matched in order, first hit wins
map27 built-in aliasesAlias → real upstream model ID. Merged key by key over the built-in table: a key here overrides the built-in entry of the same name, the rest are kept. Built-in entries cannot be deleted — override them instead
defaultnullFallback model name. Only applies when nothing matched and no prefix was stripped; it deliberately does not rescue proxy-xxx that strips to an unmapped xxx
warnUnmappedtrueWarn once per alias that strips to a name with no mapping and no default, naming the entry to add

transformers — body transforms, always on

Names only: the registry is a fixed list compiled into the proxy, so a config file can never make it execute code.

FieldDefaultMeaning
enabled[]Transformer names, applied left to right
options{}Per-transformer options, keyed by transformer name
NameOptionsEffect
noop—Changes nothing. Useful for confirming the registry is wired up
drop-fieldsfields: ["a.b"]Delete the listed dot paths
drop-empty-fieldsfields?: [...]Delete fields that are null, "", [] or {}. 0 and false survive. Without fields, every top-level key is inspected
rename-fieldsmap: {"from": "to"}Rename dot paths, creating parent objects as needed. A mapping onto itself is ignored
clamp-max-tokensmax: 4096Lower max_tokens / max_completion_tokens to at most max. Never raises a value

Order matters, because each transformer mutates the body in place: rename-fields before drop-fields is not the same as the reverse. A name that is not in the table is a startup error, not a silent no-op.

router — buckets and rules

A bucket decides two things: which model to use and which transforms to attach. enabled is false by default, so an existing config behaves exactly as before until you turn it on.

FieldDefaultMeaning
enabledfalseMaster switch. --no-router forces it off
forcednullPin every request to this bucket, ignoring all rules (--router <bucket>). It outranks enabled: false as well — naming a bucket explicitly should not be silently ignored
defaultBucketdefaultWhere requests that match no rule go
bucketsthe four built-ins, all empty{ "model": <id or null>, "transformers": [<name>] }. Custom names may be added; an undeclared name degrades to an empty bucket
rules[]Evaluated top-down, first match wins. A rule needs a bucket plus at least one matcher — zero matchers is a config error, since a rule that cannot be evaluated is not the same as one that matches everything
MatcherMatches when
pathThe request path starts with this string. Matched against the path the client sent, before request.pathRewrite
modelPrefixEither the model the client sent or the model after rewriting starts with this, so proxy-think and glm-5.3 are both usable as rule material
bodyFieldThat dot path exists and is non-empty. Add bodyFieldValue to require one exact value instead
minBytes / maxBytesRequest body size, inclusive on both ends. Bytes, not tokens — a token estimate would mean shipping a tokenizer, and a byte count is a figure you can actually tune
"router": {
  "enabled": true,
  "defaultBucket": "default",
  "buckets": {
    "think":       { "model": "glm-5.3-think" },
    "longContext": { "model": "glm-5.3-long" },
    "background":  { "model": "glm-5.3-flash", "transformers": ["clamp-max-tokens"] }
  },
  "rules": [
    { "bucket": "think",       "path": "/zen/go/v1/messages" },
    { "bucket": "longContext", "minBytes": 60000 },
    { "bucket": "background",  "modelPrefix": "proxy-haiku", "maxBytes": 4096 }
  ]
}

A bucket's transformers are appended to the global transformers.enabled list — a bucket can add transforms but cannot cancel a global one. A bucket's model is applied after alias rewriting and replaces whatever was there, so it should be a real upstream ID rather than a proxy- alias. {{model}} in an injected header sees the bucket's value, because injection runs last.

protocol — convert between protocols

The proxy detects the client's protocol from the request path (/chat/completions → chat, /messages → messages, /responses → responses) and a route decides what the upstream speaks. When the two differ, request bodies are converted before forwarding and responses converted back — JSON as a whole, SSE event by event. enabled is false by default.

FieldDefaultMeaning
enabledfalseMaster switch. --no-protocol forces it off
forcednullConvert every request to this protocol, ignoring all routes (--protocol <target>). Outranks enabled: false as well
pathsthe three /zen/go/v1/… pathsThe upstream path used per target protocol. paths.<target> replaces the forwarded path as a whole, unlike request.pathRewrite which patches the client's path
routes[]{ "model": <prefix, optional>, "target": <protocol>, "path": <upstream path, optional> }. Top-down, first match wins; a route without model matches every request
"protocol": {
  "enabled": true,
  "paths": {
    "chat": "/zen/go/v1/chat/completions",
    "messages": "/zen/go/v1/messages",
    "responses": "/zen/go/v1/responses"
  },
  "routes": [
    { "model": "minimax", "target": "messages" },
    { "model": "gpt", "target": "responses" }
  ]
}

Field normalisation is built in: max_tokens ↔ max_completion_tokens ↔ max_output_tokens; reasoning_effort ↔ thinking.budget_tokens (fixed table 1024 / 8192 / 16384, back by threshold); tool declarations and tool calls reshaped between the three formats; stop ↔ stop_sequences; chat's system message ↔ Anthropic's top-level system ↔ Responses' instructions. Fields with no equivalent are dropped by name into a logged dropped list. Anthropic requires max_tokens, so a chat request without either name gets a conservative, logged 4096 instead of a 400. Upstream 4xx/5xx error bodies are never converted — they belong to the upstream's protocol and pass through verbatim.

Template variables

VariableMeaning
{{session.id}}Session ID used for this request
{{session.count}}Request number within the session (1-based)
{{session.requestId}}Rendered from requestIdFormat, e.g. msg_3
{{model}} / {{path}} / {{method}}Rewritten model name / original path / method
{{header.x-foo}} / {{query.foo}}Client request header (lowercased) / query parameter
{{env.HOME}}Environment variable
{{uuid}} / {{random}} / {{randomHex:16}}Different on every render
{{timestamp}} / {{timestampMs}}Seconds / milliseconds since epoch

CLI reference

Every flag mixes with a config file; CLI flags win.

FlagMeaning
-c, --config <file>Read a config file (comments and trailing commas allowed)
-p, --port <n> / --host <addr>Listen port / address
-u, --upstream <url>Upstream address, e.g. https://opencode.ai or host:port
--base-path <path>Forwarding path prefix
--path-rewrite <a=>b>Path rewrite (regex), repeatable
--inject <name=value>Inject a request header, repeatable
--body-inject <k=v>Inject a body field, dot paths supported, repeatable
--model-prefix <prefix>Model name prefix to strip, repeatable
--model-map <a=b>Exact model name mapping, repeatable. Merged over the built-in alias table
--transformer <name>Attach a named body transform, repeatable. Replaces transformers.enabled rather than appending to it, the same way --model-prefix replaces stripPrefixes
--router <bucket>Force every request through one bucket, ignoring the rules
--no-routerTurn router buckets off, even if the config file enables them
--protocol <target>Convert every request to this protocol before forwarding (chat / messages / responses), overriding protocol.routes
--no-protocolTurn protocol conversion off, even if the config file enables it
--session-header <name> / --session-field <path>Append a session source location, repeatable
--session-id-format <f> / --request-id-format <t>Session ID format / per-request ID template
--no-session / --no-streamDisable session injection / disable streaming pass-through
--timeout <ms> / --max-body <bytes>Upstream timeout / request body limit
--log-level <l>Log level: silent | error | warn | info | debug
-l, --lang <en|zh>Language of console and log output (also PROXY_LANG, or the lang config key)
--log-file <f> / --no-log-fileLog file path / turn file logging off (console only)
--log-dir <dir>Directory of the default log file
--log-rotate <mode>size | daily | off
--log-keep-days <n>Delete logs older than N days (0 = keep forever)
--init [file]Write a sample config
--print-configPrint the merged config and exit
--dry-runValidate the config and print routing, injection and model resolution. No network I/O
--doctor--dry-run plus DNS/TCP/TLS reachability and listen-port checks; exits non-zero on problems
--model <id>Sample model name used by --dry-run / --doctor

Logs

Logging to disk is on by default, because the failures that are hardest to diagnose — an uncaught exception, a silent crash — leave nothing behind on stderr once the terminal is gone. The file is $LSP_HOME/logs/llm-session-proxy.log, or ~/.lsp/logs/llm-session-proxy.log when LSP_HOME is unset. Override it with --log-file <path> / --log-dir <dir>, or switch it off with --no-log-file.

log.rotateBehaviour
size (default)app.log grows to log.maxBytes, then becomes app.log.1, .2 … up to log.backups
dailyOne file per local date: app-YYYY-MM-DD.log. log.maxBytes still caps a single day's file
offNever rotate or truncate — hand the file to logrotate or similar

Archival. On startup, and at most once every six hours while running, rotated files (app.log, app.log.N, app-YYYY-MM-DD.log) older than log.keepDays — 30 days by default — are deleted. The file currently being written is never touched, and files that do not match that naming pattern, including anything else in the same directory, are left alone.

# Where will logs actually land?
llm-session-proxy --print-config | grep resolvedFile

Preflight: --dry-run and --doctor

Both flags validate the config and print what the proxy would actually do, then exit — no server, no log file, no request to the upstream. The injection table is not a description of the templates; it is the rendered result, so a misspelled template shows up here instead of in a 400 from the upstream.

$ llm-session-proxy --dry-run
llm-session-proxy v0.2.2 — dry run

Model
  sample              proxy-glm
  strip               prefix "proxy-" -> glm
  mapped              glm -> glm-5.3
  result              glm-5.3  (mapped)
  map                 27 built-in aliases, 0 overrides

Router
  enabled             no
  default bucket      default
  bucket default      (none)
  bucket background   (none)
  bucket think        (none)
  bucket longContext  (none)
  rules               (no rules)
  sample route        default (router disabled)

Transformers
  global              (none)
  effective           (none)
  available           noop, drop-fields, drop-empty-fields, rename-fields, clamp-max-tokens

Protocol
  enabled             no (pass-through)
  paths               chat=/zen/go/v1/chat/completions, messages=/zen/go/v1/messages, responses=/zen/go/v1/responses
  rules               (no rules)
  sample conversion   no conversion for the sample request

Result
  OK — the configuration is valid.
Model / resultMeaning
(mapped)The alias hit model.map. This is what you want
(from model.default)Nothing matched and nothing was stripped, so default applied
(prefix stripped, NO mapping)The prefix came off and the remainder is going upstream verbatim — the usual cause of "model not found"
(no prefix matched, forwarded as-is)The client sent a real model ID. Normal

The Router and Transformers blocks answer "which bucket would this request land in, and what would be done to it". For the config shown under router above:

Router
  enabled             yes
  default bucket      default
  bucket default      (none)
  bucket background   model=glm-5.3-flash transformers=clamp-max-tokens
  bucket think        model=glm-5.3-think
  bucket longContext  model=glm-5.3-long
  #0                  path^=/zen/go/v1/messages -> think
  #1                  bytes>=60000 -> longContext
  #2                  model~=proxy-haiku* AND bytes<=4096 -> background
  sample route        default  (default bucket, no rule matched /v1/chat/completions)

Transformers
  global              drop-empty-fields
  effective           drop-empty-fields
  available           noop, drop-fields, drop-empty-fields, rename-fields, clamp-max-tokens

#0/#1/#2 are the rules in order, with path^= meaning "path starts with" and model~= meaning "model prefix". The sample route row is the result of actually running the matcher over a sample POST /v1/chat/completions, so it lands on the fallback: nothing matched and the request goes to the default bucket. --router think overrides the whole thing: sample route think (forced by --router). effective is what will really run — the global list first, then whatever the winning bucket appends. In Chinese output these two blocks are titled 路由分桶 and 变换, deliberately distinct from the pre-existing 路由 (path rewriting) section.

The Protocol block answers "would this request be converted, and to what". With --protocol messages it shows the forced target and, on the sample conversion row, the direction plus the upstream path the converted request takes (chat -> messages /zen/go/v1/messages). In Chinese output the section is titled 协议互转.

--doctor runs the same report and adds DNS resolution, a TCP connect with handshake time, a TLS handshake when the protocol is https (an untrusted certificate is reported, not treated as a hard failure), a listen-port availability check, and a reminder that the proxy never injects credentials. It sends no HTTP request and no credentials — reachability is answered at the transport layer — so a doctor run cannot burn rate-limit quota. It exits 0 when everything is fine and 1 when something needs fixing, which makes it usable as a startup gate:

llm-session-proxy --doctor && llm-session-proxy

Local status endpoints

# Overview: session count, hit rate, injected headers, upstream target
curl http://127.0.0.1:9355/__llm_session_proxy__/status

# Per-session detail: ID, request count, last used
curl http://127.0.0.1:9355/__llm_session_proxy__/sessions

Handy when debugging: if sessions.active stays at 0, requests are not reaching the proxy at all — usually the model ID is missing its proxy- prefix and the client intercepted it.

How it works

Each request goes through a fixed sequence; the three-tier session strategy is the key part.

01 READRead the body and parse the JSON
→
02 SESSIONexplicit > fingerprint > random
→
03 REWRITEmodel name: strip, then map
→
04 ROUTEbucket model + transforms
→
05 INJECTheaders and body fields
→
06 CONVERTtranslate to the upstream protocol
→
07 FORWARDstream SSE chunk by chunk

The order of 03–06 is deliberate. Rules see the model name the client sent and the resolved one; the bucket's model override lands after alias resolution, so it can only be a real upstream ID; injection comes last-but-one, so {{model}} in a header reflects the final decision; and protocol conversion runs last of all — transforms and injection work on the client's field names, and only the finished body is translated.

The three-tier session strategy

Explicit identifier — the client sent x-opencode-session itself, or the body carries session_id. Reused as-is; this is the most accurate path.

Content fingerprint — SHA-256 over system plus the first user message. Later turns in the same conversation only append messages, so the first one never changes, the fingerprint stays stable, and the turn lands back on the same ID.

One-off random — when neither is available (a non-JSON body, say), a random ID is sent so the upstream does not return 400. These requests are kept out of the session table so it cannot grow without bound.

Stability

A single malformed request cannot take the process down. An invalid Host header, a malformed request line, or an upstream status line / header with illegal characters all become ordinary 4xx / 5xx responses, and the proxy keeps serving.

Uncaught errors land in the log file. A last-resort handler writes the full stack of uncaughtException and unhandledRejection into the log file. Node's default prints to stderr and terminates immediately, leaving the log file empty — which looks exactly like "logs are fine, the process just vanished". That is why file logging is on by default: see Logs.

Limitations

Loopback only. The proxy forwards your API key, so do not bind the listen address to 0.0.0.0.

response.stream: false breaks SSE. Keep the default true unless you genuinely want whole-response buffering.

With bufferBody: false the body cannot be rewritten — model rewriting and parameter injection stop working, and only headers can be injected.

The built-in model map is a snapshot, not a live catalogue. Upstream model IDs change without notice, and this proxy deliberately makes no network call to discover them. If an alias stops resolving, it is a one-line override in your config file — not a bug — and --dry-run shows which alias resolved to what before you send a real request.

Protocol conversion is deliberately lossy at the edges. Anthropic thinking / signature stream deltas have no chat equivalent and are dropped (by name, in the log); messages ↔ responses composes through chat rather than a direct converter. What survives is text, tool calls, finish reasons and usage — the parts clients actually act on. The transform registry remains a fixed list: the proxy never loads a transform from a path.

Roadmap

Zero dependencies, local, single process — every item has to fit that shape. Full reasoning, the comparison with comparable gateways and the explicit non-goals live in ROADMAP.md.

MilestoneThemeHighlights
v0.2.0 ✓CorrectnessA curated built-in model map (27 aliases), a warn-once notice when a proxy--prefixed alias resolves to nothing, and the --dry-run / --doctor preflight. The release that fixes “my alias silently reached the upstream verbatim”.
v0.2.1 ✓RoutingFive named in-tree transforms (noop, drop-fields, drop-empty-fields, rename-fields, clamp-max-tokens) and router buckets (default / background / think / longContext) with top-down first-match rules over path, model prefix, body field and request size. Off by default; the host the translators plug into. The load-a-plugin-from-a-path model originally sketched for this milestone was dropped on purpose.
v0.2.2 ✓Protocol reachAnthropic ↔ OpenAI ↔ Responses translation with field normalisation, streaming included — request bodies, JSON responses and event-by-event SSE transcoding. The release that stops a client’s protocol from deciding which models it can reach.
v0.3Observability and controlPrometheus /metrics, JSON log mode, OTLP/JSON export, a session and prompt-cache inspector, token and cost accounting, and a zero-build local dashboard.
v0.4Reliability under real upstreamsMultiple upstreams, fallback with jittered retry, circuit breaking, separated timeouts with a stream idle watchdog, graceful drain on shutdown.
v0.5Safety and hygieneLog redaction, regex guardrails, line-precise config schema errors, and --check-config for CI.
v1.0Distribution and guaranteesSingle-file binaries, a per-client compatibility matrix, protocol contract tests on golden SSE fixtures, a config migration tool, and semver from there on.

The v0.2 performance targets are measured by a harness that ships inside the repo: added latency of ≤ 1 ms p50 and ≤ 5 ms p99 against a local mock upstream, ≥ 1,200 RPS at 100 concurrent connections, zero process exits under the malformed-input corpus, and byte-for-byte SSE pass-through. Numbers get published together with their methodology — without it a benchmark says nothing.

Deliberately not on the roadmap: virtual keys and billing, running models locally, a hosted version, a plugin marketplace, a Kubernetes operator. Each of them pulls against “zero dependencies, local, single process”.