llm-session-proxy
A configurable local reverse proxy for LLM APIs. It fills in the session header your client never sends, injects arbitrary custom parameters, rewrites model aliases and request paths, and can convert between the OpenAI, Anthropic and Responses protocols — then forwards the request, streaming SSE responses included.
npx llm-session-proxy
The problem it solves
OpenCode Go wants a session identifier on every request, and most clients cannot set custom headers.
The upstream expects a header like this:
x-opencode-session : stable per conversation (used for prompt caching and routing)
If the client does not send it, the upstream refuses the call outright:
400 Request is missing x-opencode-session and cannot be routed efficiently
This tool moves that job onto your machine: point the client's Base URL at the local proxy
and the proxy handles the rest. Nothing in it is tied to one provider — swap
upstream and inject.headers and it fronts any OpenAI- or
Anthropic-compatible endpoint.
What it does
Every item is adjustable through a config file, an environment variable, or a CLI flag.
Automatic session IDs
Three-tier strategy: explicit client identifier → content fingerprint (system + first user message) → one-off random. Stable within a conversation, so prompt caching actually works.
Arbitrary parameter injection
Inject request headers and request body fields. Values support templates such as {{session.id}}, {{uuid}} and {{env.HOME}}.
Model alias rewriting
Prefix stripping (proxy-glm → glm-5.3) and exact mapping, and the two compose. The OpenCode Go alias table ships built in, so the proxy- prefix the client needs actually resolves — and an alias that strips down to nothing is reported by name instead of being forwarded silently.
Rule-based routing into buckets
Four buckets — default / background / think / longContext — each with its own model override and body transforms. Rules match on path prefix, model prefix, a body field or request size; they run top-down with first match winning, and the conditions inside one rule are ANDed. Off by default, so upgrading changes nothing.
Composable body transforms
Five named, in-tree transforms (drop-fields, drop-empty-fields, rename-fields, clamp-max-tokens, noop), applied globally or per bucket in a fixed order. The registry is a hard-coded list on purpose: the proxy never loads code from a path.
Protocol conversion
OpenAI Chat ↔ Anthropic Messages ↔ OpenAI Responses, for request bodies, JSON responses and event-by-event SSE transcoding. A client that speaks one protocol reaches models that speak another. Routed by model prefix, off by default.
Request path rewriting
When the client can only be pointed at /v1, one regex moves it to the path the upstream actually wants.
Unbuffered SSE pass-through
Chunks are forwarded as they arrive instead of being buffered, so streaming stays streaming. Upstream 4xx bodies come back verbatim.
Logs on by default
~/.lsp/logs/llm-session-proxy.log with size or date rotation and 30-day archival — a vanished process still leaves clues behind.
Preflight checks
--dry-run prints the effective routing, the rendered injection table and the model-resolution chain with no network I/O. --doctor adds DNS/TCP/TLS reachability and listen-port checks, and exits non-zero on problems.
Zero dependencies
Built-in Node modules only, no third-party packages, and nothing to install before npx starts.
Quick start
Requires Node.js ≥ 18.
# The default config already targets OpenCode Go — just start it
npx llm-session-proxy
# Client Base URL: http://127.0.0.1:9355/zen/go/v1
# Model ID: proxy- + the real model name, e.g. proxy-glm-5.3-flash
# API Key: your OpenCode API key
proxy- prefix is not optional.
Clients like Trae decide whether to use a custom channel based on the model ID. Use a
built-in preset name such as glm-5.3-flash and the traffic gets picked up by
the client's own cloud channel, bypassing this proxy entirely — and still returning
400. Adding the prefix forces the request through the proxy, which strips the prefix before
forwarding upstream.
npx llm-session-proxy \
--upstream https://api.deepseek.com \
--port 8788 \
--inject "x-session-id={{session.id}}" \
--inject "x-trace-id={{uuid}}" \
--model-map fast=deepseek-chat \
--model-map smart=deepseek-reasoner \
--model-prefix ""
# Client Base URL: http://127.0.0.1:8788/v1
# Model names can be short aliases like fast / smart
# Generate an annotated sample config
npx llm-session-proxy --init
# Edit it, then start
npx llm-session-proxy -c llm-session-proxy.config.json
{
// supports // and /* */ comments, and trailing commas
"listen": { "host": "127.0.0.1", "port": 9355 },
"upstream": { "host": "opencode.ai" },
"inject": {
"headers": {
"x-opencode-session": "{{session.id}}",
"x-opencode-request": "{{session.requestId}}"
}
},
"model": { "stripPrefixes": ["proxy-"] }
}
What to put in the client's Base URL
| If the client can only be given | Start the proxy with |
|---|---|
http://127.0.0.1:9355/zen/go/v1 | no extra flags (passed through as-is) |
http://127.0.0.1:9355/v1 | --path-rewrite "^/v1/=>/zen/go/v1/" |
http://127.0.0.1:9355 (no path) | --base-path /zen/go/v1 |
OpenCode Go model cheat sheet
The upstream splits models across three endpoints — whether a model is usable depends on the
protocol your client speaks, which matters more than the model name. Every client model ID
below resolves out of the box: the aliases are built into the proxy. Curious what any alias turns
into? npx llm-session-proxy --dry-run --model proxy-glm.
① /zen/go/v1/chat/completions · OpenAI-compatible, what most clients use
| Model | Client model ID | Real upstream ID |
|---|---|---|
| GLM-5.3 | proxy-glm | glm-5.3 |
| GLM-5.3-Flash | proxy-glm-flash | glm-5.3-flash |
| GLM-5.2 / 5.1 | proxy-glm-5.2 | glm-5.2 |
| Kimi K3 | proxy-kimi | kimi-k3 |
| Kimi K2.7 Code | proxy-kimi-code | kimi-k2.7-code |
| Kimi K2.6 | proxy-kimi-k2.6 | kimi-k2.6 |
| DeepSeek V4.1 Flash | proxy-deepseek | deepseek-flash |
| DeepSeek V4 Pro | proxy-deepseek-pro | deepseek-v4-pro |
| DeepSeek V4 Flash | proxy-deepseek-v4-flash | deepseek-v4-flash |
| DeepSeek V4 Flash Vision | proxy-deepseek-vision | deepseek-v4-flash-vision-exp |
| LongCat-2.0 | proxy-longcat | longcat-2.0 |
| MiMo-V2.5 / Pro | proxy-mimo | mimo-v2.5 |
| Hy3 / Hy4 preview | proxy-hy3 | hy3 |
② /zen/go/v1/responses · OpenAI Responses API, client must support that protocol
| Model | Client model ID | Real upstream ID |
|---|---|---|
| Grok 4.6 | proxy-grok | grok-4.6 |
| GPT 5.6 Luna | proxy-gpt-luna | gpt-5.6-luna |
| Muse Spark 1.3 Contributor | proxy-muse-1.3 | muse-spark-1.3-contributor |
③ /zen/go/v1/messages · Anthropic protocol, client must be able to speak it
| Model | Client model ID | Real upstream ID |
|---|---|---|
| MiniMax M3 / M2.7 / M2.5 | proxy-minimax | minimax-m3 |
| Qwen3.8 Max | proxy-qwen-max | qwen3.8-max |
| Qwen3.8 Flash | proxy-qwen-flash | qwen3.8-flash |
| Qwen3.7 / 3.6 Plus | proxy-qwen-plus | qwen3.6-plus |
Configuration reference
Precedence: defaults < config file < environment variables < CLI flags.
upstream — where requests are forwarded
| Field | Default | Meaning |
|---|---|---|
protocol | https | https or http |
host | opencode.ai | Upstream hostname; a full URL also works |
port | null | null means the protocol default |
basePath | "" | Prefix added to every forwarded path, e.g. /zen/go/v1 |
rewriteHost | true | Whether the Host header is rewritten to the upstream host |
session — session IDs
| Field | Default | Meaning |
|---|---|---|
enabled | true | Whether session ID injection is on |
headerNames | x-opencode-session, … | Read a client-supplied session ID from these headers, in order |
bodyFields | session_id, … | Read it from these body fields, in order (dot paths supported) |
contentHash.fields | ["system", …] | Fields that feed the content fingerprint |
includeFirstUserMessage | true | Whether the first user message feeds the fingerprint |
idPrefix | ses_ | Prefix for generated IDs |
idFormat | hex26 | hex26 / hex / uuid / base36 / short |
requestIdFormat | msg_{{session.count}} | Template for the per-request ID |
maxSessions | 512 | Session table cap; the oldest entries are evicted first |
ttlSeconds | 0 | Session expiry in seconds; 0 means never |
inject — parameter injection
| Field | Default | Meaning |
|---|---|---|
headers | four x-opencode-* | Headers to inject; values are templates. Set to null to skip |
body | {} | Fields merged into the request body; dot paths and templates supported |
removeBodyFields | [] | Body field paths to delete |
overwrite | true | false keeps a client-supplied header or field of the same name |
model — model name rewriting
| Field | Default | Meaning |
|---|---|---|
stripPrefixes | ["proxy-"] | Prefixes to strip; matched in order, first hit wins |
map | 27 built-in aliases | Alias → real upstream model ID. Merged key by key over the built-in table: a key here overrides the built-in entry of the same name, the rest are kept. Built-in entries cannot be deleted — override them instead |
default | null | Fallback model name. Only applies when nothing matched and no prefix was stripped; it deliberately does not rescue proxy-xxx that strips to an unmapped xxx |
warnUnmapped | true | Warn once per alias that strips to a name with no mapping and no default, naming the entry to add |
transformers — body transforms, always on
Names only: the registry is a fixed list compiled into the proxy, so a config file can never make it execute code.
| Field | Default | Meaning |
|---|---|---|
enabled | [] | Transformer names, applied left to right |
options | {} | Per-transformer options, keyed by transformer name |
| Name | Options | Effect |
|---|---|---|
noop | — | Changes nothing. Useful for confirming the registry is wired up |
drop-fields | fields: ["a.b"] | Delete the listed dot paths |
drop-empty-fields | fields?: [...] | Delete fields that are null, "", [] or {}. 0 and false survive. Without fields, every top-level key is inspected |
rename-fields | map: {"from": "to"} | Rename dot paths, creating parent objects as needed. A mapping onto itself is ignored |
clamp-max-tokens | max: 4096 | Lower max_tokens / max_completion_tokens to at most max. Never raises a value |
Order matters, because each transformer mutates the body in place: rename-fields
before drop-fields is not the same as the reverse. A name that is not in the table is
a startup error, not a silent no-op.
router — buckets and rules
A bucket decides two things: which model to use and which transforms to attach. enabled is false by default, so an existing config behaves exactly as before until you turn it on.
| Field | Default | Meaning |
|---|---|---|
enabled | false | Master switch. --no-router forces it off |
forced | null | Pin every request to this bucket, ignoring all rules (--router <bucket>). It outranks enabled: false as well — naming a bucket explicitly should not be silently ignored |
defaultBucket | default | Where requests that match no rule go |
buckets | the four built-ins, all empty | { "model": <id or null>, "transformers": [<name>] }. Custom names may be added; an undeclared name degrades to an empty bucket |
rules | [] | Evaluated top-down, first match wins. A rule needs a bucket plus at least one matcher — zero matchers is a config error, since a rule that cannot be evaluated is not the same as one that matches everything |
| Matcher | Matches when |
|---|---|
path | The request path starts with this string. Matched against the path the client sent, before request.pathRewrite |
modelPrefix | Either the model the client sent or the model after rewriting starts with this, so proxy-think and glm-5.3 are both usable as rule material |
bodyField | That dot path exists and is non-empty. Add bodyFieldValue to require one exact value instead |
minBytes / maxBytes | Request body size, inclusive on both ends. Bytes, not tokens — a token estimate would mean shipping a tokenizer, and a byte count is a figure you can actually tune |
"router": {
"enabled": true,
"defaultBucket": "default",
"buckets": {
"think": { "model": "glm-5.3-think" },
"longContext": { "model": "glm-5.3-long" },
"background": { "model": "glm-5.3-flash", "transformers": ["clamp-max-tokens"] }
},
"rules": [
{ "bucket": "think", "path": "/zen/go/v1/messages" },
{ "bucket": "longContext", "minBytes": 60000 },
{ "bucket": "background", "modelPrefix": "proxy-haiku", "maxBytes": 4096 }
]
}
A bucket's transformers are appended to the global
transformers.enabled list — a bucket can add transforms but cannot cancel a global
one. A bucket's model is applied after alias rewriting and replaces whatever was
there, so it should be a real upstream ID rather than a proxy- alias.
{{model}} in an injected header sees the bucket's value, because injection runs last.
protocol — convert between protocols
The proxy detects the client's protocol from the request path (/chat/completions → chat, /messages → messages, /responses → responses) and a route decides what the upstream speaks. When the two differ, request bodies are converted before forwarding and responses converted back — JSON as a whole, SSE event by event. enabled is false by default.
| Field | Default | Meaning |
|---|---|---|
enabled | false | Master switch. --no-protocol forces it off |
forced | null | Convert every request to this protocol, ignoring all routes (--protocol <target>). Outranks enabled: false as well |
paths | the three /zen/go/v1/… paths | The upstream path used per target protocol. paths.<target> replaces the forwarded path as a whole, unlike request.pathRewrite which patches the client's path |
routes | [] | { "model": <prefix, optional>, "target": <protocol>, "path": <upstream path, optional> }. Top-down, first match wins; a route without model matches every request |
"protocol": {
"enabled": true,
"paths": {
"chat": "/zen/go/v1/chat/completions",
"messages": "/zen/go/v1/messages",
"responses": "/zen/go/v1/responses"
},
"routes": [
{ "model": "minimax", "target": "messages" },
{ "model": "gpt", "target": "responses" }
]
}
Field normalisation is built in: max_tokens ↔ max_completion_tokens ↔
max_output_tokens; reasoning_effort ↔ thinking.budget_tokens
(fixed table 1024 / 8192 / 16384, back by threshold); tool declarations and tool calls reshaped
between the three formats; stop ↔ stop_sequences; chat's system message ↔
Anthropic's top-level system ↔ Responses' instructions. Fields with no
equivalent are dropped by name into a logged dropped list. Anthropic requires
max_tokens, so a chat request without either name gets a conservative, logged 4096
instead of a 400. Upstream 4xx/5xx error bodies are never converted — they belong to the
upstream's protocol and pass through verbatim.
Template variables
| Variable | Meaning |
|---|---|
{{session.id}} | Session ID used for this request |
{{session.count}} | Request number within the session (1-based) |
{{session.requestId}} | Rendered from requestIdFormat, e.g. msg_3 |
{{model}} / {{path}} / {{method}} | Rewritten model name / original path / method |
{{header.x-foo}} / {{query.foo}} | Client request header (lowercased) / query parameter |
{{env.HOME}} | Environment variable |
{{uuid}} / {{random}} / {{randomHex:16}} | Different on every render |
{{timestamp}} / {{timestampMs}} | Seconds / milliseconds since epoch |
CLI reference
Every flag mixes with a config file; CLI flags win.
| Flag | Meaning |
|---|---|
-c, --config <file> | Read a config file (comments and trailing commas allowed) |
-p, --port <n> / --host <addr> | Listen port / address |
-u, --upstream <url> | Upstream address, e.g. https://opencode.ai or host:port |
--base-path <path> | Forwarding path prefix |
--path-rewrite <a=>b> | Path rewrite (regex), repeatable |
--inject <name=value> | Inject a request header, repeatable |
--body-inject <k=v> | Inject a body field, dot paths supported, repeatable |
--model-prefix <prefix> | Model name prefix to strip, repeatable |
--model-map <a=b> | Exact model name mapping, repeatable. Merged over the built-in alias table |
--transformer <name> | Attach a named body transform, repeatable. Replaces transformers.enabled rather than appending to it, the same way --model-prefix replaces stripPrefixes |
--router <bucket> | Force every request through one bucket, ignoring the rules |
--no-router | Turn router buckets off, even if the config file enables them |
--protocol <target> | Convert every request to this protocol before forwarding (chat / messages / responses), overriding protocol.routes |
--no-protocol | Turn protocol conversion off, even if the config file enables it |
--session-header <name> / --session-field <path> | Append a session source location, repeatable |
--session-id-format <f> / --request-id-format <t> | Session ID format / per-request ID template |
--no-session / --no-stream | Disable session injection / disable streaming pass-through |
--timeout <ms> / --max-body <bytes> | Upstream timeout / request body limit |
--log-level <l> | Log level: silent | error | warn | info | debug |
-l, --lang <en|zh> | Language of console and log output (also PROXY_LANG, or the lang config key) |
--log-file <f> / --no-log-file | Log file path / turn file logging off (console only) |
--log-dir <dir> | Directory of the default log file |
--log-rotate <mode> | size | daily | off |
--log-keep-days <n> | Delete logs older than N days (0 = keep forever) |
--init [file] | Write a sample config |
--print-config | Print the merged config and exit |
--dry-run | Validate the config and print routing, injection and model resolution. No network I/O |
--doctor | --dry-run plus DNS/TCP/TLS reachability and listen-port checks; exits non-zero on problems |
--model <id> | Sample model name used by --dry-run / --doctor |
Logs
Logging to disk is on by default, because the failures that are hardest to diagnose — an
uncaught exception, a silent crash — leave nothing behind on stderr once the terminal is gone.
The file is $LSP_HOME/logs/llm-session-proxy.log, or
~/.lsp/logs/llm-session-proxy.log when LSP_HOME is unset.
Override it with --log-file <path> / --log-dir <dir>, or switch
it off with --no-log-file.
log.rotate | Behaviour |
|---|---|
size (default) | app.log grows to log.maxBytes, then becomes app.log.1, .2 … up to log.backups |
daily | One file per local date: app-YYYY-MM-DD.log. log.maxBytes still caps a single day's file |
off | Never rotate or truncate — hand the file to logrotate or similar |
Archival. On startup, and at most once every six hours while running, rotated files
(app.log, app.log.N, app-YYYY-MM-DD.log) older than
log.keepDays — 30 days by default — are deleted. The file currently being
written is never touched, and files that do not match that naming pattern, including anything
else in the same directory, are left alone.
# Where will logs actually land?
llm-session-proxy --print-config | grep resolvedFile
Preflight: --dry-run and --doctor
Both flags validate the config and print what the proxy would actually do, then exit — no server, no log file, no request to the upstream. The injection table is not a description of the templates; it is the rendered result, so a misspelled template shows up here instead of in a 400 from the upstream.
$ llm-session-proxy --dry-run
llm-session-proxy v0.2.2 — dry run
Model
sample proxy-glm
strip prefix "proxy-" -> glm
mapped glm -> glm-5.3
result glm-5.3 (mapped)
map 27 built-in aliases, 0 overrides
Router
enabled no
default bucket default
bucket default (none)
bucket background (none)
bucket think (none)
bucket longContext (none)
rules (no rules)
sample route default (router disabled)
Transformers
global (none)
effective (none)
available noop, drop-fields, drop-empty-fields, rename-fields, clamp-max-tokens
Protocol
enabled no (pass-through)
paths chat=/zen/go/v1/chat/completions, messages=/zen/go/v1/messages, responses=/zen/go/v1/responses
rules (no rules)
sample conversion no conversion for the sample request
Result
OK — the configuration is valid.
Model / result | Meaning |
|---|---|
(mapped) | The alias hit model.map. This is what you want |
(from model.default) | Nothing matched and nothing was stripped, so default applied |
(prefix stripped, NO mapping) | The prefix came off and the remainder is going upstream verbatim — the usual cause of "model not found" |
(no prefix matched, forwarded as-is) | The client sent a real model ID. Normal |
The Router and Transformers blocks answer "which bucket would this
request land in, and what would be done to it". For the config shown under
router above:
Router
enabled yes
default bucket default
bucket default (none)
bucket background model=glm-5.3-flash transformers=clamp-max-tokens
bucket think model=glm-5.3-think
bucket longContext model=glm-5.3-long
#0 path^=/zen/go/v1/messages -> think
#1 bytes>=60000 -> longContext
#2 model~=proxy-haiku* AND bytes<=4096 -> background
sample route default (default bucket, no rule matched /v1/chat/completions)
Transformers
global drop-empty-fields
effective drop-empty-fields
available noop, drop-fields, drop-empty-fields, rename-fields, clamp-max-tokens
#0/#1/#2 are the rules in order, with path^=
meaning "path starts with" and model~= meaning "model prefix". The
sample route row is the result of actually running the matcher over a sample
POST /v1/chat/completions, so it lands on the fallback: nothing matched and the
request goes to the default bucket. --router think overrides the whole thing:
sample route think (forced by --router). effective is what will
really run — the global list first, then whatever the winning bucket appends. In Chinese output
these two blocks are titled 路由分桶 and 变换, deliberately distinct
from the pre-existing 路由 (path rewriting) section.
The Protocol block answers "would this request be converted, and to what". With
--protocol messages it shows the forced target and, on the sample conversion
row, the direction plus the upstream path the converted request takes
(chat -> messages /zen/go/v1/messages). In Chinese output the section is titled
协议互转.
--doctor runs the same report and adds DNS resolution, a TCP connect with handshake
time, a TLS handshake when the protocol is https (an untrusted certificate is
reported, not treated as a hard failure), a listen-port availability check, and a reminder that
the proxy never injects credentials. It sends no HTTP request and no credentials —
reachability is answered at the transport layer — so a doctor run cannot burn rate-limit quota.
It exits 0 when everything is fine and 1 when something needs fixing,
which makes it usable as a startup gate:
llm-session-proxy --doctor && llm-session-proxy
Local status endpoints
# Overview: session count, hit rate, injected headers, upstream target
curl http://127.0.0.1:9355/__llm_session_proxy__/status
# Per-session detail: ID, request count, last used
curl http://127.0.0.1:9355/__llm_session_proxy__/sessions
Handy when debugging: if sessions.active stays at 0, requests are not reaching
the proxy at all — usually the model ID is missing its proxy- prefix and the
client intercepted it.
How it works
Each request goes through a fixed sequence; the three-tier session strategy is the key part.
The order of 03–06 is deliberate. Rules see the model name the client sent and the
resolved one; the bucket's model override lands after alias resolution, so it can only be a real
upstream ID; injection comes last-but-one, so {{model}} in a header reflects the
final decision; and protocol conversion runs last of all — transforms and injection work
on the client's field names, and only the finished body is translated.
The three-tier session strategy
Explicit identifier — the client sent x-opencode-session itself, or the
body carries session_id. Reused as-is; this is the most accurate path.
Content fingerprint — SHA-256 over system plus the first user message.
Later turns in the same conversation only append messages, so the first one never changes,
the fingerprint stays stable, and the turn lands back on the same ID.
One-off random — when neither is available (a non-JSON body, say), a random ID is sent so the upstream does not return 400. These requests are kept out of the session table so it cannot grow without bound.
Stability
A single malformed request cannot take the process down. An invalid Host
header, a malformed request line, or an upstream status line / header with illegal characters
all become ordinary 4xx / 5xx responses, and the proxy keeps serving.
Uncaught errors land in the log file. A last-resort handler writes the full stack of
uncaughtException and unhandledRejection into the log file. Node's
default prints to stderr and terminates immediately, leaving the log file empty — which looks
exactly like "logs are fine, the process just vanished". That is why file logging is on by
default: see Logs.
Limitations
Loopback only. The proxy forwards your API key, so do not bind the listen address to
0.0.0.0.
response.stream: false breaks SSE. Keep the default true
unless you genuinely want whole-response buffering.
With bufferBody: false the body cannot be rewritten — model rewriting and
parameter injection stop working, and only headers can be injected.
The built-in model map is a snapshot, not a live catalogue. Upstream model IDs change
without notice, and this proxy deliberately makes no network call to discover them. If an
alias stops resolving, it is a one-line override in your config file — not a bug — and
--dry-run shows which alias resolved to what before you send a real request.
Protocol conversion is deliberately lossy at the edges. Anthropic
thinking / signature stream deltas have no chat equivalent and are
dropped (by name, in the log); messages ↔ responses composes through chat rather
than a direct converter. What survives is text, tool calls, finish reasons and usage — the
parts clients actually act on. The transform registry remains a fixed list: the proxy never
loads a transform from a path.
Roadmap
Zero dependencies, local, single process — every item has to fit that shape. Full reasoning, the comparison with comparable gateways and the explicit non-goals live in ROADMAP.md.
| Milestone | Theme | Highlights |
|---|---|---|
v0.2.0 ✓ | Correctness | A curated built-in model map (27 aliases), a warn-once notice when a proxy--prefixed alias resolves to nothing, and the --dry-run / --doctor preflight. The release that fixes “my alias silently reached the upstream verbatim”. |
v0.2.1 ✓ | Routing | Five named in-tree transforms (noop, drop-fields, drop-empty-fields, rename-fields, clamp-max-tokens) and router buckets (default / background / think / longContext) with top-down first-match rules over path, model prefix, body field and request size. Off by default; the host the translators plug into. The load-a-plugin-from-a-path model originally sketched for this milestone was dropped on purpose. |
v0.2.2 ✓ | Protocol reach | Anthropic ↔ OpenAI ↔ Responses translation with field normalisation, streaming included — request bodies, JSON responses and event-by-event SSE transcoding. The release that stops a client’s protocol from deciding which models it can reach. |
v0.3 | Observability and control | Prometheus /metrics, JSON log mode, OTLP/JSON export, a session and prompt-cache inspector, token and cost accounting, and a zero-build local dashboard. |
v0.4 | Reliability under real upstreams | Multiple upstreams, fallback with jittered retry, circuit breaking, separated timeouts with a stream idle watchdog, graceful drain on shutdown. |
v0.5 | Safety and hygiene | Log redaction, regex guardrails, line-precise config schema errors, and --check-config for CI. |
v1.0 | Distribution and guarantees | Single-file binaries, a per-client compatibility matrix, protocol contract tests on golden SSE fixtures, a config migration tool, and semver from there on. |
The v0.2 performance targets are measured by a harness that ships inside the repo: added latency of ≤ 1 ms p50 and ≤ 5 ms p99 against a local mock upstream, ≥ 1,200 RPS at 100 concurrent connections, zero process exits under the malformed-input corpus, and byte-for-byte SSE pass-through. Numbers get published together with their methodology — without it a benchmark says nothing.
Deliberately not on the roadmap: virtual keys and billing, running models locally, a hosted version, a plugin marketplace, a Kubernetes operator. Each of them pulls against “zero dependencies, local, single process”.