OpenCode Go Proxy
A small, unauthenticated OpenAI-compatible proxy for OpenCode Go.
Setup
Requires Node.js 18 or newer.
npm install
cp .env.example .env
Edit .env with your Postgres connection string and optional server settings. Values already present in the process environment take precedence over .env, which makes the same configuration work locally and in hosted deployments. .env is ignored by Git.
Create keys.txt in the project root. Put one OpenCode Go API key on each line. Blank lines and comments are ignored; inline comments are supported.
go-first-key
go-second-key # optional comment
# disabled-key
For deployments such as Railway, keys can instead be supplied as a JSON environment variable. When set, OPENCODE_API_KEYS takes precedence over keys.txt:
OPENCODE_API_KEYS='["go-first-key","go-second-key"]'
To convert keys.txt into a copyable JSON value:
npm run keys:env
The global context limit defaults to 256k tokens and can be changed with MAX_CONTEXT_TOKENS.
Inference requests are limited per client IP to 2 requests per second, 100 requests per five hours, and $10 of reported upstream cost per five hours. Railway's X-Real-IP header is used to identify clients.
Set DATABASE_URL in .env (or the process environment) to a Postgres connection string. The proxy creates its requests table and index automatically on startup. Every incoming request and its complete response—including streaming responses and errors—is stored with headers, status, model, client IP, and timestamp. Successful inference requests also store a normalized training conversation containing the full chat history and generated assistant output, including reasoning, tool calls, and tool results.
npm start
The proxy listens at http://localhost:4005/oai/v1 and http://localhost:4005/ant/v1. It does not require authentication from clients. Keys are selected round-robin for each request and are sent upstream as Bearer tokens. If an upstream request fails, the proxy tries each remaining key once before returning the final failure response.
Endpoints
curl http://localhost:4005/oai/v1/models
curl http://localhost:4005/oai/v1/chat/completions \
-H 'content-type: application/json' \
-d '{"model":"kimi-k3","messages":[{"role":"user","content":"Hello"}]}'
Codex and other OpenAI Responses API clients can use POST /oai/v1/responses. Responses requests are translated to the upstream Chat Completions API and responses, including streaming events, are translated back to the Responses format.
Streaming requests are supported with "stream":true and are passed through as Server-Sent Events. The model list is fetched from OpenCode Go, so it can include models whose upstream transport is not OpenAI Chat Completions.
The Anthropic-compatible adapter exposes POST /ant/v1/messages and GET /ant/v1/models. Messages, tools, JSON responses, and streaming events are translated to and from OpenAI Chat Completions. Anthropic clients should use their normal Messages API request format and can use the OpenCode model IDs returned by the models endpoint.
For Claude Code, either set ANTHROPIC_BASE_URL=http://localhost:4005/ant or use the displayed http://localhost:4005/ant/v1 value; both forms are supported.
Exporting training data
Exports use JSONL: each line is one inference request containing its complete input history followed by the generated assistant message. The normalized OpenAI-style messages preserve system/user/assistant roles, reasoning_content, assistant tool_calls, and tool results. Non-chat routes and failed inference requests are excluded.
Each line also has metadata with the exact tool definitions supplied on that request, tool choice, model, API format, endpoint, streaming mode, generation parameters, caller-supplied request metadata, response ID/model, finish reason, usage, HTTP status, duration, request ID, and timestamp. Tool metadata remains in its original OpenAI, Responses, or Anthropic format so no provider-specific schema information is lost.
{"messages":[{"role":"user","content":"What is 2+2?"},{"role":"assistant","content":"4","reasoning_content":"Adding the values gives four."}],"metadata":{"schema_version":1,"api":"openai_chat_completions","model":"model-a","stream":false,"tools":[],"response_status":200}}
Omit the limit to export every training request, newest first:
npm run --silent export > requests.jsonl
Pass a positive limit to export that many of the most recent training requests:
npm run --silent export -- 100 > recent-requests.jsonl
# Equivalent: npm run --silent export -- --limit 100
The export command uses the same DATABASE_URL as the server. --silent suppresses npm's banner, and the script's progress message is written to standard error, so redirected JSONL remains valid. Rows collected before normalized training storage was added are normalized from their saved raw request and response during export.
Tests
npm test