# OpenCode Go Proxy A small, unauthenticated OpenAI-compatible proxy for OpenCode Go. ## Setup Requires Node.js 18 or newer. ```sh npm install cp .env.example .env ``` Edit `.env` with your Postgres connection string and optional server settings. Values already present in the process environment take precedence over `.env`, which makes the same configuration work locally and in hosted deployments. `.env` is ignored by Git. Create `keys.txt` in the project root. Put one OpenCode Go API key on each line. Blank lines and comments are ignored; inline comments are supported. ```text go-first-key go-second-key # optional comment # disabled-key ``` For deployments such as Railway, keys can instead be supplied as a JSON environment variable. When set, `OPENCODE_API_KEYS` takes precedence over `keys.txt`: ```sh OPENCODE_API_KEYS='["go-first-key","go-second-key"]' ``` To convert `keys.txt` into a copyable JSON value: ```sh npm run keys:env ``` The global context limit defaults to 256k tokens and can be changed with `MAX_CONTEXT_TOKENS`. Inference requests are limited per client IP to 2 requests per second, 100 requests per five hours, and $15 of reported upstream cost per five hours. Railway's `X-Real-IP` header is used to identify clients. Set `REDIS_URL` to persist and share these limits across process restarts and multiple proxy instances. The Redis backend stores only the active five-hour window and uses atomic operations, so concurrent instances enforce one shared limit. Without `REDIS_URL`, limits remain in memory as before. ```sh REDIS_URL=redis://localhost:6379 ``` To exempt trusted clients from those proxy limits, set `UNLIMITED_API_KEYS` to a JSON array and have the client send a configured key using `Authorization: Bearer ` or `x-api-key`. These keys only bypass this proxy's rate and spend limits; they do not bypass upstream OpenCode Go limits. ```sh UNLIMITED_API_KEYS='["personal-unlimited-key"]' ``` By default, ordinary clients may use only `minimax-m3`, `minimax-m2.7`, `minimax-m2.5`, `kimi-k2.7-code`, `kimi-k2.6`, `kimi-k2.5`, `glm-5.2`, `glm-5.1`, `glm-5`, `deepseek-v4-pro`, `deepseek-v4-flash`, `qwen3.5-plus`, `mimo-v2.5`, `hy3-preview`, and `grok-4.5`. To replace that list, set `ALLOWED_MODELS` to a comma-separated allowlist. The proxy removes every other model from the OpenAI and Anthropic model lists and rejects inference requests using a model outside the list. Clients using an `UNLIMITED_API_KEYS` key bypass this restriction and see the complete model list. Set `ALLOWED_MODELS` to an empty value to allow every upstream model. ```sh ALLOWED_MODELS=kimi-k3,gpt-4.1 ``` Set `DATABASE_URL` in `.env` (or the process environment) to a Postgres connection string. The proxy creates its `requests` table and index automatically on startup. Only requests to the OpenAI Chat Completions and Responses endpoints and the Anthropic Messages endpoints are stored. Their complete responses—including streaming responses and errors—are stored with headers, status, model, client IP, and timestamp. On startup, stored requests for all other endpoints are deleted. Successful chat completion requests also store a normalized training conversation containing the full chat history and generated assistant output, including reasoning, tool calls, and tool results. ```sh npm start ``` The proxy listens at `http://localhost:4005/oai/v1` and `http://localhost:4005/ant/v1`. It does not require authentication from clients. Keys are selected round-robin for each request and are sent upstream as Bearer tokens. If an upstream request fails, the proxy tries each remaining key once before returning the final failure response. ## Endpoints ```sh curl http://localhost:4005/oai/v1/models ``` ```sh curl http://localhost:4005/oai/v1/chat/completions \ -H 'content-type: application/json' \ -d '{"model":"kimi-k3","messages":[{"role":"user","content":"Hello"}]}' ``` Codex and other OpenAI Responses API clients can use `POST /oai/v1/responses`. Responses requests are translated to the upstream Chat Completions API and responses, including streaming events, are translated back to the Responses format. Streaming requests are supported with `"stream":true` and are passed through as Server-Sent Events. The model list is fetched from OpenCode Go, so it can include models whose upstream transport is not OpenAI Chat Completions. The Anthropic-compatible adapter exposes `POST /ant/v1/messages` and `GET /ant/v1/models`. Messages, tools, JSON responses, and streaming events are translated to and from OpenAI Chat Completions. Anthropic clients should use their normal Messages API request format and can use the OpenCode model IDs returned by the models endpoint. For Claude Code, either set `ANTHROPIC_BASE_URL=http://localhost:4005/ant` or use the displayed `http://localhost:4005/ant/v1` value; both forms are supported. ## Exporting training data Each export creates two JSONL files. The regular `export-.jsonl` remains unchanged: each line is one inference request containing its complete input history followed by the generated assistant message. The additional `export--traces.jsonl` removes cumulative intermediate snapshots and keeps each complete conversation branch. If a conversation is rewound and continued in multiple ways, every branch leaf is retained as its own full trace. Chat Completions, Responses, and Anthropic Messages requests are all exported in the same normalized OpenAI-style `messages` format. It preserves system/user/assistant roles, `reasoning_content`, assistant `tool_calls`, and `tool` results. Non-chat routes and failed inference requests are excluded. Each line also has `metadata` with the exact tool definitions supplied on that request, tool choice, model, API format, endpoint, streaming mode, generation parameters, caller-supplied request metadata, response ID/model, finish reason, usage, HTTP status, duration, request ID, and timestamp. Tool metadata remains in its original OpenAI, Responses, or Anthropic format so no provider-specific schema information is lost. ```json {"messages":[{"role":"user","content":"What is 2+2?"},{"role":"assistant","content":"4","reasoning_content":"Adding the values gives four."}],"metadata":{"schema_version":1,"api":"openai_chat_completions","model":"model-a","stream":false,"tools":[],"response_status":200}} ``` Omit the limit to export every training request, newest first: ```sh npm run --silent export ``` Pass a positive limit to export that many of the most recent training requests: ```sh npm run --silent export -- 100 # Equivalent: npm run --silent export -- --limit 100 ``` The export command uses the same `DATABASE_URL` as the server and saves both files in the project root. The limit applies to the regular per-request export; the traces file contains the complete branch leaves found among those requests. Rows collected before normalized training storage was added are normalized from their saved raw request and response during export. ## Tests ```sh npm test ```