Files
opencode-proxy/README.md
T
2026-07-21 15:51:16 -05:00

6.6 KiB

OpenCode Go Proxy

A small, unauthenticated OpenAI-compatible proxy for OpenCode Go.

Setup

Requires Node.js 18 or newer.

npm install
cp .env.example .env

Edit .env with your Postgres connection string and optional server settings. Values already present in the process environment take precedence over .env, which makes the same configuration work locally and in hosted deployments. .env is ignored by Git.

Create keys.txt in the project root. Put one OpenCode Go API key on each line. Blank lines and comments are ignored; inline comments are supported.

go-first-key
go-second-key # optional comment
# disabled-key

For deployments such as Railway, keys can instead be supplied as a JSON environment variable. When set, OPENCODE_API_KEYS takes precedence over keys.txt:

OPENCODE_API_KEYS='["go-first-key","go-second-key"]'

To convert keys.txt into a copyable JSON value:

npm run keys:env

The global context limit defaults to 256k tokens and can be changed with MAX_CONTEXT_TOKENS.

Inference requests are limited per client IP to 2 requests per second, 100 requests per five hours, and $15 of reported upstream cost per five hours. Railway's X-Real-IP header is used to identify clients.

To exempt trusted clients from those proxy limits, set UNLIMITED_API_KEYS to a JSON array and have the client send a configured key using Authorization: Bearer <key> or x-api-key. These keys only bypass this proxy's rate and spend limits; they do not bypass upstream OpenCode Go limits.

UNLIMITED_API_KEYS='["personal-unlimited-key"]'

By default, ordinary clients may use only minimax-m3, minimax-m2.7, minimax-m2.5, kimi-k2.7-code, kimi-k2.6, kimi-k2.5, glm-5.2, glm-5.1, glm-5, deepseek-v4-pro, deepseek-v4-flash, qwen3.5-plus, mimo-v2.5, hy3-preview, and grok-4.5. To replace that list, set ALLOWED_MODELS to a comma-separated allowlist. The proxy removes every other model from the OpenAI and Anthropic model lists and rejects inference requests using a model outside the list. Clients using an UNLIMITED_API_KEYS key bypass this restriction and see the complete model list. Set ALLOWED_MODELS to an empty value to allow every upstream model.

ALLOWED_MODELS=kimi-k3,gpt-4.1

Set DATABASE_URL in .env (or the process environment) to a Postgres connection string. The proxy creates its requests table and index automatically on startup. Only requests to the OpenAI Chat Completions and Responses endpoints and the Anthropic Messages endpoints are stored. Their complete responses—including streaming responses and errors—are stored with headers, status, model, client IP, and timestamp. On startup, stored requests for all other endpoints are deleted. Successful chat completion requests also store a normalized training conversation containing the full chat history and generated assistant output, including reasoning, tool calls, and tool results.

npm start

The proxy listens at http://localhost:4005/oai/v1 and http://localhost:4005/ant/v1. It does not require authentication from clients. Keys are selected round-robin for each request and are sent upstream as Bearer tokens. If an upstream request fails, the proxy tries each remaining key once before returning the final failure response.

Endpoints

curl http://localhost:4005/oai/v1/models
curl http://localhost:4005/oai/v1/chat/completions \
  -H 'content-type: application/json' \
  -d '{"model":"kimi-k3","messages":[{"role":"user","content":"Hello"}]}'

Codex and other OpenAI Responses API clients can use POST /oai/v1/responses. Responses requests are translated to the upstream Chat Completions API and responses, including streaming events, are translated back to the Responses format.

Streaming requests are supported with "stream":true and are passed through as Server-Sent Events. The model list is fetched from OpenCode Go, so it can include models whose upstream transport is not OpenAI Chat Completions.

The Anthropic-compatible adapter exposes POST /ant/v1/messages and GET /ant/v1/models. Messages, tools, JSON responses, and streaming events are translated to and from OpenAI Chat Completions. Anthropic clients should use their normal Messages API request format and can use the OpenCode model IDs returned by the models endpoint.

For Claude Code, either set ANTHROPIC_BASE_URL=http://localhost:4005/ant or use the displayed http://localhost:4005/ant/v1 value; both forms are supported.

Exporting training data

Each export creates two JSONL files. The regular export-<timestamp>.jsonl remains unchanged: each line is one inference request containing its complete input history followed by the generated assistant message. The additional export-<timestamp>-traces.jsonl removes cumulative intermediate snapshots and keeps each complete conversation branch. If a conversation is rewound and continued in multiple ways, every branch leaf is retained as its own full trace.

Chat Completions, Responses, and Anthropic Messages requests are all exported in the same normalized OpenAI-style messages format. It preserves system/user/assistant roles, reasoning_content, assistant tool_calls, and tool results. Non-chat routes and failed inference requests are excluded.

Each line also has metadata with the exact tool definitions supplied on that request, tool choice, model, API format, endpoint, streaming mode, generation parameters, caller-supplied request metadata, response ID/model, finish reason, usage, HTTP status, duration, request ID, and timestamp. Tool metadata remains in its original OpenAI, Responses, or Anthropic format so no provider-specific schema information is lost.

{"messages":[{"role":"user","content":"What is 2+2?"},{"role":"assistant","content":"4","reasoning_content":"Adding the values gives four."}],"metadata":{"schema_version":1,"api":"openai_chat_completions","model":"model-a","stream":false,"tools":[],"response_status":200}}

Omit the limit to export every training request, newest first:

npm run --silent export

Pass a positive limit to export that many of the most recent training requests:

npm run --silent export -- 100
# Equivalent: npm run --silent export -- --limit 100

The export command uses the same DATABASE_URL as the server and saves both files in the project root. The limit applies to the regular per-request export; the traces file contains the complete branch leaves found among those requests. Rows collected before normalized training storage was added are normalized from their saved raw request and response during export.

Tests

npm test