adding model limits
This commit is contained in:
@@ -43,6 +43,12 @@ To exempt trusted clients from those proxy limits, set `UNLIMITED_API_KEYS` to a
|
||||
UNLIMITED_API_KEYS='["personal-unlimited-key"]'
|
||||
```
|
||||
|
||||
By default, ordinary clients may use only `minimax-m3`, `minimax-m2.7`, `minimax-m2.5`, `kimi-k2.7-code`, `kimi-k2.6`, `kimi-k2.5`, `glm-5.2`, `glm-5.1`, `glm-5`, `deepseek-v4-pro`, `deepseek-v4-flash`, `qwen3.5-plus`, `mimo-v2.5`, `hy3-preview`, and `grok-4.5`. To replace that list, set `ALLOWED_MODELS` to a comma-separated allowlist. The proxy removes every other model from the OpenAI and Anthropic model lists and rejects inference requests using a model outside the list. Clients using an `UNLIMITED_API_KEYS` key bypass this restriction and see the complete model list. Set `ALLOWED_MODELS` to an empty value to allow every upstream model.
|
||||
|
||||
```sh
|
||||
ALLOWED_MODELS=kimi-k3,gpt-4.1
|
||||
```
|
||||
|
||||
Set `DATABASE_URL` in `.env` (or the process environment) to a Postgres connection string. The proxy creates its `requests` table and index automatically on startup. Only requests to the OpenAI Chat Completions and Responses endpoints and the Anthropic Messages endpoints are stored. Their complete responses—including streaming responses and errors—are stored with headers, status, model, client IP, and timestamp. On startup, stored requests for all other endpoints are deleted. Successful chat completion requests also store a normalized training conversation containing the full chat history and generated assistant output, including reasoning, tool calls, and tool results.
|
||||
|
||||
```sh
|
||||
|
||||
Reference in New Issue
Block a user