adding model limits

This commit is contained in:
2026-07-21 15:43:26 -05:00
parent 8229d878da
commit a1f2f13b1d
4 changed files with 132 additions and 17 deletions
+6
View File
@@ -43,6 +43,12 @@ To exempt trusted clients from those proxy limits, set `UNLIMITED_API_KEYS` to a
UNLIMITED_API_KEYS='["personal-unlimited-key"]'
```
By default, ordinary clients may use only `minimax-m3`, `minimax-m2.7`, `minimax-m2.5`, `kimi-k2.7-code`, `kimi-k2.6`, `kimi-k2.5`, `glm-5.2`, `glm-5.1`, `glm-5`, `deepseek-v4-pro`, `deepseek-v4-flash`, `qwen3.5-plus`, `mimo-v2.5`, `hy3-preview`, and `grok-4.5`. To replace that list, set `ALLOWED_MODELS` to a comma-separated allowlist. The proxy removes every other model from the OpenAI and Anthropic model lists and rejects inference requests using a model outside the list. Clients using an `UNLIMITED_API_KEYS` key bypass this restriction and see the complete model list. Set `ALLOWED_MODELS` to an empty value to allow every upstream model.
```sh
ALLOWED_MODELS=kimi-k3,gpt-4.1
```
Set `DATABASE_URL` in `.env` (or the process environment) to a Postgres connection string. The proxy creates its `requests` table and index automatically on startup. Only requests to the OpenAI Chat Completions and Responses endpoints and the Anthropic Messages endpoints are stored. Their complete responses—including streaming responses and errors—are stored with headers, status, model, client IP, and timestamp. On startup, stored requests for all other endpoints are deleted. Successful chat completion requests also store a normalized training conversation containing the full chat history and generated assistant output, including reasoning, tool calls, and tool results.
```sh