- Bump vllm[flashinfer] to 0.20.0 in Dockerfile
- Remove io_processor param from OpenAIServingRender (dropped in 0.20.0)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Fixes FDE-174. Some model stores (e.g. RunPod pre-cached network volumes)
normalize repo IDs to lowercase. HuggingFace Hub caches using the original
casing, so MODEL_NAME=Qwen/Qwen2.5-Coder-32B-Instruct-AWQ would miss a
cache stored as models--qwen--qwen2.5-coder-32b-instruct-awq/ and attempt
a redundant download that fails on limited container storage.
If the exact-case HF cache directory is absent but a lowercase variant
exists, the latest snapshot path is returned directly so vLLM loads from
disk. Absolute paths and models with no lowercase cache are unchanged.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Fixes FDE-194. Previously a malformed LORA_MODULES value was swallowed at
info level and the engine would start with no LoRA adapters, causing 500s
on any request using an adapter model name (e.g. npc-sim-*).
Changes:
- Log at error level when LORA_MODULES cannot be parsed as JSON
- Log at error level when individual adapter dicts fail LoRAModulePath validation
- Log a final error when all adapters fail to load so the cause is obvious
- Accept a single adapter dict (not just an array) for convenience
- Return early when LORA_MODULES is unset to skip unnecessary parsing
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Bump vllm[flashinfer] to 0.19.1 in Dockerfile
- Add OpenAIServingRender (new required dependency in 0.19.x serving layer)
- Pass openai_serving_render to all four serving class constructors
- Remove log_error_stack param (removed upstream in 0.19.x)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- engine.py: drop UnboundLocalError-prone isinstance(response, ...) checks in
except blocks of _handle_responses_request and _handle_messages_request;
emit SSE-shaped error frames mid-stream instead of raw dicts; add missing
blank line between handlers.
- engine_args.py: restructure LMCache HMA guard so the warning branch is
actually reachable when user explicitly sets disable_hybrid_kv_cache_manager=False,
and correct the inverted message (HMA must be disabled = True).
- requirements.txt: drop stray whitespace in transformers version specifier.