fix: address review comments on responses/messages handlers and lmcache guard

- engine.py: drop UnboundLocalError-prone isinstance(response, ...) checks in
  except blocks of _handle_responses_request and _handle_messages_request;
  emit SSE-shaped error frames mid-stream instead of raw dicts; add missing
  blank line between handlers.
- engine_args.py: restructure LMCache HMA guard so the warning branch is
  actually reachable when user explicitly sets disable_hybrid_kv_cache_manager=False,
  and correct the inverted message (HMA must be disabled = True).
- requirements.txt: drop stray whitespace in transformers version specifier.
This commit is contained in:
Tim Pietrusky
2026-04-23 11:28:50 +02:00
parent f299204770
commit 3403889528
3 changed files with 29 additions and 33 deletions
+1 -1
View File
@@ -9,7 +9,7 @@ typing-extensions>=4.8.0
pydantic
pydantic-settings
hf-transfer
transformers>= 4.57.0,< 5
transformers>=4.57.0,<5
bitsandbytes>=0.45.0
kernels
torch-c-dlpack-ext