fix: address review comments on responses/messages handlers and lmcache guard
- engine.py: drop UnboundLocalError-prone isinstance(response, ...) checks in except blocks of _handle_responses_request and _handle_messages_request; emit SSE-shaped error frames mid-stream instead of raw dicts; add missing blank line between handlers. - engine_args.py: restructure LMCache HMA guard so the warning branch is actually reachable when user explicitly sets disable_hybrid_kv_cache_manager=False, and correct the inverted message (HMA must be disabled = True). - requirements.txt: drop stray whitespace in transformers version specifier.
This commit is contained in:
@@ -9,7 +9,7 @@ typing-extensions>=4.8.0
|
||||
pydantic
|
||||
pydantic-settings
|
||||
hf-transfer
|
||||
transformers>= 4.57.0,< 5
|
||||
transformers>=4.57.0,<5
|
||||
bitsandbytes>=0.45.0
|
||||
kernels
|
||||
torch-c-dlpack-ext
|
||||
|
||||
Reference in New Issue
Block a user