- engine.py: drop UnboundLocalError-prone isinstance(response, ...) checks in except blocks of _handle_responses_request and _handle_messages_request; emit SSE-shaped error frames mid-stream instead of raw dicts; add missing blank line between handlers. - engine_args.py: restructure LMCache HMA guard so the warning branch is actually reachable when user explicitly sets disable_hybrid_kv_cache_manager=False, and correct the inverted message (HMA must be disabled = True). - requirements.txt: drop stray whitespace in transformers version specifier.
16 lines
209 B
Plaintext
16 lines
209 B
Plaintext
ray
|
|
pandas
|
|
pyarrow
|
|
runpod
|
|
huggingface-hub
|
|
lmcache==0.4.1
|
|
packaging>=24.2
|
|
typing-extensions>=4.8.0
|
|
pydantic
|
|
pydantic-settings
|
|
hf-transfer
|
|
transformers>=4.57.0,<5
|
|
bitsandbytes>=0.45.0
|
|
kernels
|
|
torch-c-dlpack-ext
|