Release / release (push) Waiting to run
Three coupled fixes verified end-to-end on a private fork (TimPietruskyRunPod/worker-vllm v0.1.3 → both hub tests passing): 1. tests.json allowedCudaVersions: 12.x → 13.0 The Dockerfile and hub.json moved to CUDA 13.0 in v2.20.0 (#288, #289), but tests.json was still pinned to 12.5–12.9, so the test pod was scheduled on a GPU with driver < 13.0 and container init failed at the nvidia-container-cli hook with "unsatisfied condition: cuda>=13.0". 2. requirements.txt kernels<0.15 huggingface/kernels v0.15.1 tightened LayerRepository to require a revision or version argument (https://github.com/huggingface/kernels/pull/544). transformers >=5 still constructs LayerRepository(repo_id=..., layer_name=...) without either, so worker import raised ValueError during `from transformers import ...`. 0.14.1 is the last safe release. 3. tests.json timeout 30000 → 300000 vLLM cold start (torch.compile + FlashInfer warmup) on RTX 4090 for SmolLM2-135M takes ~60–70s before the first request can be served. The previous 30s per-test timeout fired before the worker came up, producing "context cancelled or timed out: context deadline exceeded" for every test even when the worker was healthy. 300s gives enough headroom for cold start + the actual inference call. Refs: DR-1161
44 lines
991 B
JSON
44 lines
991 B
JSON
{
|
|
"tests": [
|
|
{
|
|
"name": "basic_inference_test",
|
|
"input": {
|
|
"prompt": "Write a short poem about artificial intelligence."
|
|
},
|
|
"timeout": 300000
|
|
},
|
|
{
|
|
"name": "openai_messages_test",
|
|
"input": {
|
|
"openai_route": "/v1/chat/completions",
|
|
"openai_input": {
|
|
"messages": [
|
|
{
|
|
"role": "system",
|
|
"content": "You are a helpful assistant that writes concise responses."
|
|
},
|
|
{
|
|
"role": "user",
|
|
"content": "Explain what a neural network is in one sentence."
|
|
}
|
|
],
|
|
"max_tokens": 200,
|
|
"temperature": 0.1
|
|
}
|
|
},
|
|
"timeout": 300000
|
|
}
|
|
],
|
|
"config": {
|
|
"gpuTypeId": "NVIDIA GeForce RTX 4090",
|
|
"gpuCount": 1,
|
|
"env": [
|
|
{
|
|
"key": "MODEL_NAME",
|
|
"value": "HuggingFaceTB/SmolLM2-135M-Instruct"
|
|
}
|
|
],
|
|
"allowedCudaVersions": ["13.0"]
|
|
}
|
|
}
|