* VLLM upgrade to 0.12.0 and compatibility fixes
* MAX_NUM_BATCHED_TOKENS fix and CUDA tester
* Sys kill worker instead of marking as failed
* upgrade to vllm 0.12.0
* Update to vllm 0.15.0 and lora fix
* Update for HUB and removal of deprected env variables
* reverted docker-bake changes
* removed leftovers
* Update src/handler.py
Co-authored-by: Dj Isaac <contact@dejaydev.com>
* Update src/utils.py
Co-authored-by: Dj Isaac <contact@dejaydev.com>
* Update src/handler.py
Co-authored-by: Dj Isaac <contact@dejaydev.com>
* Clean up of docs and comments in code
* nit: lowercase p
* nit: lowercase p
---------
Co-authored-by: Dj Isaac <contact@dejaydev.com>
Co-authored-by: chrisvela <chris.vela@runpod.io>
* fix: update CUDA to 12.4.1 for Blackwell GPU support
- Update Dockerfile base image from CUDA 12.1.0 to 12.4.1
- Update ldconfig path to cuda-12.4
- Update FlashInfer installation to use flashinfer-python package
- Add NVIDIA B200 (Blackwell) to supported gpuIds in hub.json
This fixes the "imagePullAsync: failed to get self-hosted image registry auth"
error when deploying on Blackwell GPUs (RTX PRO 6000, B200) by aligning
the Docker image CUDA version with the allowedCudaVersions in hub.json.
Fixes: DR-1118
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* revert: remove NVIDIA B200 from default gpuIds
The gpuIds in hub.json controls default GPU selection for deployments,
not GPU compatibility. The CUDA 12.4 upgrade is sufficient to enable
Blackwell GPU support.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* fix: remove FlashInfer to avoid JIT compilation errors
FlashInfer requires nvcc to JIT-compile CUDA kernels at runtime for
new GPU architectures (like Blackwell SM 10.0). Since we use the CUDA
base image without the toolkit, nvcc is not available.
vLLM will use its built-in fallback sampling methods instead.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
remove unsupported cuda versions (12.1-12.3) from hub.json and tests.json
to fix compatibility issues with worker deployment
- hub.json: remove 12.1, 12.2, 12.3 from allowedCudaVersions
- tests.json: remove 12.1, 12.2, 12.3, 12.4 from allowedCudaVersions
refs: AE-1452
* feat: moved config into docs; added banner; auto detect "messages" in input
* docs: moved config into docs
* chore: added .DS_Store
* chore: get the original stuff working again
* chore: remove all changes
* docs: reduced toc and added small config table
---------
Co-authored-by: Tim Pietrusky <tim.pietrusky@runpod.io>
* docs: remove outdated video; remove old info; added missing config for tools
* ci: use proper release for dev (pr only) and production (release only)
* ci(hub): added openai example; use smollm2 as base model
* docs: added conventions to be able to work with ai ide's
* chore: remove outdated stuff
* chore: update copyright to 2025
* ci: added github permissions
* feat: added gpuIds, gputCount and allowedCudaVersions; removed default value for LOAD_FORMAT to check which influence this has on the ui
---------
Co-authored-by: Tim Pietrusky <tim.pietrusky@runpod.io>