# TrueCluster Prototype heterogeneous pipeline-parallel LLM inference cluster. See [`SPEC.md`](SPEC.md). ## Install ```bash pip install -e . ``` ### Nvidia CUDA install If a node has an Nvidia GPU, it must install a CUDA-enabled PyTorch build. If you see: ```text Torch not compiled with CUDA enabled ``` then the node installed the CPU-only PyTorch package. Recommended fix on a macOS/Linux Nvidia node: ```bash source .venv/bin/activate pip uninstall -y torch torchvision torchaudio pip install --index-url https://download.pytorch.org/whl/cu121 torch torchvision torchaudio pip install -e . ``` Recommended fix on a Windows Nvidia node using PowerShell: ```powershell .\.venv\Scripts\Activate.ps1 python -m pip uninstall -y torch torchvision torchaudio python -m pip install --index-url https://download.pytorch.org/whl/cu121 torch torchvision torchaudio python -m pip install -e . ``` If PowerShell blocks venv activation, run this once in the same PowerShell window: ```powershell Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass .\.venv\Scripts\Activate.ps1 ``` Windows Command Prompt alternative: ```bat .venv\Scripts\activate.bat python -m pip uninstall -y torch torchvision torchaudio python -m pip install --index-url https://download.pytorch.org/whl/cu121 torch torchvision torchaudio python -m pip install -e . ``` For newer CUDA builds, PyTorch may also provide `cu124` or `cu126` wheels. Check https://pytorch.org/get-started/locally/ if `cu121` is not appropriate. Verify CUDA support: ```bash python - <<'PY' import torch print('torch:', torch.__version__) print('cuda available:', torch.cuda.is_available()) print('cuda version:', torch.version.cuda) print('gpu:', torch.cuda.get_device_name(0) if torch.cuda.is_available() else None) PY ``` Then run the node with: ```bash truecluster node --cluster-host YOUR_CLUSTER_IP --cluster-port 7001 --device cuda:0 ``` ## Run a cluster ```bash truecluster cluster --model Qwen/Qwen2.5-0.5B-Instruct --max-nodes 1 --quant fp16 ``` ## Run a node Nodes download/resolve the cluster model from HuggingFace themselves and load only the assigned layer range. ```bash truecluster node --cluster-host 127.0.0.1 --cluster-port 7001 --device auto ``` ## Generate ```bash curl http://127.0.0.1:8000/v1/completions \ -H 'content-type: application/json' \ -d '{"model":"Qwen/Qwen2.5-0.5B-Instruct","prompt":"Hello","max_tokens":32}' ```