2026-06-04 18:11:26 -05:00
2026-06-04 18:11:26 -05:00
2026-06-04 18:11:26 -05:00
2026-06-04 19:03:03 -05:00

TrueCluster

Prototype heterogeneous pipeline-parallel LLM inference cluster.

See SPEC.md.

Install

pip install -e .

Nvidia CUDA install

If a node has an Nvidia GPU, it must install a CUDA-enabled PyTorch build. If you see:

Torch not compiled with CUDA enabled

then the node installed the CPU-only PyTorch package.

Recommended fix on a macOS/Linux Nvidia node:

source .venv/bin/activate
pip uninstall -y torch torchvision torchaudio
pip install --index-url https://download.pytorch.org/whl/cu121 torch torchvision torchaudio
pip install -e .

Recommended fix on a Windows Nvidia node using PowerShell:

.\.venv\Scripts\Activate.ps1
python -m pip uninstall -y torch torchvision torchaudio
python -m pip install --index-url https://download.pytorch.org/whl/cu121 torch torchvision torchaudio
python -m pip install -e .

If PowerShell blocks venv activation, run this once in the same PowerShell window:

Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass
.\.venv\Scripts\Activate.ps1

Windows Command Prompt alternative:

.venv\Scripts\activate.bat
python -m pip uninstall -y torch torchvision torchaudio
python -m pip install --index-url https://download.pytorch.org/whl/cu121 torch torchvision torchaudio
python -m pip install -e .

For newer CUDA builds, PyTorch may also provide cu124 or cu126 wheels. Check https://pytorch.org/get-started/locally/ if cu121 is not appropriate.

Verify CUDA support:

python - <<'PY'
import torch
print('torch:', torch.__version__)
print('cuda available:', torch.cuda.is_available())
print('cuda version:', torch.version.cuda)
print('gpu:', torch.cuda.get_device_name(0) if torch.cuda.is_available() else None)
PY

Then run the node with:

truecluster node --cluster-host YOUR_CLUSTER_IP --cluster-port 7001 --device cuda:0

Run a cluster

truecluster cluster --model Qwen/Qwen2.5-0.5B-Instruct --max-nodes 1 --quant fp16

Run a node

Nodes download/resolve the cluster model from HuggingFace themselves and load only the assigned layer range.

truecluster node --cluster-host 127.0.0.1 --cluster-port 7001 --device auto

Generate

curl http://127.0.0.1:8000/v1/completions \
  -H 'content-type: application/json' \
  -d '{"model":"Qwen/Qwen2.5-0.5B-Instruct","prompt":"Hello","max_tokens":32}'
S
Description
No description provided
Readme
92 KiB
Languages
Python 100%