# TrueCluster Prototype heterogeneous pipeline-parallel LLM inference cluster. See [`SPEC.md`](SPEC.md). ## Install ```bash pip install -e . ``` ## Run a cluster ```bash truecluster cluster --model Qwen/Qwen2.5-0.5B-Instruct --max-nodes 1 --quant fp16 ``` ## Run a node Nodes download/resolve the cluster model from HuggingFace themselves and load only the assigned layer range. ```bash truecluster node --cluster-host 127.0.0.1 --cluster-port 7001 --device auto ``` ## Generate ```bash curl http://127.0.0.1:8000/v1/completions \ -H 'content-type: application/json' \ -d '{"model":"Qwen/Qwen2.5-0.5B-Instruct","prompt":"Hello","max_tokens":32}' ```