22 KiB
TrueCluster Prototype Specification
1. Goal
TrueCluster is a Python 3.10 prototype for distributed LLM inference across a small cluster of heterogeneous machines. It allows a coordinator/cluster process to load a HuggingFace safetensors model config/tokenizer/coordinator-owned tensors, split transformer layers across connected worker nodes, tell each node which layer range to load from HuggingFace, and expose a simple OpenAI-compatible HTTP API for generation.
The initial target is small Qwen2/Qwen2.5-style decoder-only models around 0.5B-1.5B parameters, with first-class support for mixed Nvidia CUDA and Apple Silicon Mac MPS nodes in the same cluster. Qwen3.5 hybrid linear-attention models are a later target and are not part of the first fp16 prototype.
Primary prototype goals:
- Python 3.10.
- Easy CLI for running a cluster or node.
- Coordinator loads one model at a time.
- Nodes connect to the coordinator by host/port.
- Coordinator sends layer assignments to nodes over sockets.
- Nodes download/resolve the full HuggingFace model themselves, then load only their assigned layers.
- Support HuggingFace safetensors models.
- Initial architecture target: Qwen2/Qwen2.5-style causal language models.
- Support
fp16, portableint8, and portableint4weight-only quantization modes. - OpenAI-compatible unauthenticated HTTP API.
- Efficient generation by transferring activations during inference, not weights.
2. Non-goals for the Initial Prototype
The first prototype will intentionally avoid several advanced features:
- No multi-model serving.
- No authentication.
- No request batching.
- No tensor parallelism across machines.
- No expert parallelism.
- No dynamic model hot-swapping.
- No high-performance CUDA-only quantization kernels as a requirement.
- No coordinator-to-node weight transfer. Nodes are expected to have HuggingFace/model access.
- No web UI.
- No streaming responses in the first milestone.
- No fault-tolerant recovery during an active generation.
These can be added after a correct baseline works.
3. High-Level Architecture
TrueCluster uses pipeline-parallel inference.
The coordinator owns:
- CLI entry point for
cluster. - HTTP API.
- Tokenizer.
- Sampling logic.
- HuggingFace model metadata/config loading.
- Coordinator-owned tensor loading from safetensors.
- Model split planning.
- Embedding layer.
- Final normalization.
- LM head.
- Node registry and orchestration.
Each node owns:
- CLI entry point for
node. - Persistent socket connection to coordinator.
- Device detection and selection.
- One contiguous range of transformer layers.
- KV cache for its assigned layers.
- Local forward execution on CUDA, MPS, or CPU.
Generation path:
HTTP request
-> coordinator tokenizes prompt
-> coordinator runs embedding
-> hidden states sent to node 1
-> node 1 runs assigned layers
-> hidden states sent to node 2
-> ...
-> final hidden states returned to coordinator
-> coordinator runs final norm + lm_head
-> coordinator samples next token
-> repeat until completion
Weights are not transferred over the cluster socket. During assignment, the coordinator sends model id, config, and layer range. Each node downloads/resolves the model from HuggingFace or a matching local path and loads only its assigned layer tensors. During generation, only hidden states and small metadata are passed between coordinator and nodes.
4. Execution Model
4.1 Pipeline Parallelism
The model is split by complete transformer blocks. Each node receives a contiguous set of layers:
coordinator:
embed_tokens
final_norm
lm_head
node 1:
layers 0-7
node 2:
layers 8-15
node 3:
layers 16-23
This is simpler and more reliable than tensor parallelism for heterogeneous machines and normal Ethernet/Wi-Fi networks.
4.2 KV Cache Ownership
KV cache is stored on the node that owns the relevant layers.
For example:
node 1 cache: layers 0-7
node 2 cache: layers 8-15
node 3 cache: layers 16-23
During prefill, each node creates cache entries for its layers. During decode, each node appends one token of keys/values to its cache.
4.3 Request Concurrency
Initial prototype supports one active generation at a time per cluster.
Reason: distributed KV cache management is much easier with a single active request. Later versions can introduce request IDs, cache slots, and batching.
5. CLI Design
Use typer for the CLI.
Package command:
truecluster
5.1 Cluster Command
Example:
truecluster cluster \
--model Qwen/Qwen3.5-0.8B \
--node-host 0.0.0.0 \
--node-port 7001 \
--api-host 0.0.0.0 \
--api-port 8000 \
--max-nodes 4 \
--quant fp16
Options:
--model TEXT HuggingFace model id or local path.
--node-host TEXT Host/IP for worker-node socket server. Default: 0.0.0.0
--node-port INT Port for worker-node socket server. Default: 7001
--api-host TEXT Host/IP for HTTP API. Default: 0.0.0.0
--api-port INT HTTP API port. Default: 8000
--max-nodes INT Maximum number of worker nodes to use.
--quant [fp16|int8|int4] Weight mode. Default: fp16
--dtype [fp16|bf16|fp32] Compute dtype preference. Default: fp16
--target-node-memory-gb FLOAT Optional planning hint if node memory is unknown.
--trust-remote-code BOOL HuggingFace trust_remote_code. Default: false
--hf-cache-dir PATH Optional HuggingFace cache directory.
--log-level TEXT Default: info
Behavior:
- Resolve/download model.
- Load config and tokenizer.
- Build safetensors metadata index and load only coordinator-owned tensors.
- Start node socket server.
- Start HTTP API.
- Wait for enough nodes.
- Assign layer ranges.
- Nodes download/resolve the model and load assigned layers.
- Mark cluster ready.
5.2 Node Command
Example:
truecluster node \
--cluster-host 192.168.1.50 \
--cluster-port 7001 \
--device auto
Options:
--cluster-host TEXT Coordinator node socket host.
--cluster-port INT Coordinator node socket port.
--device TEXT auto, cuda, cuda:0, mps, or cpu. Default: auto
--node-id TEXT Optional stable node id.
--hf-cache-dir PATH Optional HuggingFace cache directory for node downloads.
--work-dir PATH Reserved for future temporary/cache files.
--max-memory-gb FLOAT Optional memory capability override.
--log-level TEXT Default: info
Behavior:
- Detect hardware and PyTorch backends.
- Connect to coordinator socket.
- Send
HELLOcapability message. - Wait for assignment.
- Receive config and assigned layer range.
- Download/resolve the model locally and build local model shard.
- Mark itself ready.
- Serve prefill/decode requests over the persistent connection.
6. HTTP API
Use FastAPI and Uvicorn.
The API is unauthenticated for the prototype.
6.1 GET /v1/models
Returns the single loaded model.
Example response:
{
"object": "list",
"data": [
{
"id": "Qwen/Qwen3.5-0.8B",
"object": "model",
"created": 0,
"owned_by": "truecluster"
}
]
}
6.2 POST /v1/completions
Supported request fields initially:
{
"model": "Qwen/Qwen3.5-0.8B",
"prompt": "Hello",
"max_tokens": 64,
"temperature": 0.7,
"top_p": 0.95,
"stop": null
}
Response should be OpenAI-compatible enough for common clients:
{
"id": "cmpl-...",
"object": "text_completion",
"created": 0,
"model": "Qwen/Qwen3.5-0.8B",
"choices": [
{
"text": " world",
"index": 0,
"logprobs": null,
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 1,
"completion_tokens": 1,
"total_tokens": 2
}
}
6.3 POST /v1/chat/completions
Supported request fields initially:
{
"model": "Qwen/Qwen3.5-0.8B",
"messages": [
{"role": "user", "content": "Hello"}
],
"max_tokens": 64,
"temperature": 0.7,
"top_p": 0.95,
"stop": null,
"stream": false
}
The coordinator should use the HuggingFace tokenizer chat template if available.
stream: true may return a clear unsupported error in the initial prototype.
6.4 Not-Ready Error
If a generation request arrives before enough nodes are connected and loaded:
{
"error": {
"message": "Model is not ready. Required nodes: 3, connected ready nodes: 1",
"type": "cluster_not_ready"
}
}
HTTP status: 503.
7. Node Socket Protocol
Use persistent TCP sockets with asyncio streams.
Encoding:
[8-byte unsigned big-endian payload length][msgpack payload]
Large tensor payloads are sent as chunked binary data inside protocol messages or as msgpack metadata followed by raw bytes.
7.1 Message Envelope
Every message should include:
{
"type": "MESSAGE_TYPE",
"request_id": "optional-request-id",
"seq": 1,
"payload": {}
}
7.2 Core Message Types
Coordinator/node lifecycle:
HELLO
HELLO_ACK
ASSIGNMENT
LOAD_COMPLETE
LOAD_FAILED
PING
PONG
ERROR
Inference:
CLEAR_CACHE
RUN_PREFILL
RUN_DECODE
HIDDEN_STATE
INFERENCE_ERROR
7.3 HELLO
Sent by node immediately after connection.
Example:
{
"type": "HELLO",
"payload": {
"node_id": "macbook-pro-1",
"hostname": "macbook-pro.local",
"python_version": "3.10.13",
"torch_version": "2.x",
"platform": "darwin",
"devices": [
{
"id": "mps",
"type": "mps",
"name": "Apple Silicon MPS",
"total_memory": null,
"free_memory": null
}
],
"selected_device": "mps",
"max_memory_bytes": null
}
}
CUDA example device:
{
"id": "cuda:0",
"type": "cuda",
"name": "NVIDIA GeForce RTX 4090",
"total_memory": 25757220864,
"free_memory": 23000000000
}
7.4 ASSIGNMENT
Sent by coordinator after planning.
{
"type": "ASSIGNMENT",
"payload": {
"model_id": "Qwen/Qwen3.5-0.8B",
"architecture": "qwen",
"quant": "fp16",
"compute_dtype": "fp16",
"layer_start": 0,
"layer_end_exclusive": 8,
"config": {},
"tensor_count": 128,
"total_weight_bytes": 123456789
}
}
7.5 Model Loading on Nodes
The coordinator does not transfer model weights. After ASSIGNMENT, each node resolves/downloads model_id itself using HuggingFace cache semantics and loads only tensors whose names match its assigned layer range.
7.6 Inference Messages
RUN_PREFILL:
{
"type": "RUN_PREFILL",
"request_id": "req-1",
"payload": {
"position_start": 0,
"input_length": 42,
"hidden_state": "tensor payload"
}
}
RUN_DECODE:
{
"type": "RUN_DECODE",
"request_id": "req-1",
"payload": {
"position": 42,
"hidden_state": "tensor payload"
}
}
HIDDEN_STATE:
{
"type": "HIDDEN_STATE",
"request_id": "req-1",
"payload": {
"hidden_state": "tensor payload"
}
}
8. Model Support
8.1 Initial Architecture
Initial implementation should support Qwen2/Qwen2.5-style decoder-only causal LMs. The first implementation does not support Qwen3.5 hybrid models with linear_attention/GatedDeltaNet layers.
Required components:
- Token embedding.
- Stacked transformer decoder blocks.
- RMSNorm.
- Rotary position embeddings.
- Grouped-query attention.
- Causal attention mask.
- Gated MLP/SwiGLU.
- Final RMSNorm.
- LM head.
- Tied or untied output embeddings.
8.2 HuggingFace Files
Coordinator should support models with:
config.json
tokenizer.json/tokenizer.model/tokenizer_config.json
model.safetensors or model-00001-of-000xx.safetensors
model.safetensors.index.json, if sharded
Use libraries:
huggingface_hubsafetensorstransformersfor tokenizer/config only where possibletorch
8.3 Weight Name Mapping
For Qwen-style models, expected tensor names include patterns like:
model.embed_tokens.weight
model.layers.{i}.input_layernorm.weight
model.layers.{i}.self_attn.q_proj.weight
model.layers.{i}.self_attn.k_proj.weight
model.layers.{i}.self_attn.v_proj.weight
model.layers.{i}.self_attn.o_proj.weight
model.layers.{i}.post_attention_layernorm.weight
model.layers.{i}.mlp.gate_proj.weight
model.layers.{i}.mlp.up_proj.weight
model.layers.{i}.mlp.down_proj.weight
model.norm.weight
lm_head.weight
The model loader should validate required tensors before accepting a model.
9. Model Planning and Splitting
The coordinator computes layer assignments from model metadata and node capacity.
9.1 Inputs
- Number of transformer layers.
- Per-layer tensor byte sizes.
- Quantization mode.
- Max nodes.
- Connected node capabilities.
- Optional target memory hint.
9.2 Rules
- Split only on full transformer layer boundaries.
- Assign contiguous layer ranges.
- Preserve layer order.
- Coordinator keeps embeddings, final norm, and lm head.
- Required nodes must be connected and loaded before generation.
- If insufficient nodes are available, API returns
cluster_not_ready.
9.3 First Planner Algorithm
Simple deterministic version:
- Compute total transformer layer bytes after quantization.
- Estimate bytes per layer.
- Determine number of partitions as
min(max_nodes, num_layers). - If node memory information is available, reduce/increase partitions so each assignment fits.
- Otherwise split evenly by layer byte size.
- Assign partitions to the first compatible ready nodes.
9.4 Future Planner Improvements
- Benchmark node speed and assign more layers to faster GPUs.
- Prefer CUDA nodes for larger shards.
- Consider network latency and bandwidth.
- Replicate small layers for resilience.
- Rebalance between generations.
10. Quantization
Quantization must work on CUDA, MPS, and CPU. Therefore the baseline implementation should avoid CUDA-only dependencies such as bitsandbytes.
Supported modes:
fp16
int8
int4
10.1 fp16
- Store weights as
torch.float16. - Compute in
float16by default. - Works on CUDA and MPS.
- CPU fallback may use
float32internally if needed.
10.2 Portable int8
Use symmetric per-output-channel weight-only quantization.
For a linear weight W shaped [out_features, in_features]:
scale[out_features] = max(abs(W[row])) / 127
qweight[row] = round(W[row] / scale[row]).clamp(-127, 127).int8
Forward path:
W_dequant = qweight.float() * scale[:, None]
y = x @ W_dequant.T
This is portable but not maximally fast.
10.3 Portable int4
Use group-wise weight-only quantization.
Suggested default group size: 128.
Store:
packed_qweight: uint8
scale: float16/float32 per group
zero_point: optional
metadata: original shape, group size, packing order
Forward path:
- Unpack int4 values.
- Dequantize to compute dtype.
- Perform normal PyTorch matmul.
This is designed for correctness and portability, not peak speed.
10.4 Quantization Timing
Preferred prototype behavior:
- Coordinator loads original safetensors.
- Nodes quantize their locally loaded assigned tensors if
int8orint4is selected. - Coordinator also quantizes/loads its own embedding/lm_head as needed.
11. Generation Algorithm
11.1 Prefill
For prompt token IDs of length N:
- Coordinator computes embeddings:
[1, N, hidden_size]. - Coordinator sends hidden state to first node with position start
0. - Each node runs its assigned layers across the full sequence.
- Each node initializes KV cache for its layers.
- Final node returns hidden state to coordinator.
- Coordinator applies final norm and lm head to the last token.
- Coordinator samples next token.
11.2 Decode
For each generated token:
- Coordinator embeds last token:
[1, 1, hidden_size]. - Coordinator sends hidden state to first node with current position.
- Each node runs one-token decode using local KV cache.
- Each node appends to its KV cache.
- Final node returns hidden state to coordinator.
- Coordinator computes logits and samples next token.
- Stop if EOS, stop sequence, or
max_tokensreached.
11.3 Sampling
Initial sampler supports:
- Greedy when
temperature == 0. - Temperature scaling.
- Top-p nucleus sampling.
- EOS handling.
- Stop strings after decoding.
Future additions:
- Top-k.
- Repetition penalty.
- Frequency/presence penalties.
- Logprobs.
12. Device Support
12.1 Device Auto Detection
Node device priority when --device auto:
- CUDA if available.
- MPS if available.
- CPU fallback.
12.2 CUDA
Use:
torch.cuda.is_available()
torch.cuda.get_device_properties(index)
torch.cuda.mem_get_info(index)
12.3 Apple MPS
Use:
torch.backends.mps.is_available()
torch.device("mps")
MPS memory reporting is limited, so allow --max-memory-gb override.
12.4 CPU
CPU is allowed for testing and fallback, but may be slow.
13. Package Structure
Recommended source tree:
truecluster/
__init__.py
cli.py
cluster/
__init__.py
server.py # node TCP server
api.py # FastAPI OpenAI-compatible API
planner.py # layer splitting
scheduler.py # generation orchestration
model_store.py # HF/safetensors loading
sampler.py
state.py
node/
__init__.py
client.py # connects to cluster
runtime.py # owns assigned layers + cache
device.py
model/
__init__.py
qwen.py # minimal Qwen implementation
layers.py
rotary.py
kv_cache.py
quant.py
tensor_names.py
protocol/
__init__.py
framing.py
messages.py
tensors.py
tests/
test_single_node_matches_transformers.py
test_protocol.py
test_quant.py
test_planner.py
Project metadata:
pyproject.toml
README.md
SPEC.md
14. Suggested Dependencies
Runtime:
torch
transformers
huggingface_hub
safetensors
fastapi
uvicorn[standard]
typer
msgpack
pydantic
numpy
tqdm
Development/test:
pytest
pytest-asyncio
httpx
ruff
mypy optional
Python version:
>=3.10,<3.13
15. Validation and Testing
15.1 Correctness Test Against Transformers
Most important validation:
- Load the target model with HuggingFace Transformers locally.
- Load the same model through TrueCluster with one local node.
- Run the same prompt.
- Compare final logits before sampling.
- Assert max difference is within tolerance for selected dtype.
15.2 Multi-node Local Test
Run on one machine:
truecluster cluster --model ... --max-nodes 2
truecluster node --cluster-host 127.0.0.1 --cluster-port 7001 --device cpu
truecluster node --cluster-host 127.0.0.1 --cluster-port 7001 --device cpu
Verify:
- Both nodes receive different layer ranges.
- Prefill works.
- Decode works.
- Output matches single-node output within tolerance.
15.3 Mixed Hardware Test
Example:
coordinator: Mac or Linux host
node 1: Nvidia CUDA machine
node 2: Apple Silicon Mac MPS machine
Verify:
- Both nodes connect.
- Assignments are sent.
- Generation completes.
15.4 Quantization Tests
For int8 and int4:
- Quantize/dequantize synthetic tensors.
- Check shape preservation.
- Check error bounds.
- Run short generation and verify no crashes.
16. Implementation Phases
Phase 1: Local Single-Process Model Proof
Deliverables:
- Minimal Qwen model implementation.
- Safetensors loading.
- Local full-model forward.
- Logit comparison against Transformers.
Phase 2: One Node Distributed Inference
Deliverables:
- Socket protocol.
- Coordinator assigns all transformer layers to one local node.
- Node loads layers and runs them.
- Coordinator keeps embedding/final norm/lm head.
/v1/completionsworks.
Phase 3: Multi-node Layer Split
Deliverables:
- Planner assigns contiguous layer ranges.
- Multiple nodes are supported.
- Distributed KV cache works.
- Not-ready errors work.
Phase 4: Mixed CUDA/MPS Support
Deliverables:
- Device detection.
- CUDA execution.
- MPS execution.
- CPU fallback.
- Mixed Nvidia/Mac cluster generation test.
Phase 5: Portable Quantization
Deliverables:
fp16baseline.- Portable
int8linear. - Portable
int4linear. - Node-local quantized layer loading.
- CLI
--quantoption.
Phase 6: API Polish
Deliverables:
/v1/models./v1/chat/completions.- Stop sequence support.
- Usage accounting.
- Better OpenAI-compatible errors.
17. Initial Acceptance Criteria
A prototype is considered working when:
- A cluster can be started with a HuggingFace safetensors Qwen-style model.
- A node can connect to the cluster with no local model files.
- The cluster sends a layer assignment to the node over the socket.
- The node downloads/resolves the model and loads assigned layers on CUDA, MPS, or CPU.
/v1/modelsreturns the loaded model./v1/completionsgenerates text through the distributed pipeline./v1/chat/completionsworks for simple chat prompts.- If insufficient nodes are ready, API returns a clear
cluster_not_readyerror. - One local-node output matches Transformers logits within reasonable dtype tolerance.
- Multi-node local CPU test works.
18. Key Design Decisions
- Use pipeline parallelism, not tensor parallelism.
- Use contiguous layer ranges only.
- Keep tokenizer, embeddings, final norm, lm head, and sampler on the coordinator.
- Do not send weights over the node socket; nodes load weights locally from HuggingFace/cache.
- Send hidden states during generation.
- Store KV cache on worker nodes.
- Implement portable quantization instead of relying on CUDA-only libraries.
- Start with Qwen2/Qwen2.5-style causal LMs only.
- Optimize for correctness and clean architecture before speed.