# TrueCluster Prototype Specification ## 1. Goal TrueCluster is a Python 3.10 prototype for distributed LLM inference across a small cluster of heterogeneous machines. It allows a coordinator/cluster process to load a HuggingFace safetensors model config/tokenizer/coordinator-owned tensors, split transformer layers across connected worker nodes, tell each node which layer range to load from HuggingFace, and expose a simple OpenAI-compatible HTTP API for generation. The initial target is small Qwen2/Qwen2.5-style decoder-only models around 0.5B-1.5B parameters, with first-class support for mixed Nvidia CUDA and Apple Silicon Mac MPS nodes in the same cluster. Qwen3.5 hybrid linear-attention models are a later target and are not part of the first fp16 prototype. Primary prototype goals: - Python 3.10. - Easy CLI for running a cluster or node. - Coordinator loads one model at a time. - Nodes connect to the coordinator by host/port. - Coordinator sends layer assignments to nodes over sockets. - Nodes download/resolve the full HuggingFace model themselves, then load only their assigned layers. - Support HuggingFace safetensors models. - Initial architecture target: Qwen2/Qwen2.5-style causal language models. - Support `fp16`, portable `int8`, and portable `int4` weight-only quantization modes. - OpenAI-compatible unauthenticated HTTP API. - Efficient generation by transferring activations during inference, not weights. ## 2. Non-goals for the Initial Prototype The first prototype will intentionally avoid several advanced features: - No multi-model serving. - No authentication. - No request batching. - No tensor parallelism across machines. - No expert parallelism. - No dynamic model hot-swapping. - No high-performance CUDA-only quantization kernels as a requirement. - No coordinator-to-node weight transfer. Nodes are expected to have HuggingFace/model access. - No web UI. - No streaming responses in the first milestone. - No fault-tolerant recovery during an active generation. These can be added after a correct baseline works. ## 3. High-Level Architecture TrueCluster uses pipeline-parallel inference. The coordinator owns: - CLI entry point for `cluster`. - HTTP API. - Tokenizer. - Sampling logic. - HuggingFace model metadata/config loading. - Coordinator-owned tensor loading from safetensors. - Model split planning. - Embedding layer. - Final normalization. - LM head. - Node registry and orchestration. Each node owns: - CLI entry point for `node`. - Persistent socket connection to coordinator. - Device detection and selection. - One contiguous range of transformer layers. - KV cache for its assigned layers. - Local forward execution on CUDA, MPS, or CPU. Generation path: ```text HTTP request -> coordinator tokenizes prompt -> coordinator runs embedding -> hidden states sent to node 1 -> node 1 runs assigned layers -> hidden states sent to node 2 -> ... -> final hidden states returned to coordinator -> coordinator runs final norm + lm_head -> coordinator samples next token -> repeat until completion ``` Weights are not transferred over the cluster socket. During assignment, the coordinator sends model id, config, and layer range. Each node downloads/resolves the model from HuggingFace or a matching local path and loads only its assigned layer tensors. During generation, only hidden states and small metadata are passed between coordinator and nodes. ## 4. Execution Model ### 4.1 Pipeline Parallelism The model is split by complete transformer blocks. Each node receives a contiguous set of layers: ```text coordinator: embed_tokens final_norm lm_head node 1: layers 0-7 node 2: layers 8-15 node 3: layers 16-23 ``` This is simpler and more reliable than tensor parallelism for heterogeneous machines and normal Ethernet/Wi-Fi networks. ### 4.2 KV Cache Ownership KV cache is stored on the node that owns the relevant layers. For example: ```text node 1 cache: layers 0-7 node 2 cache: layers 8-15 node 3 cache: layers 16-23 ``` During prefill, each node creates cache entries for its layers. During decode, each node appends one token of keys/values to its cache. ### 4.3 Request Concurrency Initial prototype supports one active generation at a time per cluster. Reason: distributed KV cache management is much easier with a single active request. Later versions can introduce request IDs, cache slots, and batching. ## 5. CLI Design Use `typer` for the CLI. Package command: ```bash truecluster ``` ### 5.1 Cluster Command Example: ```bash truecluster cluster \ --model Qwen/Qwen3.5-0.8B \ --node-host 0.0.0.0 \ --node-port 7001 \ --api-host 0.0.0.0 \ --api-port 8000 \ --max-nodes 4 \ --quant fp16 ``` Options: ```text --model TEXT HuggingFace model id or local path. --node-host TEXT Host/IP for worker-node socket server. Default: 0.0.0.0 --node-port INT Port for worker-node socket server. Default: 7001 --api-host TEXT Host/IP for HTTP API. Default: 0.0.0.0 --api-port INT HTTP API port. Default: 8000 --max-nodes INT Maximum number of worker nodes to use. --quant [fp16|int8|int4] Weight mode. Default: fp16 --dtype [fp16|bf16|fp32] Compute dtype preference. Default: fp16 --target-node-memory-gb FLOAT Optional planning hint if node memory is unknown. --trust-remote-code BOOL HuggingFace trust_remote_code. Default: false --hf-cache-dir PATH Optional HuggingFace cache directory. --log-level TEXT Default: info ``` Behavior: 1. Resolve/download model. 2. Load config and tokenizer. 3. Build safetensors metadata index and load only coordinator-owned tensors. 4. Start node socket server. 5. Start HTTP API. 6. Wait for enough nodes. 7. Assign layer ranges. 8. Nodes download/resolve the model and load assigned layers. 9. Mark cluster ready. ### 5.2 Node Command Example: ```bash truecluster node \ --cluster-host 192.168.1.50 \ --cluster-port 7001 \ --device auto ``` Options: ```text --cluster-host TEXT Coordinator node socket host. --cluster-port INT Coordinator node socket port. --device TEXT auto, cuda, cuda:0, mps, or cpu. Default: auto --node-id TEXT Optional stable node id. --hf-cache-dir PATH Optional HuggingFace cache directory for node downloads. --work-dir PATH Reserved for future temporary/cache files. --max-memory-gb FLOAT Optional memory capability override. --log-level TEXT Default: info ``` Behavior: 1. Detect hardware and PyTorch backends. 2. Connect to coordinator socket. 3. Send `HELLO` capability message. 4. Wait for assignment. 5. Receive config and assigned layer range. 6. Download/resolve the model locally and build local model shard. 7. Mark itself ready. 8. Serve prefill/decode requests over the persistent connection. ## 6. HTTP API Use FastAPI and Uvicorn. The API is unauthenticated for the prototype. ### 6.1 `GET /v1/models` Returns the single loaded model. Example response: ```json { "object": "list", "data": [ { "id": "Qwen/Qwen3.5-0.8B", "object": "model", "created": 0, "owned_by": "truecluster" } ] } ``` ### 6.2 `POST /v1/completions` Supported request fields initially: ```json { "model": "Qwen/Qwen3.5-0.8B", "prompt": "Hello", "max_tokens": 64, "temperature": 0.7, "top_p": 0.95, "stop": null } ``` Response should be OpenAI-compatible enough for common clients: ```json { "id": "cmpl-...", "object": "text_completion", "created": 0, "model": "Qwen/Qwen3.5-0.8B", "choices": [ { "text": " world", "index": 0, "logprobs": null, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 1, "completion_tokens": 1, "total_tokens": 2 } } ``` ### 6.3 `POST /v1/chat/completions` Supported request fields initially: ```json { "model": "Qwen/Qwen3.5-0.8B", "messages": [ {"role": "user", "content": "Hello"} ], "max_tokens": 64, "temperature": 0.7, "top_p": 0.95, "stop": null, "stream": false } ``` The coordinator should use the HuggingFace tokenizer chat template if available. `stream: true` may return a clear unsupported error in the initial prototype. ### 6.4 Not-Ready Error If a generation request arrives before enough nodes are connected and loaded: ```json { "error": { "message": "Model is not ready. Required nodes: 3, connected ready nodes: 1", "type": "cluster_not_ready" } } ``` HTTP status: `503`. ## 7. Node Socket Protocol Use persistent TCP sockets with asyncio streams. Encoding: ```text [8-byte unsigned big-endian payload length][msgpack payload] ``` Large tensor payloads are sent as chunked binary data inside protocol messages or as msgpack metadata followed by raw bytes. ### 7.1 Message Envelope Every message should include: ```json { "type": "MESSAGE_TYPE", "request_id": "optional-request-id", "seq": 1, "payload": {} } ``` ### 7.2 Core Message Types Coordinator/node lifecycle: ```text HELLO HELLO_ACK ASSIGNMENT LOAD_COMPLETE LOAD_FAILED PING PONG ERROR ``` Inference: ```text CLEAR_CACHE RUN_PREFILL RUN_DECODE HIDDEN_STATE INFERENCE_ERROR ``` ### 7.3 `HELLO` Sent by node immediately after connection. Example: ```json { "type": "HELLO", "payload": { "node_id": "macbook-pro-1", "hostname": "macbook-pro.local", "python_version": "3.10.13", "torch_version": "2.x", "platform": "darwin", "devices": [ { "id": "mps", "type": "mps", "name": "Apple Silicon MPS", "total_memory": null, "free_memory": null } ], "selected_device": "mps", "max_memory_bytes": null } } ``` CUDA example device: ```json { "id": "cuda:0", "type": "cuda", "name": "NVIDIA GeForce RTX 4090", "total_memory": 25757220864, "free_memory": 23000000000 } ``` ### 7.4 `ASSIGNMENT` Sent by coordinator after planning. ```json { "type": "ASSIGNMENT", "payload": { "model_id": "Qwen/Qwen3.5-0.8B", "architecture": "qwen", "quant": "fp16", "compute_dtype": "fp16", "layer_start": 0, "layer_end_exclusive": 8, "config": {}, "tensor_count": 128, "total_weight_bytes": 123456789 } } ``` ### 7.5 Model Loading on Nodes The coordinator does not transfer model weights. After `ASSIGNMENT`, each node resolves/downloads `model_id` itself using HuggingFace cache semantics and loads only tensors whose names match its assigned layer range. ### 7.6 Inference Messages `RUN_PREFILL`: ```json { "type": "RUN_PREFILL", "request_id": "req-1", "payload": { "position_start": 0, "input_length": 42, "hidden_state": "tensor payload" } } ``` `RUN_DECODE`: ```json { "type": "RUN_DECODE", "request_id": "req-1", "payload": { "position": 42, "hidden_state": "tensor payload" } } ``` `HIDDEN_STATE`: ```json { "type": "HIDDEN_STATE", "request_id": "req-1", "payload": { "hidden_state": "tensor payload" } } ``` ## 8. Model Support ### 8.1 Initial Architecture Initial implementation should support Qwen2/Qwen2.5-style decoder-only causal LMs. The first implementation does not support Qwen3.5 hybrid models with `linear_attention`/GatedDeltaNet layers. Required components: - Token embedding. - Stacked transformer decoder blocks. - RMSNorm. - Rotary position embeddings. - Grouped-query attention. - Causal attention mask. - Gated MLP/SwiGLU. - Final RMSNorm. - LM head. - Tied or untied output embeddings. ### 8.2 HuggingFace Files Coordinator should support models with: ```text config.json tokenizer.json/tokenizer.model/tokenizer_config.json model.safetensors or model-00001-of-000xx.safetensors model.safetensors.index.json, if sharded ``` Use libraries: - `huggingface_hub` - `safetensors` - `transformers` for tokenizer/config only where possible - `torch` ### 8.3 Weight Name Mapping For Qwen-style models, expected tensor names include patterns like: ```text model.embed_tokens.weight model.layers.{i}.input_layernorm.weight model.layers.{i}.self_attn.q_proj.weight model.layers.{i}.self_attn.k_proj.weight model.layers.{i}.self_attn.v_proj.weight model.layers.{i}.self_attn.o_proj.weight model.layers.{i}.post_attention_layernorm.weight model.layers.{i}.mlp.gate_proj.weight model.layers.{i}.mlp.up_proj.weight model.layers.{i}.mlp.down_proj.weight model.norm.weight lm_head.weight ``` The model loader should validate required tensors before accepting a model. ## 9. Model Planning and Splitting The coordinator computes layer assignments from model metadata and node capacity. ### 9.1 Inputs - Number of transformer layers. - Per-layer tensor byte sizes. - Quantization mode. - Max nodes. - Connected node capabilities. - Optional target memory hint. ### 9.2 Rules - Split only on full transformer layer boundaries. - Assign contiguous layer ranges. - Preserve layer order. - Coordinator keeps embeddings, final norm, and lm head. - Required nodes must be connected and loaded before generation. - If insufficient nodes are available, API returns `cluster_not_ready`. ### 9.3 First Planner Algorithm Simple deterministic version: 1. Compute total transformer layer bytes after quantization. 2. Estimate bytes per layer. 3. Determine number of partitions as `min(max_nodes, num_layers)`. 4. If node memory information is available, reduce/increase partitions so each assignment fits. 5. Otherwise split evenly by layer byte size. 6. Assign partitions to the first compatible ready nodes. ### 9.4 Future Planner Improvements - Benchmark node speed and assign more layers to faster GPUs. - Prefer CUDA nodes for larger shards. - Consider network latency and bandwidth. - Replicate small layers for resilience. - Rebalance between generations. ## 10. Quantization Quantization must work on CUDA, MPS, and CPU. Therefore the baseline implementation should avoid CUDA-only dependencies such as bitsandbytes. Supported modes: ```text fp16 int8 int4 ``` ### 10.1 `fp16` - Store weights as `torch.float16`. - Compute in `float16` by default. - Works on CUDA and MPS. - CPU fallback may use `float32` internally if needed. ### 10.2 Portable `int8` Use symmetric per-output-channel weight-only quantization. For a linear weight `W` shaped `[out_features, in_features]`: ```text scale[out_features] = max(abs(W[row])) / 127 qweight[row] = round(W[row] / scale[row]).clamp(-127, 127).int8 ``` Forward path: ```text W_dequant = qweight.float() * scale[:, None] y = x @ W_dequant.T ``` This is portable but not maximally fast. ### 10.3 Portable `int4` Use group-wise weight-only quantization. Suggested default group size: `128`. Store: ```text packed_qweight: uint8 scale: float16/float32 per group zero_point: optional metadata: original shape, group size, packing order ``` Forward path: 1. Unpack int4 values. 2. Dequantize to compute dtype. 3. Perform normal PyTorch matmul. This is designed for correctness and portability, not peak speed. ### 10.4 Quantization Timing Preferred prototype behavior: - Coordinator loads original safetensors. - Nodes quantize their locally loaded assigned tensors if `int8` or `int4` is selected. - Coordinator also quantizes/loads its own embedding/lm_head as needed. ## 11. Generation Algorithm ### 11.1 Prefill For prompt token IDs of length `N`: 1. Coordinator computes embeddings: `[1, N, hidden_size]`. 2. Coordinator sends hidden state to first node with position start `0`. 3. Each node runs its assigned layers across the full sequence. 4. Each node initializes KV cache for its layers. 5. Final node returns hidden state to coordinator. 6. Coordinator applies final norm and lm head to the last token. 7. Coordinator samples next token. ### 11.2 Decode For each generated token: 1. Coordinator embeds last token: `[1, 1, hidden_size]`. 2. Coordinator sends hidden state to first node with current position. 3. Each node runs one-token decode using local KV cache. 4. Each node appends to its KV cache. 5. Final node returns hidden state to coordinator. 6. Coordinator computes logits and samples next token. 7. Stop if EOS, stop sequence, or `max_tokens` reached. ### 11.3 Sampling Initial sampler supports: - Greedy when `temperature == 0`. - Temperature scaling. - Top-p nucleus sampling. - EOS handling. - Stop strings after decoding. Future additions: - Top-k. - Repetition penalty. - Frequency/presence penalties. - Logprobs. ## 12. Device Support ### 12.1 Device Auto Detection Node device priority when `--device auto`: 1. CUDA if available. 2. MPS if available. 3. CPU fallback. ### 12.2 CUDA Use: ```python torch.cuda.is_available() torch.cuda.get_device_properties(index) torch.cuda.mem_get_info(index) ``` ### 12.3 Apple MPS Use: ```python torch.backends.mps.is_available() torch.device("mps") ``` MPS memory reporting is limited, so allow `--max-memory-gb` override. ### 12.4 CPU CPU is allowed for testing and fallback, but may be slow. ## 13. Package Structure Recommended source tree: ```text truecluster/ __init__.py cli.py cluster/ __init__.py server.py # node TCP server api.py # FastAPI OpenAI-compatible API planner.py # layer splitting scheduler.py # generation orchestration model_store.py # HF/safetensors loading sampler.py state.py node/ __init__.py client.py # connects to cluster runtime.py # owns assigned layers + cache device.py model/ __init__.py qwen.py # minimal Qwen implementation layers.py rotary.py kv_cache.py quant.py tensor_names.py protocol/ __init__.py framing.py messages.py tensors.py tests/ test_single_node_matches_transformers.py test_protocol.py test_quant.py test_planner.py ``` Project metadata: ```text pyproject.toml README.md SPEC.md ``` ## 14. Suggested Dependencies Runtime: ```text torch transformers huggingface_hub safetensors fastapi uvicorn[standard] typer msgpack pydantic numpy tqdm ``` Development/test: ```text pytest pytest-asyncio httpx ruff mypy optional ``` Python version: ```text >=3.10,<3.13 ``` ## 15. Validation and Testing ### 15.1 Correctness Test Against Transformers Most important validation: 1. Load the target model with HuggingFace Transformers locally. 2. Load the same model through TrueCluster with one local node. 3. Run the same prompt. 4. Compare final logits before sampling. 5. Assert max difference is within tolerance for selected dtype. ### 15.2 Multi-node Local Test Run on one machine: ```bash truecluster cluster --model ... --max-nodes 2 truecluster node --cluster-host 127.0.0.1 --cluster-port 7001 --device cpu truecluster node --cluster-host 127.0.0.1 --cluster-port 7001 --device cpu ``` Verify: - Both nodes receive different layer ranges. - Prefill works. - Decode works. - Output matches single-node output within tolerance. ### 15.3 Mixed Hardware Test Example: ```text coordinator: Mac or Linux host node 1: Nvidia CUDA machine node 2: Apple Silicon Mac MPS machine ``` Verify: - Both nodes connect. - Assignments are sent. - Generation completes. ### 15.4 Quantization Tests For `int8` and `int4`: - Quantize/dequantize synthetic tensors. - Check shape preservation. - Check error bounds. - Run short generation and verify no crashes. ## 16. Implementation Phases ### Phase 1: Local Single-Process Model Proof Deliverables: - Minimal Qwen model implementation. - Safetensors loading. - Local full-model forward. - Logit comparison against Transformers. ### Phase 2: One Node Distributed Inference Deliverables: - Socket protocol. - Coordinator assigns all transformer layers to one local node. - Node loads layers and runs them. - Coordinator keeps embedding/final norm/lm head. - `/v1/completions` works. ### Phase 3: Multi-node Layer Split Deliverables: - Planner assigns contiguous layer ranges. - Multiple nodes are supported. - Distributed KV cache works. - Not-ready errors work. ### Phase 4: Mixed CUDA/MPS Support Deliverables: - Device detection. - CUDA execution. - MPS execution. - CPU fallback. - Mixed Nvidia/Mac cluster generation test. ### Phase 5: Portable Quantization Deliverables: - `fp16` baseline. - Portable `int8` linear. - Portable `int4` linear. - Node-local quantized layer loading. - CLI `--quant` option. ### Phase 6: API Polish Deliverables: - `/v1/models`. - `/v1/chat/completions`. - Stop sequence support. - Usage accounting. - Better OpenAI-compatible errors. ## 17. Initial Acceptance Criteria A prototype is considered working when: 1. A cluster can be started with a HuggingFace safetensors Qwen-style model. 2. A node can connect to the cluster with no local model files. 3. The cluster sends a layer assignment to the node over the socket. 4. The node downloads/resolves the model and loads assigned layers on CUDA, MPS, or CPU. 5. `/v1/models` returns the loaded model. 6. `/v1/completions` generates text through the distributed pipeline. 7. `/v1/chat/completions` works for simple chat prompts. 8. If insufficient nodes are ready, API returns a clear `cluster_not_ready` error. 9. One local-node output matches Transformers logits within reasonable dtype tolerance. 10. Multi-node local CPU test works. ## 18. Key Design Decisions - Use pipeline parallelism, not tensor parallelism. - Use contiguous layer ranges only. - Keep tokenizer, embeddings, final norm, lm head, and sampler on the coordinator. - Do not send weights over the node socket; nodes load weights locally from HuggingFace/cache. - Send hidden states during generation. - Store KV cache on worker nodes. - Implement portable quantization instead of relying on CUDA-only libraries. - Start with Qwen2/Qwen2.5-style causal LMs only. - Optimize for correctness and clean architecture before speed.