Files
2026-06-04 19:03:03 -05:00

22 KiB

TrueCluster Prototype Specification

1. Goal

TrueCluster is a Python 3.10 prototype for distributed LLM inference across a small cluster of heterogeneous machines. It allows a coordinator/cluster process to load a HuggingFace safetensors model config/tokenizer/coordinator-owned tensors, split transformer layers across connected worker nodes, tell each node which layer range to load from HuggingFace, and expose a simple OpenAI-compatible HTTP API for generation.

The initial target is small Qwen2/Qwen2.5-style decoder-only models around 0.5B-1.5B parameters, with first-class support for mixed Nvidia CUDA and Apple Silicon Mac MPS nodes in the same cluster. Qwen3.5 hybrid linear-attention models are a later target and are not part of the first fp16 prototype.

Primary prototype goals:

  • Python 3.10.
  • Easy CLI for running a cluster or node.
  • Coordinator loads one model at a time.
  • Nodes connect to the coordinator by host/port.
  • Coordinator sends layer assignments to nodes over sockets.
  • Nodes download/resolve the full HuggingFace model themselves, then load only their assigned layers.
  • Support HuggingFace safetensors models.
  • Initial architecture target: Qwen2/Qwen2.5-style causal language models.
  • Support fp16, portable int8, and portable int4 weight-only quantization modes.
  • OpenAI-compatible unauthenticated HTTP API.
  • Efficient generation by transferring activations during inference, not weights.

2. Non-goals for the Initial Prototype

The first prototype will intentionally avoid several advanced features:

  • No multi-model serving.
  • No authentication.
  • No request batching.
  • No tensor parallelism across machines.
  • No expert parallelism.
  • No dynamic model hot-swapping.
  • No high-performance CUDA-only quantization kernels as a requirement.
  • No coordinator-to-node weight transfer. Nodes are expected to have HuggingFace/model access.
  • No web UI.
  • No streaming responses in the first milestone.
  • No fault-tolerant recovery during an active generation.

These can be added after a correct baseline works.

3. High-Level Architecture

TrueCluster uses pipeline-parallel inference.

The coordinator owns:

  • CLI entry point for cluster.
  • HTTP API.
  • Tokenizer.
  • Sampling logic.
  • HuggingFace model metadata/config loading.
  • Coordinator-owned tensor loading from safetensors.
  • Model split planning.
  • Embedding layer.
  • Final normalization.
  • LM head.
  • Node registry and orchestration.

Each node owns:

  • CLI entry point for node.
  • Persistent socket connection to coordinator.
  • Device detection and selection.
  • One contiguous range of transformer layers.
  • KV cache for its assigned layers.
  • Local forward execution on CUDA, MPS, or CPU.

Generation path:

HTTP request
  -> coordinator tokenizes prompt
  -> coordinator runs embedding
  -> hidden states sent to node 1
  -> node 1 runs assigned layers
  -> hidden states sent to node 2
  -> ...
  -> final hidden states returned to coordinator
  -> coordinator runs final norm + lm_head
  -> coordinator samples next token
  -> repeat until completion

Weights are not transferred over the cluster socket. During assignment, the coordinator sends model id, config, and layer range. Each node downloads/resolves the model from HuggingFace or a matching local path and loads only its assigned layer tensors. During generation, only hidden states and small metadata are passed between coordinator and nodes.

4. Execution Model

4.1 Pipeline Parallelism

The model is split by complete transformer blocks. Each node receives a contiguous set of layers:

coordinator:
  embed_tokens
  final_norm
  lm_head

node 1:
  layers 0-7

node 2:
  layers 8-15

node 3:
  layers 16-23

This is simpler and more reliable than tensor parallelism for heterogeneous machines and normal Ethernet/Wi-Fi networks.

4.2 KV Cache Ownership

KV cache is stored on the node that owns the relevant layers.

For example:

node 1 cache: layers 0-7
node 2 cache: layers 8-15
node 3 cache: layers 16-23

During prefill, each node creates cache entries for its layers. During decode, each node appends one token of keys/values to its cache.

4.3 Request Concurrency

Initial prototype supports one active generation at a time per cluster.

Reason: distributed KV cache management is much easier with a single active request. Later versions can introduce request IDs, cache slots, and batching.

5. CLI Design

Use typer for the CLI.

Package command:

truecluster

5.1 Cluster Command

Example:

truecluster cluster \
  --model Qwen/Qwen3.5-0.8B \
  --node-host 0.0.0.0 \
  --node-port 7001 \
  --api-host 0.0.0.0 \
  --api-port 8000 \
  --max-nodes 4 \
  --quant fp16

Options:

--model TEXT                 HuggingFace model id or local path.
--node-host TEXT             Host/IP for worker-node socket server. Default: 0.0.0.0
--node-port INT              Port for worker-node socket server. Default: 7001
--api-host TEXT              Host/IP for HTTP API. Default: 0.0.0.0
--api-port INT               HTTP API port. Default: 8000
--max-nodes INT              Maximum number of worker nodes to use.
--quant [fp16|int8|int4]     Weight mode. Default: fp16
--dtype [fp16|bf16|fp32]     Compute dtype preference. Default: fp16
--target-node-memory-gb FLOAT Optional planning hint if node memory is unknown.
--trust-remote-code BOOL     HuggingFace trust_remote_code. Default: false
--hf-cache-dir PATH          Optional HuggingFace cache directory.
--log-level TEXT             Default: info

Behavior:

  1. Resolve/download model.
  2. Load config and tokenizer.
  3. Build safetensors metadata index and load only coordinator-owned tensors.
  4. Start node socket server.
  5. Start HTTP API.
  6. Wait for enough nodes.
  7. Assign layer ranges.
  8. Nodes download/resolve the model and load assigned layers.
  9. Mark cluster ready.

5.2 Node Command

Example:

truecluster node \
  --cluster-host 192.168.1.50 \
  --cluster-port 7001 \
  --device auto

Options:

--cluster-host TEXT          Coordinator node socket host.
--cluster-port INT           Coordinator node socket port.
--device TEXT                auto, cuda, cuda:0, mps, or cpu. Default: auto
--node-id TEXT               Optional stable node id.
--hf-cache-dir PATH          Optional HuggingFace cache directory for node downloads.
--work-dir PATH              Reserved for future temporary/cache files.
--max-memory-gb FLOAT        Optional memory capability override.
--log-level TEXT             Default: info

Behavior:

  1. Detect hardware and PyTorch backends.
  2. Connect to coordinator socket.
  3. Send HELLO capability message.
  4. Wait for assignment.
  5. Receive config and assigned layer range.
  6. Download/resolve the model locally and build local model shard.
  7. Mark itself ready.
  8. Serve prefill/decode requests over the persistent connection.

6. HTTP API

Use FastAPI and Uvicorn.

The API is unauthenticated for the prototype.

6.1 GET /v1/models

Returns the single loaded model.

Example response:

{
  "object": "list",
  "data": [
    {
      "id": "Qwen/Qwen3.5-0.8B",
      "object": "model",
      "created": 0,
      "owned_by": "truecluster"
    }
  ]
}

6.2 POST /v1/completions

Supported request fields initially:

{
  "model": "Qwen/Qwen3.5-0.8B",
  "prompt": "Hello",
  "max_tokens": 64,
  "temperature": 0.7,
  "top_p": 0.95,
  "stop": null
}

Response should be OpenAI-compatible enough for common clients:

{
  "id": "cmpl-...",
  "object": "text_completion",
  "created": 0,
  "model": "Qwen/Qwen3.5-0.8B",
  "choices": [
    {
      "text": " world",
      "index": 0,
      "logprobs": null,
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 1,
    "completion_tokens": 1,
    "total_tokens": 2
  }
}

6.3 POST /v1/chat/completions

Supported request fields initially:

{
  "model": "Qwen/Qwen3.5-0.8B",
  "messages": [
    {"role": "user", "content": "Hello"}
  ],
  "max_tokens": 64,
  "temperature": 0.7,
  "top_p": 0.95,
  "stop": null,
  "stream": false
}

The coordinator should use the HuggingFace tokenizer chat template if available.

stream: true may return a clear unsupported error in the initial prototype.

6.4 Not-Ready Error

If a generation request arrives before enough nodes are connected and loaded:

{
  "error": {
    "message": "Model is not ready. Required nodes: 3, connected ready nodes: 1",
    "type": "cluster_not_ready"
  }
}

HTTP status: 503.

7. Node Socket Protocol

Use persistent TCP sockets with asyncio streams.

Encoding:

[8-byte unsigned big-endian payload length][msgpack payload]

Large tensor payloads are sent as chunked binary data inside protocol messages or as msgpack metadata followed by raw bytes.

7.1 Message Envelope

Every message should include:

{
  "type": "MESSAGE_TYPE",
  "request_id": "optional-request-id",
  "seq": 1,
  "payload": {}
}

7.2 Core Message Types

Coordinator/node lifecycle:

HELLO
HELLO_ACK
ASSIGNMENT
LOAD_COMPLETE
LOAD_FAILED
PING
PONG
ERROR

Inference:

CLEAR_CACHE
RUN_PREFILL
RUN_DECODE
HIDDEN_STATE
INFERENCE_ERROR

7.3 HELLO

Sent by node immediately after connection.

Example:

{
  "type": "HELLO",
  "payload": {
    "node_id": "macbook-pro-1",
    "hostname": "macbook-pro.local",
    "python_version": "3.10.13",
    "torch_version": "2.x",
    "platform": "darwin",
    "devices": [
      {
        "id": "mps",
        "type": "mps",
        "name": "Apple Silicon MPS",
        "total_memory": null,
        "free_memory": null
      }
    ],
    "selected_device": "mps",
    "max_memory_bytes": null
  }
}

CUDA example device:

{
  "id": "cuda:0",
  "type": "cuda",
  "name": "NVIDIA GeForce RTX 4090",
  "total_memory": 25757220864,
  "free_memory": 23000000000
}

7.4 ASSIGNMENT

Sent by coordinator after planning.

{
  "type": "ASSIGNMENT",
  "payload": {
    "model_id": "Qwen/Qwen3.5-0.8B",
    "architecture": "qwen",
    "quant": "fp16",
    "compute_dtype": "fp16",
    "layer_start": 0,
    "layer_end_exclusive": 8,
    "config": {},
    "tensor_count": 128,
    "total_weight_bytes": 123456789
  }
}

7.5 Model Loading on Nodes

The coordinator does not transfer model weights. After ASSIGNMENT, each node resolves/downloads model_id itself using HuggingFace cache semantics and loads only tensors whose names match its assigned layer range.

7.6 Inference Messages

RUN_PREFILL:

{
  "type": "RUN_PREFILL",
  "request_id": "req-1",
  "payload": {
    "position_start": 0,
    "input_length": 42,
    "hidden_state": "tensor payload"
  }
}

RUN_DECODE:

{
  "type": "RUN_DECODE",
  "request_id": "req-1",
  "payload": {
    "position": 42,
    "hidden_state": "tensor payload"
  }
}

HIDDEN_STATE:

{
  "type": "HIDDEN_STATE",
  "request_id": "req-1",
  "payload": {
    "hidden_state": "tensor payload"
  }
}

8. Model Support

8.1 Initial Architecture

Initial implementation should support Qwen2/Qwen2.5-style decoder-only causal LMs. The first implementation does not support Qwen3.5 hybrid models with linear_attention/GatedDeltaNet layers.

Required components:

  • Token embedding.
  • Stacked transformer decoder blocks.
  • RMSNorm.
  • Rotary position embeddings.
  • Grouped-query attention.
  • Causal attention mask.
  • Gated MLP/SwiGLU.
  • Final RMSNorm.
  • LM head.
  • Tied or untied output embeddings.

8.2 HuggingFace Files

Coordinator should support models with:

config.json
tokenizer.json/tokenizer.model/tokenizer_config.json
model.safetensors or model-00001-of-000xx.safetensors
model.safetensors.index.json, if sharded

Use libraries:

  • huggingface_hub
  • safetensors
  • transformers for tokenizer/config only where possible
  • torch

8.3 Weight Name Mapping

For Qwen-style models, expected tensor names include patterns like:

model.embed_tokens.weight
model.layers.{i}.input_layernorm.weight
model.layers.{i}.self_attn.q_proj.weight
model.layers.{i}.self_attn.k_proj.weight
model.layers.{i}.self_attn.v_proj.weight
model.layers.{i}.self_attn.o_proj.weight
model.layers.{i}.post_attention_layernorm.weight
model.layers.{i}.mlp.gate_proj.weight
model.layers.{i}.mlp.up_proj.weight
model.layers.{i}.mlp.down_proj.weight
model.norm.weight
lm_head.weight

The model loader should validate required tensors before accepting a model.

9. Model Planning and Splitting

The coordinator computes layer assignments from model metadata and node capacity.

9.1 Inputs

  • Number of transformer layers.
  • Per-layer tensor byte sizes.
  • Quantization mode.
  • Max nodes.
  • Connected node capabilities.
  • Optional target memory hint.

9.2 Rules

  • Split only on full transformer layer boundaries.
  • Assign contiguous layer ranges.
  • Preserve layer order.
  • Coordinator keeps embeddings, final norm, and lm head.
  • Required nodes must be connected and loaded before generation.
  • If insufficient nodes are available, API returns cluster_not_ready.

9.3 First Planner Algorithm

Simple deterministic version:

  1. Compute total transformer layer bytes after quantization.
  2. Estimate bytes per layer.
  3. Determine number of partitions as min(max_nodes, num_layers).
  4. If node memory information is available, reduce/increase partitions so each assignment fits.
  5. Otherwise split evenly by layer byte size.
  6. Assign partitions to the first compatible ready nodes.

9.4 Future Planner Improvements

  • Benchmark node speed and assign more layers to faster GPUs.
  • Prefer CUDA nodes for larger shards.
  • Consider network latency and bandwidth.
  • Replicate small layers for resilience.
  • Rebalance between generations.

10. Quantization

Quantization must work on CUDA, MPS, and CPU. Therefore the baseline implementation should avoid CUDA-only dependencies such as bitsandbytes.

Supported modes:

fp16
int8
int4

10.1 fp16

  • Store weights as torch.float16.
  • Compute in float16 by default.
  • Works on CUDA and MPS.
  • CPU fallback may use float32 internally if needed.

10.2 Portable int8

Use symmetric per-output-channel weight-only quantization.

For a linear weight W shaped [out_features, in_features]:

scale[out_features] = max(abs(W[row])) / 127
qweight[row] = round(W[row] / scale[row]).clamp(-127, 127).int8

Forward path:

W_dequant = qweight.float() * scale[:, None]
y = x @ W_dequant.T

This is portable but not maximally fast.

10.3 Portable int4

Use group-wise weight-only quantization.

Suggested default group size: 128.

Store:

packed_qweight: uint8
scale: float16/float32 per group
zero_point: optional
metadata: original shape, group size, packing order

Forward path:

  1. Unpack int4 values.
  2. Dequantize to compute dtype.
  3. Perform normal PyTorch matmul.

This is designed for correctness and portability, not peak speed.

10.4 Quantization Timing

Preferred prototype behavior:

  • Coordinator loads original safetensors.
  • Nodes quantize their locally loaded assigned tensors if int8 or int4 is selected.
  • Coordinator also quantizes/loads its own embedding/lm_head as needed.

11. Generation Algorithm

11.1 Prefill

For prompt token IDs of length N:

  1. Coordinator computes embeddings: [1, N, hidden_size].
  2. Coordinator sends hidden state to first node with position start 0.
  3. Each node runs its assigned layers across the full sequence.
  4. Each node initializes KV cache for its layers.
  5. Final node returns hidden state to coordinator.
  6. Coordinator applies final norm and lm head to the last token.
  7. Coordinator samples next token.

11.2 Decode

For each generated token:

  1. Coordinator embeds last token: [1, 1, hidden_size].
  2. Coordinator sends hidden state to first node with current position.
  3. Each node runs one-token decode using local KV cache.
  4. Each node appends to its KV cache.
  5. Final node returns hidden state to coordinator.
  6. Coordinator computes logits and samples next token.
  7. Stop if EOS, stop sequence, or max_tokens reached.

11.3 Sampling

Initial sampler supports:

  • Greedy when temperature == 0.
  • Temperature scaling.
  • Top-p nucleus sampling.
  • EOS handling.
  • Stop strings after decoding.

Future additions:

  • Top-k.
  • Repetition penalty.
  • Frequency/presence penalties.
  • Logprobs.

12. Device Support

12.1 Device Auto Detection

Node device priority when --device auto:

  1. CUDA if available.
  2. MPS if available.
  3. CPU fallback.

12.2 CUDA

Use:

torch.cuda.is_available()
torch.cuda.get_device_properties(index)
torch.cuda.mem_get_info(index)

12.3 Apple MPS

Use:

torch.backends.mps.is_available()
torch.device("mps")

MPS memory reporting is limited, so allow --max-memory-gb override.

12.4 CPU

CPU is allowed for testing and fallback, but may be slow.

13. Package Structure

Recommended source tree:

truecluster/
  __init__.py
  cli.py

  cluster/
    __init__.py
    server.py          # node TCP server
    api.py             # FastAPI OpenAI-compatible API
    planner.py         # layer splitting
    scheduler.py       # generation orchestration
    model_store.py     # HF/safetensors loading
    sampler.py
    state.py

  node/
    __init__.py
    client.py          # connects to cluster
    runtime.py         # owns assigned layers + cache
    device.py

  model/
    __init__.py
    qwen.py            # minimal Qwen implementation
    layers.py
    rotary.py
    kv_cache.py
    quant.py
    tensor_names.py

  protocol/
    __init__.py
    framing.py
    messages.py
    tensors.py

  tests/
    test_single_node_matches_transformers.py
    test_protocol.py
    test_quant.py
    test_planner.py

Project metadata:

pyproject.toml
README.md
SPEC.md

14. Suggested Dependencies

Runtime:

torch
transformers
huggingface_hub
safetensors
fastapi
uvicorn[standard]
typer
msgpack
pydantic
numpy
tqdm

Development/test:

pytest
pytest-asyncio
httpx
ruff
mypy optional

Python version:

>=3.10,<3.13

15. Validation and Testing

15.1 Correctness Test Against Transformers

Most important validation:

  1. Load the target model with HuggingFace Transformers locally.
  2. Load the same model through TrueCluster with one local node.
  3. Run the same prompt.
  4. Compare final logits before sampling.
  5. Assert max difference is within tolerance for selected dtype.

15.2 Multi-node Local Test

Run on one machine:

truecluster cluster --model ... --max-nodes 2
truecluster node --cluster-host 127.0.0.1 --cluster-port 7001 --device cpu
truecluster node --cluster-host 127.0.0.1 --cluster-port 7001 --device cpu

Verify:

  • Both nodes receive different layer ranges.
  • Prefill works.
  • Decode works.
  • Output matches single-node output within tolerance.

15.3 Mixed Hardware Test

Example:

coordinator: Mac or Linux host
node 1: Nvidia CUDA machine
node 2: Apple Silicon Mac MPS machine

Verify:

  • Both nodes connect.
  • Assignments are sent.
  • Generation completes.

15.4 Quantization Tests

For int8 and int4:

  • Quantize/dequantize synthetic tensors.
  • Check shape preservation.
  • Check error bounds.
  • Run short generation and verify no crashes.

16. Implementation Phases

Phase 1: Local Single-Process Model Proof

Deliverables:

  • Minimal Qwen model implementation.
  • Safetensors loading.
  • Local full-model forward.
  • Logit comparison against Transformers.

Phase 2: One Node Distributed Inference

Deliverables:

  • Socket protocol.
  • Coordinator assigns all transformer layers to one local node.
  • Node loads layers and runs them.
  • Coordinator keeps embedding/final norm/lm head.
  • /v1/completions works.

Phase 3: Multi-node Layer Split

Deliverables:

  • Planner assigns contiguous layer ranges.
  • Multiple nodes are supported.
  • Distributed KV cache works.
  • Not-ready errors work.

Phase 4: Mixed CUDA/MPS Support

Deliverables:

  • Device detection.
  • CUDA execution.
  • MPS execution.
  • CPU fallback.
  • Mixed Nvidia/Mac cluster generation test.

Phase 5: Portable Quantization

Deliverables:

  • fp16 baseline.
  • Portable int8 linear.
  • Portable int4 linear.
  • Node-local quantized layer loading.
  • CLI --quant option.

Phase 6: API Polish

Deliverables:

  • /v1/models.
  • /v1/chat/completions.
  • Stop sequence support.
  • Usage accounting.
  • Better OpenAI-compatible errors.

17. Initial Acceptance Criteria

A prototype is considered working when:

  1. A cluster can be started with a HuggingFace safetensors Qwen-style model.
  2. A node can connect to the cluster with no local model files.
  3. The cluster sends a layer assignment to the node over the socket.
  4. The node downloads/resolves the model and loads assigned layers on CUDA, MPS, or CPU.
  5. /v1/models returns the loaded model.
  6. /v1/completions generates text through the distributed pipeline.
  7. /v1/chat/completions works for simple chat prompts.
  8. If insufficient nodes are ready, API returns a clear cluster_not_ready error.
  9. One local-node output matches Transformers logits within reasonable dtype tolerance.
  10. Multi-node local CPU test works.

18. Key Design Decisions

  • Use pipeline parallelism, not tensor parallelism.
  • Use contiguous layer ranges only.
  • Keep tokenizer, embeddings, final norm, lm head, and sampler on the coordinator.
  • Do not send weights over the node socket; nodes load weights locally from HuggingFace/cache.
  • Send hidden states during generation.
  • Store KV cache on worker nodes.
  • Implement portable quantization instead of relying on CUDA-only libraries.
  • Start with Qwen2/Qwen2.5-style causal LMs only.
  • Optimize for correctness and clean architecture before speed.