980 lines
21 KiB
Markdown
980 lines
21 KiB
Markdown
# TrueCluster Prototype Specification
|
|
|
|
## 1. Goal
|
|
|
|
TrueCluster is a Python 3.10 prototype for distributed LLM inference across a small cluster of heterogeneous machines. It allows a coordinator/cluster process to load a HuggingFace safetensors model, split transformer layers across connected worker nodes, send each node only the weights it needs over a socket connection, and expose a simple OpenAI-compatible HTTP API for generation.
|
|
|
|
The initial target is small Qwen2/Qwen2.5-style decoder-only models around 0.5B-1.5B parameters, with first-class support for mixed Nvidia CUDA and Apple Silicon Mac MPS nodes in the same cluster. Qwen3.5 hybrid linear-attention models are a later target and are not part of the first fp16 prototype.
|
|
|
|
Primary prototype goals:
|
|
|
|
- Python 3.10.
|
|
- Easy CLI for running a cluster or node.
|
|
- Coordinator loads one model at a time.
|
|
- Nodes connect to the coordinator by host/port.
|
|
- Coordinator sends assigned model shards to nodes over sockets.
|
|
- Nodes do not need local model files.
|
|
- Support HuggingFace safetensors models.
|
|
- Initial architecture target: Qwen2/Qwen2.5-style causal language models.
|
|
- Support `fp16`, portable `int8`, and portable `int4` weight-only quantization modes.
|
|
- OpenAI-compatible unauthenticated HTTP API.
|
|
- Efficient generation by transferring activations during inference, not weights.
|
|
|
|
## 2. Non-goals for the Initial Prototype
|
|
|
|
The first prototype will intentionally avoid several advanced features:
|
|
|
|
- No multi-model serving.
|
|
- No authentication.
|
|
- No request batching.
|
|
- No tensor parallelism across machines.
|
|
- No expert parallelism.
|
|
- No dynamic model hot-swapping.
|
|
- No high-performance CUDA-only quantization kernels as a requirement.
|
|
- No dependency on every node having HuggingFace access.
|
|
- No web UI.
|
|
- No streaming responses in the first milestone.
|
|
- No fault-tolerant recovery during an active generation.
|
|
|
|
These can be added after a correct baseline works.
|
|
|
|
## 3. High-Level Architecture
|
|
|
|
TrueCluster uses pipeline-parallel inference.
|
|
|
|
The coordinator owns:
|
|
|
|
- CLI entry point for `cluster`.
|
|
- HTTP API.
|
|
- Tokenizer.
|
|
- Sampling logic.
|
|
- HuggingFace model metadata/config loading.
|
|
- Model weight loading from safetensors.
|
|
- Model split planning.
|
|
- Embedding layer.
|
|
- Final normalization.
|
|
- LM head.
|
|
- Node registry and orchestration.
|
|
|
|
Each node owns:
|
|
|
|
- CLI entry point for `node`.
|
|
- Persistent socket connection to coordinator.
|
|
- Device detection and selection.
|
|
- One contiguous range of transformer layers.
|
|
- KV cache for its assigned layers.
|
|
- Local forward execution on CUDA, MPS, or CPU.
|
|
|
|
Generation path:
|
|
|
|
```text
|
|
HTTP request
|
|
-> coordinator tokenizes prompt
|
|
-> coordinator runs embedding
|
|
-> hidden states sent to node 1
|
|
-> node 1 runs assigned layers
|
|
-> hidden states sent to node 2
|
|
-> ...
|
|
-> final hidden states returned to coordinator
|
|
-> coordinator runs final norm + lm_head
|
|
-> coordinator samples next token
|
|
-> repeat until completion
|
|
```
|
|
|
|
Weights are transferred once during assignment. During generation, only hidden states and small metadata are passed between coordinator and nodes.
|
|
|
|
## 4. Execution Model
|
|
|
|
### 4.1 Pipeline Parallelism
|
|
|
|
The model is split by complete transformer blocks. Each node receives a contiguous set of layers:
|
|
|
|
```text
|
|
coordinator:
|
|
embed_tokens
|
|
final_norm
|
|
lm_head
|
|
|
|
node 1:
|
|
layers 0-7
|
|
|
|
node 2:
|
|
layers 8-15
|
|
|
|
node 3:
|
|
layers 16-23
|
|
```
|
|
|
|
This is simpler and more reliable than tensor parallelism for heterogeneous machines and normal Ethernet/Wi-Fi networks.
|
|
|
|
### 4.2 KV Cache Ownership
|
|
|
|
KV cache is stored on the node that owns the relevant layers.
|
|
|
|
For example:
|
|
|
|
```text
|
|
node 1 cache: layers 0-7
|
|
node 2 cache: layers 8-15
|
|
node 3 cache: layers 16-23
|
|
```
|
|
|
|
During prefill, each node creates cache entries for its layers. During decode, each node appends one token of keys/values to its cache.
|
|
|
|
### 4.3 Request Concurrency
|
|
|
|
Initial prototype supports one active generation at a time per cluster.
|
|
|
|
Reason: distributed KV cache management is much easier with a single active request. Later versions can introduce request IDs, cache slots, and batching.
|
|
|
|
## 5. CLI Design
|
|
|
|
Use `typer` for the CLI.
|
|
|
|
Package command:
|
|
|
|
```bash
|
|
truecluster
|
|
```
|
|
|
|
### 5.1 Cluster Command
|
|
|
|
Example:
|
|
|
|
```bash
|
|
truecluster cluster \
|
|
--model Qwen/Qwen3.5-0.8B \
|
|
--node-host 0.0.0.0 \
|
|
--node-port 7001 \
|
|
--api-host 0.0.0.0 \
|
|
--api-port 8000 \
|
|
--max-nodes 4 \
|
|
--quant fp16
|
|
```
|
|
|
|
Options:
|
|
|
|
```text
|
|
--model TEXT HuggingFace model id or local path.
|
|
--node-host TEXT Host/IP for worker-node socket server. Default: 0.0.0.0
|
|
--node-port INT Port for worker-node socket server. Default: 7001
|
|
--api-host TEXT Host/IP for HTTP API. Default: 0.0.0.0
|
|
--api-port INT HTTP API port. Default: 8000
|
|
--max-nodes INT Maximum number of worker nodes to use.
|
|
--quant [fp16|int8|int4] Weight mode. Default: fp16
|
|
--dtype [fp16|bf16|fp32] Compute dtype preference. Default: fp16
|
|
--target-node-memory-gb FLOAT Optional planning hint if node memory is unknown.
|
|
--trust-remote-code BOOL HuggingFace trust_remote_code. Default: false
|
|
--hf-cache-dir PATH Optional HuggingFace cache directory.
|
|
--log-level TEXT Default: info
|
|
```
|
|
|
|
Behavior:
|
|
|
|
1. Resolve/download model.
|
|
2. Load config and tokenizer.
|
|
3. Load safetensors into coordinator RAM or memory-mapped index.
|
|
4. Build model tensor index.
|
|
5. Start node socket server.
|
|
6. Start HTTP API.
|
|
7. Wait for enough nodes.
|
|
8. Assign layer ranges and transmit weights.
|
|
9. Mark cluster ready.
|
|
|
|
### 5.2 Node Command
|
|
|
|
Example:
|
|
|
|
```bash
|
|
truecluster node \
|
|
--cluster-host 192.168.1.50 \
|
|
--cluster-port 7001 \
|
|
--device auto
|
|
```
|
|
|
|
Options:
|
|
|
|
```text
|
|
--cluster-host TEXT Coordinator node socket host.
|
|
--cluster-port INT Coordinator node socket port.
|
|
--device TEXT auto, cuda, cuda:0, mps, or cpu. Default: auto
|
|
--node-id TEXT Optional stable node id.
|
|
--work-dir PATH Temporary local directory for received weights/cache.
|
|
--max-memory-gb FLOAT Optional memory capability override.
|
|
--log-level TEXT Default: info
|
|
```
|
|
|
|
Behavior:
|
|
|
|
1. Detect hardware and PyTorch backends.
|
|
2. Connect to coordinator socket.
|
|
3. Send `HELLO` capability message.
|
|
4. Wait for assignment.
|
|
5. Receive config and layer weights.
|
|
6. Build local model shard.
|
|
7. Mark itself ready.
|
|
8. Serve prefill/decode requests over the persistent connection.
|
|
|
|
## 6. HTTP API
|
|
|
|
Use FastAPI and Uvicorn.
|
|
|
|
The API is unauthenticated for the prototype.
|
|
|
|
### 6.1 `GET /v1/models`
|
|
|
|
Returns the single loaded model.
|
|
|
|
Example response:
|
|
|
|
```json
|
|
{
|
|
"object": "list",
|
|
"data": [
|
|
{
|
|
"id": "Qwen/Qwen3.5-0.8B",
|
|
"object": "model",
|
|
"created": 0,
|
|
"owned_by": "truecluster"
|
|
}
|
|
]
|
|
}
|
|
```
|
|
|
|
### 6.2 `POST /v1/completions`
|
|
|
|
Supported request fields initially:
|
|
|
|
```json
|
|
{
|
|
"model": "Qwen/Qwen3.5-0.8B",
|
|
"prompt": "Hello",
|
|
"max_tokens": 64,
|
|
"temperature": 0.7,
|
|
"top_p": 0.95,
|
|
"stop": null
|
|
}
|
|
```
|
|
|
|
Response should be OpenAI-compatible enough for common clients:
|
|
|
|
```json
|
|
{
|
|
"id": "cmpl-...",
|
|
"object": "text_completion",
|
|
"created": 0,
|
|
"model": "Qwen/Qwen3.5-0.8B",
|
|
"choices": [
|
|
{
|
|
"text": " world",
|
|
"index": 0,
|
|
"logprobs": null,
|
|
"finish_reason": "stop"
|
|
}
|
|
],
|
|
"usage": {
|
|
"prompt_tokens": 1,
|
|
"completion_tokens": 1,
|
|
"total_tokens": 2
|
|
}
|
|
}
|
|
```
|
|
|
|
### 6.3 `POST /v1/chat/completions`
|
|
|
|
Supported request fields initially:
|
|
|
|
```json
|
|
{
|
|
"model": "Qwen/Qwen3.5-0.8B",
|
|
"messages": [
|
|
{"role": "user", "content": "Hello"}
|
|
],
|
|
"max_tokens": 64,
|
|
"temperature": 0.7,
|
|
"top_p": 0.95,
|
|
"stop": null,
|
|
"stream": false
|
|
}
|
|
```
|
|
|
|
The coordinator should use the HuggingFace tokenizer chat template if available.
|
|
|
|
`stream: true` may return a clear unsupported error in the initial prototype.
|
|
|
|
### 6.4 Not-Ready Error
|
|
|
|
If a generation request arrives before enough nodes are connected and loaded:
|
|
|
|
```json
|
|
{
|
|
"error": {
|
|
"message": "Model is not ready. Required nodes: 3, connected ready nodes: 1",
|
|
"type": "cluster_not_ready"
|
|
}
|
|
}
|
|
```
|
|
|
|
HTTP status: `503`.
|
|
|
|
## 7. Node Socket Protocol
|
|
|
|
Use persistent TCP sockets with asyncio streams.
|
|
|
|
Encoding:
|
|
|
|
```text
|
|
[8-byte unsigned big-endian payload length][msgpack payload]
|
|
```
|
|
|
|
Large tensor payloads are sent as chunked binary data inside protocol messages or as msgpack metadata followed by raw bytes.
|
|
|
|
### 7.1 Message Envelope
|
|
|
|
Every message should include:
|
|
|
|
```json
|
|
{
|
|
"type": "MESSAGE_TYPE",
|
|
"request_id": "optional-request-id",
|
|
"seq": 1,
|
|
"payload": {}
|
|
}
|
|
```
|
|
|
|
### 7.2 Core Message Types
|
|
|
|
Coordinator/node lifecycle:
|
|
|
|
```text
|
|
HELLO
|
|
HELLO_ACK
|
|
ASSIGNMENT
|
|
WEIGHT_CHUNK
|
|
WEIGHTS_COMPLETE
|
|
LOAD_COMPLETE
|
|
LOAD_FAILED
|
|
PING
|
|
PONG
|
|
ERROR
|
|
```
|
|
|
|
Inference:
|
|
|
|
```text
|
|
CLEAR_CACHE
|
|
RUN_PREFILL
|
|
RUN_DECODE
|
|
HIDDEN_STATE
|
|
INFERENCE_ERROR
|
|
```
|
|
|
|
### 7.3 `HELLO`
|
|
|
|
Sent by node immediately after connection.
|
|
|
|
Example:
|
|
|
|
```json
|
|
{
|
|
"type": "HELLO",
|
|
"payload": {
|
|
"node_id": "macbook-pro-1",
|
|
"hostname": "macbook-pro.local",
|
|
"python_version": "3.10.13",
|
|
"torch_version": "2.x",
|
|
"platform": "darwin",
|
|
"devices": [
|
|
{
|
|
"id": "mps",
|
|
"type": "mps",
|
|
"name": "Apple Silicon MPS",
|
|
"total_memory": null,
|
|
"free_memory": null
|
|
}
|
|
],
|
|
"selected_device": "mps",
|
|
"max_memory_bytes": null
|
|
}
|
|
}
|
|
```
|
|
|
|
CUDA example device:
|
|
|
|
```json
|
|
{
|
|
"id": "cuda:0",
|
|
"type": "cuda",
|
|
"name": "NVIDIA GeForce RTX 4090",
|
|
"total_memory": 25757220864,
|
|
"free_memory": 23000000000
|
|
}
|
|
```
|
|
|
|
### 7.4 `ASSIGNMENT`
|
|
|
|
Sent by coordinator after planning.
|
|
|
|
```json
|
|
{
|
|
"type": "ASSIGNMENT",
|
|
"payload": {
|
|
"model_id": "Qwen/Qwen3.5-0.8B",
|
|
"architecture": "qwen",
|
|
"quant": "fp16",
|
|
"compute_dtype": "fp16",
|
|
"layer_start": 0,
|
|
"layer_end_exclusive": 8,
|
|
"config": {},
|
|
"tensor_count": 128,
|
|
"total_weight_bytes": 123456789
|
|
}
|
|
}
|
|
```
|
|
|
|
### 7.5 Tensor Transfer
|
|
|
|
Each tensor chunk message includes metadata:
|
|
|
|
```json
|
|
{
|
|
"type": "WEIGHT_CHUNK",
|
|
"payload": {
|
|
"tensor_name": "model.layers.0.self_attn.q_proj.weight",
|
|
"dtype": "float16",
|
|
"shape": [1024, 1024],
|
|
"chunk_index": 0,
|
|
"chunk_count": 4,
|
|
"offset": 0,
|
|
"data": "binary payload or external raw section"
|
|
}
|
|
}
|
|
```
|
|
|
|
For very large tensors, the preferred format is:
|
|
|
|
```text
|
|
[frame length][msgpack metadata][raw bytes referenced by metadata]
|
|
```
|
|
|
|
The implementation should keep this hidden behind `protocol/tensors.py`.
|
|
|
|
### 7.6 Inference Messages
|
|
|
|
`RUN_PREFILL`:
|
|
|
|
```json
|
|
{
|
|
"type": "RUN_PREFILL",
|
|
"request_id": "req-1",
|
|
"payload": {
|
|
"position_start": 0,
|
|
"input_length": 42,
|
|
"hidden_state": "tensor payload"
|
|
}
|
|
}
|
|
```
|
|
|
|
`RUN_DECODE`:
|
|
|
|
```json
|
|
{
|
|
"type": "RUN_DECODE",
|
|
"request_id": "req-1",
|
|
"payload": {
|
|
"position": 42,
|
|
"hidden_state": "tensor payload"
|
|
}
|
|
}
|
|
```
|
|
|
|
`HIDDEN_STATE`:
|
|
|
|
```json
|
|
{
|
|
"type": "HIDDEN_STATE",
|
|
"request_id": "req-1",
|
|
"payload": {
|
|
"hidden_state": "tensor payload"
|
|
}
|
|
}
|
|
```
|
|
|
|
## 8. Model Support
|
|
|
|
### 8.1 Initial Architecture
|
|
|
|
Initial implementation should support Qwen2/Qwen2.5-style decoder-only causal LMs. The first implementation does not support Qwen3.5 hybrid models with `linear_attention`/GatedDeltaNet layers.
|
|
|
|
Required components:
|
|
|
|
- Token embedding.
|
|
- Stacked transformer decoder blocks.
|
|
- RMSNorm.
|
|
- Rotary position embeddings.
|
|
- Grouped-query attention.
|
|
- Causal attention mask.
|
|
- Gated MLP/SwiGLU.
|
|
- Final RMSNorm.
|
|
- LM head.
|
|
- Tied or untied output embeddings.
|
|
|
|
### 8.2 HuggingFace Files
|
|
|
|
Coordinator should support models with:
|
|
|
|
```text
|
|
config.json
|
|
tokenizer.json/tokenizer.model/tokenizer_config.json
|
|
model.safetensors or model-00001-of-000xx.safetensors
|
|
model.safetensors.index.json, if sharded
|
|
```
|
|
|
|
Use libraries:
|
|
|
|
- `huggingface_hub`
|
|
- `safetensors`
|
|
- `transformers` for tokenizer/config only where possible
|
|
- `torch`
|
|
|
|
### 8.3 Weight Name Mapping
|
|
|
|
For Qwen-style models, expected tensor names include patterns like:
|
|
|
|
```text
|
|
model.embed_tokens.weight
|
|
model.layers.{i}.input_layernorm.weight
|
|
model.layers.{i}.self_attn.q_proj.weight
|
|
model.layers.{i}.self_attn.k_proj.weight
|
|
model.layers.{i}.self_attn.v_proj.weight
|
|
model.layers.{i}.self_attn.o_proj.weight
|
|
model.layers.{i}.post_attention_layernorm.weight
|
|
model.layers.{i}.mlp.gate_proj.weight
|
|
model.layers.{i}.mlp.up_proj.weight
|
|
model.layers.{i}.mlp.down_proj.weight
|
|
model.norm.weight
|
|
lm_head.weight
|
|
```
|
|
|
|
The model loader should validate required tensors before accepting a model.
|
|
|
|
## 9. Model Planning and Splitting
|
|
|
|
The coordinator computes layer assignments from model metadata and node capacity.
|
|
|
|
### 9.1 Inputs
|
|
|
|
- Number of transformer layers.
|
|
- Per-layer tensor byte sizes.
|
|
- Quantization mode.
|
|
- Max nodes.
|
|
- Connected node capabilities.
|
|
- Optional target memory hint.
|
|
|
|
### 9.2 Rules
|
|
|
|
- Split only on full transformer layer boundaries.
|
|
- Assign contiguous layer ranges.
|
|
- Preserve layer order.
|
|
- Coordinator keeps embeddings, final norm, and lm head.
|
|
- Required nodes must be connected and loaded before generation.
|
|
- If insufficient nodes are available, API returns `cluster_not_ready`.
|
|
|
|
### 9.3 First Planner Algorithm
|
|
|
|
Simple deterministic version:
|
|
|
|
1. Compute total transformer layer bytes after quantization.
|
|
2. Estimate bytes per layer.
|
|
3. Determine number of partitions as `min(max_nodes, num_layers)`.
|
|
4. If node memory information is available, reduce/increase partitions so each assignment fits.
|
|
5. Otherwise split evenly by layer byte size.
|
|
6. Assign partitions to the first compatible ready nodes.
|
|
|
|
### 9.4 Future Planner Improvements
|
|
|
|
- Benchmark node speed and assign more layers to faster GPUs.
|
|
- Prefer CUDA nodes for larger shards.
|
|
- Consider network latency and bandwidth.
|
|
- Replicate small layers for resilience.
|
|
- Rebalance between generations.
|
|
|
|
## 10. Quantization
|
|
|
|
Quantization must work on CUDA, MPS, and CPU. Therefore the baseline implementation should avoid CUDA-only dependencies such as bitsandbytes.
|
|
|
|
Supported modes:
|
|
|
|
```text
|
|
fp16
|
|
int8
|
|
int4
|
|
```
|
|
|
|
### 10.1 `fp16`
|
|
|
|
- Store weights as `torch.float16`.
|
|
- Compute in `float16` by default.
|
|
- Works on CUDA and MPS.
|
|
- CPU fallback may use `float32` internally if needed.
|
|
|
|
### 10.2 Portable `int8`
|
|
|
|
Use symmetric per-output-channel weight-only quantization.
|
|
|
|
For a linear weight `W` shaped `[out_features, in_features]`:
|
|
|
|
```text
|
|
scale[out_features] = max(abs(W[row])) / 127
|
|
qweight[row] = round(W[row] / scale[row]).clamp(-127, 127).int8
|
|
```
|
|
|
|
Forward path:
|
|
|
|
```text
|
|
W_dequant = qweight.float() * scale[:, None]
|
|
y = x @ W_dequant.T
|
|
```
|
|
|
|
This is portable but not maximally fast.
|
|
|
|
### 10.3 Portable `int4`
|
|
|
|
Use group-wise weight-only quantization.
|
|
|
|
Suggested default group size: `128`.
|
|
|
|
Store:
|
|
|
|
```text
|
|
packed_qweight: uint8
|
|
scale: float16/float32 per group
|
|
zero_point: optional
|
|
metadata: original shape, group size, packing order
|
|
```
|
|
|
|
Forward path:
|
|
|
|
1. Unpack int4 values.
|
|
2. Dequantize to compute dtype.
|
|
3. Perform normal PyTorch matmul.
|
|
|
|
This is designed for correctness and portability, not peak speed.
|
|
|
|
### 10.4 Quantization Timing
|
|
|
|
Preferred prototype behavior:
|
|
|
|
- Coordinator loads original safetensors.
|
|
- Coordinator quantizes tensors before sending to nodes if `int8` or `int4` is selected.
|
|
- Nodes receive already-quantized tensors plus quantization metadata.
|
|
- Coordinator also quantizes/loads its own embedding/lm_head as needed.
|
|
|
|
## 11. Generation Algorithm
|
|
|
|
### 11.1 Prefill
|
|
|
|
For prompt token IDs of length `N`:
|
|
|
|
1. Coordinator computes embeddings: `[1, N, hidden_size]`.
|
|
2. Coordinator sends hidden state to first node with position start `0`.
|
|
3. Each node runs its assigned layers across the full sequence.
|
|
4. Each node initializes KV cache for its layers.
|
|
5. Final node returns hidden state to coordinator.
|
|
6. Coordinator applies final norm and lm head to the last token.
|
|
7. Coordinator samples next token.
|
|
|
|
### 11.2 Decode
|
|
|
|
For each generated token:
|
|
|
|
1. Coordinator embeds last token: `[1, 1, hidden_size]`.
|
|
2. Coordinator sends hidden state to first node with current position.
|
|
3. Each node runs one-token decode using local KV cache.
|
|
4. Each node appends to its KV cache.
|
|
5. Final node returns hidden state to coordinator.
|
|
6. Coordinator computes logits and samples next token.
|
|
7. Stop if EOS, stop sequence, or `max_tokens` reached.
|
|
|
|
### 11.3 Sampling
|
|
|
|
Initial sampler supports:
|
|
|
|
- Greedy when `temperature == 0`.
|
|
- Temperature scaling.
|
|
- Top-p nucleus sampling.
|
|
- EOS handling.
|
|
- Stop strings after decoding.
|
|
|
|
Future additions:
|
|
|
|
- Top-k.
|
|
- Repetition penalty.
|
|
- Frequency/presence penalties.
|
|
- Logprobs.
|
|
|
|
## 12. Device Support
|
|
|
|
### 12.1 Device Auto Detection
|
|
|
|
Node device priority when `--device auto`:
|
|
|
|
1. CUDA if available.
|
|
2. MPS if available.
|
|
3. CPU fallback.
|
|
|
|
### 12.2 CUDA
|
|
|
|
Use:
|
|
|
|
```python
|
|
torch.cuda.is_available()
|
|
torch.cuda.get_device_properties(index)
|
|
torch.cuda.mem_get_info(index)
|
|
```
|
|
|
|
### 12.3 Apple MPS
|
|
|
|
Use:
|
|
|
|
```python
|
|
torch.backends.mps.is_available()
|
|
torch.device("mps")
|
|
```
|
|
|
|
MPS memory reporting is limited, so allow `--max-memory-gb` override.
|
|
|
|
### 12.4 CPU
|
|
|
|
CPU is allowed for testing and fallback, but may be slow.
|
|
|
|
## 13. Package Structure
|
|
|
|
Recommended source tree:
|
|
|
|
```text
|
|
truecluster/
|
|
__init__.py
|
|
cli.py
|
|
|
|
cluster/
|
|
__init__.py
|
|
server.py # node TCP server
|
|
api.py # FastAPI OpenAI-compatible API
|
|
planner.py # layer splitting
|
|
scheduler.py # generation orchestration
|
|
model_store.py # HF/safetensors loading
|
|
sampler.py
|
|
state.py
|
|
|
|
node/
|
|
__init__.py
|
|
client.py # connects to cluster
|
|
runtime.py # owns assigned layers + cache
|
|
device.py
|
|
|
|
model/
|
|
__init__.py
|
|
qwen.py # minimal Qwen implementation
|
|
layers.py
|
|
rotary.py
|
|
kv_cache.py
|
|
quant.py
|
|
tensor_names.py
|
|
|
|
protocol/
|
|
__init__.py
|
|
framing.py
|
|
messages.py
|
|
tensors.py
|
|
|
|
tests/
|
|
test_single_node_matches_transformers.py
|
|
test_protocol.py
|
|
test_quant.py
|
|
test_planner.py
|
|
```
|
|
|
|
Project metadata:
|
|
|
|
```text
|
|
pyproject.toml
|
|
README.md
|
|
SPEC.md
|
|
```
|
|
|
|
## 14. Suggested Dependencies
|
|
|
|
Runtime:
|
|
|
|
```text
|
|
torch
|
|
transformers
|
|
huggingface_hub
|
|
safetensors
|
|
fastapi
|
|
uvicorn[standard]
|
|
typer
|
|
msgpack
|
|
pydantic
|
|
numpy
|
|
tqdm
|
|
```
|
|
|
|
Development/test:
|
|
|
|
```text
|
|
pytest
|
|
pytest-asyncio
|
|
httpx
|
|
ruff
|
|
mypy optional
|
|
```
|
|
|
|
Python version:
|
|
|
|
```text
|
|
>=3.10,<3.13
|
|
```
|
|
|
|
## 15. Validation and Testing
|
|
|
|
### 15.1 Correctness Test Against Transformers
|
|
|
|
Most important validation:
|
|
|
|
1. Load the target model with HuggingFace Transformers locally.
|
|
2. Load the same model through TrueCluster with one local node.
|
|
3. Run the same prompt.
|
|
4. Compare final logits before sampling.
|
|
5. Assert max difference is within tolerance for selected dtype.
|
|
|
|
### 15.2 Multi-node Local Test
|
|
|
|
Run on one machine:
|
|
|
|
```bash
|
|
truecluster cluster --model ... --max-nodes 2
|
|
truecluster node --cluster-host 127.0.0.1 --cluster-port 7001 --device cpu
|
|
truecluster node --cluster-host 127.0.0.1 --cluster-port 7001 --device cpu
|
|
```
|
|
|
|
Verify:
|
|
|
|
- Both nodes receive different layer ranges.
|
|
- Prefill works.
|
|
- Decode works.
|
|
- Output matches single-node output within tolerance.
|
|
|
|
### 15.3 Mixed Hardware Test
|
|
|
|
Example:
|
|
|
|
```text
|
|
coordinator: Mac or Linux host
|
|
node 1: Nvidia CUDA machine
|
|
node 2: Apple Silicon Mac MPS machine
|
|
```
|
|
|
|
Verify:
|
|
|
|
- Both nodes connect.
|
|
- Assignments are sent.
|
|
- Generation completes.
|
|
|
|
### 15.4 Quantization Tests
|
|
|
|
For `int8` and `int4`:
|
|
|
|
- Quantize/dequantize synthetic tensors.
|
|
- Check shape preservation.
|
|
- Check error bounds.
|
|
- Run short generation and verify no crashes.
|
|
|
|
## 16. Implementation Phases
|
|
|
|
### Phase 1: Local Single-Process Model Proof
|
|
|
|
Deliverables:
|
|
|
|
- Minimal Qwen model implementation.
|
|
- Safetensors loading.
|
|
- Local full-model forward.
|
|
- Logit comparison against Transformers.
|
|
|
|
### Phase 2: One Node Distributed Inference
|
|
|
|
Deliverables:
|
|
|
|
- Socket protocol.
|
|
- Coordinator sends all transformer layers to one local node.
|
|
- Node loads layers and runs them.
|
|
- Coordinator keeps embedding/final norm/lm head.
|
|
- `/v1/completions` works.
|
|
|
|
### Phase 3: Multi-node Layer Split
|
|
|
|
Deliverables:
|
|
|
|
- Planner assigns contiguous layer ranges.
|
|
- Multiple nodes are supported.
|
|
- Distributed KV cache works.
|
|
- Not-ready errors work.
|
|
|
|
### Phase 4: Mixed CUDA/MPS Support
|
|
|
|
Deliverables:
|
|
|
|
- Device detection.
|
|
- CUDA execution.
|
|
- MPS execution.
|
|
- CPU fallback.
|
|
- Mixed Nvidia/Mac cluster generation test.
|
|
|
|
### Phase 5: Portable Quantization
|
|
|
|
Deliverables:
|
|
|
|
- `fp16` baseline.
|
|
- Portable `int8` linear.
|
|
- Portable `int4` linear.
|
|
- Quantized weight transfer.
|
|
- CLI `--quant` option.
|
|
|
|
### Phase 6: API Polish
|
|
|
|
Deliverables:
|
|
|
|
- `/v1/models`.
|
|
- `/v1/chat/completions`.
|
|
- Stop sequence support.
|
|
- Usage accounting.
|
|
- Better OpenAI-compatible errors.
|
|
|
|
## 17. Initial Acceptance Criteria
|
|
|
|
A prototype is considered working when:
|
|
|
|
1. A cluster can be started with a HuggingFace safetensors Qwen-style model.
|
|
2. A node can connect to the cluster with no local model files.
|
|
3. The cluster sends layer weights to the node over the socket.
|
|
4. The node loads assigned layers on CUDA, MPS, or CPU.
|
|
5. `/v1/models` returns the loaded model.
|
|
6. `/v1/completions` generates text through the distributed pipeline.
|
|
7. `/v1/chat/completions` works for simple chat prompts.
|
|
8. If insufficient nodes are ready, API returns a clear `cluster_not_ready` error.
|
|
9. One local-node output matches Transformers logits within reasonable dtype tolerance.
|
|
10. Multi-node local CPU test works.
|
|
|
|
## 18. Key Design Decisions
|
|
|
|
- Use pipeline parallelism, not tensor parallelism.
|
|
- Use contiguous layer ranges only.
|
|
- Keep tokenizer, embeddings, final norm, lm head, and sampler on the coordinator.
|
|
- Send weights once at node assignment time.
|
|
- Send hidden states during generation.
|
|
- Store KV cache on worker nodes.
|
|
- Implement portable quantization instead of relying on CUDA-only libraries.
|
|
- Start with Qwen2/Qwen2.5-style causal LMs only.
|
|
- Optimize for correctness and clean architecture before speed.
|