Removing dumb idea
This commit is contained in:
@@ -2,7 +2,7 @@
|
||||
|
||||
## 1. Goal
|
||||
|
||||
TrueCluster is a Python 3.10 prototype for distributed LLM inference across a small cluster of heterogeneous machines. It allows a coordinator/cluster process to load a HuggingFace safetensors model, split transformer layers across connected worker nodes, send each node only the weights it needs over a socket connection, and expose a simple OpenAI-compatible HTTP API for generation.
|
||||
TrueCluster is a Python 3.10 prototype for distributed LLM inference across a small cluster of heterogeneous machines. It allows a coordinator/cluster process to load a HuggingFace safetensors model config/tokenizer/coordinator-owned tensors, split transformer layers across connected worker nodes, tell each node which layer range to load from HuggingFace, and expose a simple OpenAI-compatible HTTP API for generation.
|
||||
|
||||
The initial target is small Qwen2/Qwen2.5-style decoder-only models around 0.5B-1.5B parameters, with first-class support for mixed Nvidia CUDA and Apple Silicon Mac MPS nodes in the same cluster. Qwen3.5 hybrid linear-attention models are a later target and are not part of the first fp16 prototype.
|
||||
|
||||
@@ -12,8 +12,8 @@ Primary prototype goals:
|
||||
- Easy CLI for running a cluster or node.
|
||||
- Coordinator loads one model at a time.
|
||||
- Nodes connect to the coordinator by host/port.
|
||||
- Coordinator sends assigned model shards to nodes over sockets.
|
||||
- Nodes do not need local model files.
|
||||
- Coordinator sends layer assignments to nodes over sockets.
|
||||
- Nodes download/resolve the full HuggingFace model themselves, then load only their assigned layers.
|
||||
- Support HuggingFace safetensors models.
|
||||
- Initial architecture target: Qwen2/Qwen2.5-style causal language models.
|
||||
- Support `fp16`, portable `int8`, and portable `int4` weight-only quantization modes.
|
||||
@@ -31,7 +31,7 @@ The first prototype will intentionally avoid several advanced features:
|
||||
- No expert parallelism.
|
||||
- No dynamic model hot-swapping.
|
||||
- No high-performance CUDA-only quantization kernels as a requirement.
|
||||
- No dependency on every node having HuggingFace access.
|
||||
- No coordinator-to-node weight transfer. Nodes are expected to have HuggingFace/model access.
|
||||
- No web UI.
|
||||
- No streaming responses in the first milestone.
|
||||
- No fault-tolerant recovery during an active generation.
|
||||
@@ -49,7 +49,7 @@ The coordinator owns:
|
||||
- Tokenizer.
|
||||
- Sampling logic.
|
||||
- HuggingFace model metadata/config loading.
|
||||
- Model weight loading from safetensors.
|
||||
- Coordinator-owned tensor loading from safetensors.
|
||||
- Model split planning.
|
||||
- Embedding layer.
|
||||
- Final normalization.
|
||||
@@ -81,7 +81,7 @@ HTTP request
|
||||
-> repeat until completion
|
||||
```
|
||||
|
||||
Weights are transferred once during assignment. During generation, only hidden states and small metadata are passed between coordinator and nodes.
|
||||
Weights are not transferred over the cluster socket. During assignment, the coordinator sends model id, config, and layer range. Each node downloads/resolves the model from HuggingFace or a matching local path and loads only its assigned layer tensors. During generation, only hidden states and small metadata are passed between coordinator and nodes.
|
||||
|
||||
## 4. Execution Model
|
||||
|
||||
@@ -173,12 +173,12 @@ Behavior:
|
||||
|
||||
1. Resolve/download model.
|
||||
2. Load config and tokenizer.
|
||||
3. Load safetensors into coordinator RAM or memory-mapped index.
|
||||
4. Build model tensor index.
|
||||
5. Start node socket server.
|
||||
6. Start HTTP API.
|
||||
7. Wait for enough nodes.
|
||||
8. Assign layer ranges and transmit weights.
|
||||
3. Build safetensors metadata index and load only coordinator-owned tensors.
|
||||
4. Start node socket server.
|
||||
5. Start HTTP API.
|
||||
6. Wait for enough nodes.
|
||||
7. Assign layer ranges.
|
||||
8. Nodes download/resolve the model and load assigned layers.
|
||||
9. Mark cluster ready.
|
||||
|
||||
### 5.2 Node Command
|
||||
@@ -199,7 +199,8 @@ Options:
|
||||
--cluster-port INT Coordinator node socket port.
|
||||
--device TEXT auto, cuda, cuda:0, mps, or cpu. Default: auto
|
||||
--node-id TEXT Optional stable node id.
|
||||
--work-dir PATH Temporary local directory for received weights/cache.
|
||||
--hf-cache-dir PATH Optional HuggingFace cache directory for node downloads.
|
||||
--work-dir PATH Reserved for future temporary/cache files.
|
||||
--max-memory-gb FLOAT Optional memory capability override.
|
||||
--log-level TEXT Default: info
|
||||
```
|
||||
@@ -210,8 +211,8 @@ Behavior:
|
||||
2. Connect to coordinator socket.
|
||||
3. Send `HELLO` capability message.
|
||||
4. Wait for assignment.
|
||||
5. Receive config and layer weights.
|
||||
6. Build local model shard.
|
||||
5. Receive config and assigned layer range.
|
||||
6. Download/resolve the model locally and build local model shard.
|
||||
7. Mark itself ready.
|
||||
8. Serve prefill/decode requests over the persistent connection.
|
||||
|
||||
@@ -350,8 +351,6 @@ Coordinator/node lifecycle:
|
||||
HELLO
|
||||
HELLO_ACK
|
||||
ASSIGNMENT
|
||||
WEIGHT_CHUNK
|
||||
WEIGHTS_COMPLETE
|
||||
LOAD_COMPLETE
|
||||
LOAD_FAILED
|
||||
PING
|
||||
@@ -432,32 +431,9 @@ Sent by coordinator after planning.
|
||||
}
|
||||
```
|
||||
|
||||
### 7.5 Tensor Transfer
|
||||
### 7.5 Model Loading on Nodes
|
||||
|
||||
Each tensor chunk message includes metadata:
|
||||
|
||||
```json
|
||||
{
|
||||
"type": "WEIGHT_CHUNK",
|
||||
"payload": {
|
||||
"tensor_name": "model.layers.0.self_attn.q_proj.weight",
|
||||
"dtype": "float16",
|
||||
"shape": [1024, 1024],
|
||||
"chunk_index": 0,
|
||||
"chunk_count": 4,
|
||||
"offset": 0,
|
||||
"data": "binary payload or external raw section"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
For very large tensors, the preferred format is:
|
||||
|
||||
```text
|
||||
[frame length][msgpack metadata][raw bytes referenced by metadata]
|
||||
```
|
||||
|
||||
The implementation should keep this hidden behind `protocol/tensors.py`.
|
||||
The coordinator does not transfer model weights. After `ASSIGNMENT`, each node resolves/downloads `model_id` itself using HuggingFace cache semantics and loads only tensors whose names match its assigned layer range.
|
||||
|
||||
### 7.6 Inference Messages
|
||||
|
||||
@@ -666,8 +642,7 @@ This is designed for correctness and portability, not peak speed.
|
||||
Preferred prototype behavior:
|
||||
|
||||
- Coordinator loads original safetensors.
|
||||
- Coordinator quantizes tensors before sending to nodes if `int8` or `int4` is selected.
|
||||
- Nodes receive already-quantized tensors plus quantization metadata.
|
||||
- Nodes quantize their locally loaded assigned tensors if `int8` or `int4` is selected.
|
||||
- Coordinator also quantizes/loads its own embedding/lm_head as needed.
|
||||
|
||||
## 11. Generation Algorithm
|
||||
@@ -907,7 +882,7 @@ Deliverables:
|
||||
Deliverables:
|
||||
|
||||
- Socket protocol.
|
||||
- Coordinator sends all transformer layers to one local node.
|
||||
- Coordinator assigns all transformer layers to one local node.
|
||||
- Node loads layers and runs them.
|
||||
- Coordinator keeps embedding/final norm/lm head.
|
||||
- `/v1/completions` works.
|
||||
@@ -938,7 +913,7 @@ Deliverables:
|
||||
- `fp16` baseline.
|
||||
- Portable `int8` linear.
|
||||
- Portable `int4` linear.
|
||||
- Quantized weight transfer.
|
||||
- Node-local quantized layer loading.
|
||||
- CLI `--quant` option.
|
||||
|
||||
### Phase 6: API Polish
|
||||
@@ -957,8 +932,8 @@ A prototype is considered working when:
|
||||
|
||||
1. A cluster can be started with a HuggingFace safetensors Qwen-style model.
|
||||
2. A node can connect to the cluster with no local model files.
|
||||
3. The cluster sends layer weights to the node over the socket.
|
||||
4. The node loads assigned layers on CUDA, MPS, or CPU.
|
||||
3. The cluster sends a layer assignment to the node over the socket.
|
||||
4. The node downloads/resolves the model and loads assigned layers on CUDA, MPS, or CPU.
|
||||
5. `/v1/models` returns the loaded model.
|
||||
6. `/v1/completions` generates text through the distributed pipeline.
|
||||
7. `/v1/chat/completions` works for simple chat prompts.
|
||||
@@ -971,7 +946,7 @@ A prototype is considered working when:
|
||||
- Use pipeline parallelism, not tensor parallelism.
|
||||
- Use contiguous layer ranges only.
|
||||
- Keep tokenizer, embeddings, final norm, lm head, and sampler on the coordinator.
|
||||
- Send weights once at node assignment time.
|
||||
- Do not send weights over the node socket; nodes load weights locally from HuggingFace/cache.
|
||||
- Send hidden states during generation.
|
||||
- Store KV cache on worker nodes.
|
||||
- Implement portable quantization instead of relying on CUDA-only libraries.
|
||||
|
||||
Reference in New Issue
Block a user