Removing dumb idea

This commit is contained in:
2026-06-04 19:03:03 -05:00
parent 6e54d9eb5e
commit 04eafb5446
7 changed files with 184 additions and 186 deletions
+24 -49
View File
@@ -2,7 +2,7 @@
## 1. Goal
TrueCluster is a Python 3.10 prototype for distributed LLM inference across a small cluster of heterogeneous machines. It allows a coordinator/cluster process to load a HuggingFace safetensors model, split transformer layers across connected worker nodes, send each node only the weights it needs over a socket connection, and expose a simple OpenAI-compatible HTTP API for generation.
TrueCluster is a Python 3.10 prototype for distributed LLM inference across a small cluster of heterogeneous machines. It allows a coordinator/cluster process to load a HuggingFace safetensors model config/tokenizer/coordinator-owned tensors, split transformer layers across connected worker nodes, tell each node which layer range to load from HuggingFace, and expose a simple OpenAI-compatible HTTP API for generation.
The initial target is small Qwen2/Qwen2.5-style decoder-only models around 0.5B-1.5B parameters, with first-class support for mixed Nvidia CUDA and Apple Silicon Mac MPS nodes in the same cluster. Qwen3.5 hybrid linear-attention models are a later target and are not part of the first fp16 prototype.
@@ -12,8 +12,8 @@ Primary prototype goals:
- Easy CLI for running a cluster or node.
- Coordinator loads one model at a time.
- Nodes connect to the coordinator by host/port.
- Coordinator sends assigned model shards to nodes over sockets.
- Nodes do not need local model files.
- Coordinator sends layer assignments to nodes over sockets.
- Nodes download/resolve the full HuggingFace model themselves, then load only their assigned layers.
- Support HuggingFace safetensors models.
- Initial architecture target: Qwen2/Qwen2.5-style causal language models.
- Support `fp16`, portable `int8`, and portable `int4` weight-only quantization modes.
@@ -31,7 +31,7 @@ The first prototype will intentionally avoid several advanced features:
- No expert parallelism.
- No dynamic model hot-swapping.
- No high-performance CUDA-only quantization kernels as a requirement.
- No dependency on every node having HuggingFace access.
- No coordinator-to-node weight transfer. Nodes are expected to have HuggingFace/model access.
- No web UI.
- No streaming responses in the first milestone.
- No fault-tolerant recovery during an active generation.
@@ -49,7 +49,7 @@ The coordinator owns:
- Tokenizer.
- Sampling logic.
- HuggingFace model metadata/config loading.
- Model weight loading from safetensors.
- Coordinator-owned tensor loading from safetensors.
- Model split planning.
- Embedding layer.
- Final normalization.
@@ -81,7 +81,7 @@ HTTP request
-> repeat until completion
```
Weights are transferred once during assignment. During generation, only hidden states and small metadata are passed between coordinator and nodes.
Weights are not transferred over the cluster socket. During assignment, the coordinator sends model id, config, and layer range. Each node downloads/resolves the model from HuggingFace or a matching local path and loads only its assigned layer tensors. During generation, only hidden states and small metadata are passed between coordinator and nodes.
## 4. Execution Model
@@ -173,12 +173,12 @@ Behavior:
1. Resolve/download model.
2. Load config and tokenizer.
3. Load safetensors into coordinator RAM or memory-mapped index.
4. Build model tensor index.
5. Start node socket server.
6. Start HTTP API.
7. Wait for enough nodes.
8. Assign layer ranges and transmit weights.
3. Build safetensors metadata index and load only coordinator-owned tensors.
4. Start node socket server.
5. Start HTTP API.
6. Wait for enough nodes.
7. Assign layer ranges.
8. Nodes download/resolve the model and load assigned layers.
9. Mark cluster ready.
### 5.2 Node Command
@@ -199,7 +199,8 @@ Options:
--cluster-port INT Coordinator node socket port.
--device TEXT auto, cuda, cuda:0, mps, or cpu. Default: auto
--node-id TEXT Optional stable node id.
--work-dir PATH Temporary local directory for received weights/cache.
--hf-cache-dir PATH Optional HuggingFace cache directory for node downloads.
--work-dir PATH Reserved for future temporary/cache files.
--max-memory-gb FLOAT Optional memory capability override.
--log-level TEXT Default: info
```
@@ -210,8 +211,8 @@ Behavior:
2. Connect to coordinator socket.
3. Send `HELLO` capability message.
4. Wait for assignment.
5. Receive config and layer weights.
6. Build local model shard.
5. Receive config and assigned layer range.
6. Download/resolve the model locally and build local model shard.
7. Mark itself ready.
8. Serve prefill/decode requests over the persistent connection.
@@ -350,8 +351,6 @@ Coordinator/node lifecycle:
HELLO
HELLO_ACK
ASSIGNMENT
WEIGHT_CHUNK
WEIGHTS_COMPLETE
LOAD_COMPLETE
LOAD_FAILED
PING
@@ -432,32 +431,9 @@ Sent by coordinator after planning.
}
```
### 7.5 Tensor Transfer
### 7.5 Model Loading on Nodes
Each tensor chunk message includes metadata:
```json
{
"type": "WEIGHT_CHUNK",
"payload": {
"tensor_name": "model.layers.0.self_attn.q_proj.weight",
"dtype": "float16",
"shape": [1024, 1024],
"chunk_index": 0,
"chunk_count": 4,
"offset": 0,
"data": "binary payload or external raw section"
}
}
```
For very large tensors, the preferred format is:
```text
[frame length][msgpack metadata][raw bytes referenced by metadata]
```
The implementation should keep this hidden behind `protocol/tensors.py`.
The coordinator does not transfer model weights. After `ASSIGNMENT`, each node resolves/downloads `model_id` itself using HuggingFace cache semantics and loads only tensors whose names match its assigned layer range.
### 7.6 Inference Messages
@@ -666,8 +642,7 @@ This is designed for correctness and portability, not peak speed.
Preferred prototype behavior:
- Coordinator loads original safetensors.
- Coordinator quantizes tensors before sending to nodes if `int8` or `int4` is selected.
- Nodes receive already-quantized tensors plus quantization metadata.
- Nodes quantize their locally loaded assigned tensors if `int8` or `int4` is selected.
- Coordinator also quantizes/loads its own embedding/lm_head as needed.
## 11. Generation Algorithm
@@ -907,7 +882,7 @@ Deliverables:
Deliverables:
- Socket protocol.
- Coordinator sends all transformer layers to one local node.
- Coordinator assigns all transformer layers to one local node.
- Node loads layers and runs them.
- Coordinator keeps embedding/final norm/lm head.
- `/v1/completions` works.
@@ -938,7 +913,7 @@ Deliverables:
- `fp16` baseline.
- Portable `int8` linear.
- Portable `int4` linear.
- Quantized weight transfer.
- Node-local quantized layer loading.
- CLI `--quant` option.
### Phase 6: API Polish
@@ -957,8 +932,8 @@ A prototype is considered working when:
1. A cluster can be started with a HuggingFace safetensors Qwen-style model.
2. A node can connect to the cluster with no local model files.
3. The cluster sends layer weights to the node over the socket.
4. The node loads assigned layers on CUDA, MPS, or CPU.
3. The cluster sends a layer assignment to the node over the socket.
4. The node downloads/resolves the model and loads assigned layers on CUDA, MPS, or CPU.
5. `/v1/models` returns the loaded model.
6. `/v1/completions` generates text through the distributed pipeline.
7. `/v1/chat/completions` works for simple chat prompts.
@@ -971,7 +946,7 @@ A prototype is considered working when:
- Use pipeline parallelism, not tensor parallelism.
- Use contiguous layer ranges only.
- Keep tokenizer, embeddings, final norm, lm head, and sampler on the coordinator.
- Send weights once at node assignment time.
- Do not send weights over the node socket; nodes load weights locally from HuggingFace/cache.
- Send hidden states during generation.
- Store KV cache on worker nodes.
- Implement portable quantization instead of relying on CUDA-only libraries.