Inital commit
This commit is contained in:
@@ -0,0 +1,979 @@
|
||||
# TrueCluster Prototype Specification
|
||||
|
||||
## 1. Goal
|
||||
|
||||
TrueCluster is a Python 3.10 prototype for distributed LLM inference across a small cluster of heterogeneous machines. It allows a coordinator/cluster process to load a HuggingFace safetensors model, split transformer layers across connected worker nodes, send each node only the weights it needs over a socket connection, and expose a simple OpenAI-compatible HTTP API for generation.
|
||||
|
||||
The initial target is small Qwen2/Qwen2.5-style decoder-only models around 0.5B-1.5B parameters, with first-class support for mixed Nvidia CUDA and Apple Silicon Mac MPS nodes in the same cluster. Qwen3.5 hybrid linear-attention models are a later target and are not part of the first fp16 prototype.
|
||||
|
||||
Primary prototype goals:
|
||||
|
||||
- Python 3.10.
|
||||
- Easy CLI for running a cluster or node.
|
||||
- Coordinator loads one model at a time.
|
||||
- Nodes connect to the coordinator by host/port.
|
||||
- Coordinator sends assigned model shards to nodes over sockets.
|
||||
- Nodes do not need local model files.
|
||||
- Support HuggingFace safetensors models.
|
||||
- Initial architecture target: Qwen2/Qwen2.5-style causal language models.
|
||||
- Support `fp16`, portable `int8`, and portable `int4` weight-only quantization modes.
|
||||
- OpenAI-compatible unauthenticated HTTP API.
|
||||
- Efficient generation by transferring activations during inference, not weights.
|
||||
|
||||
## 2. Non-goals for the Initial Prototype
|
||||
|
||||
The first prototype will intentionally avoid several advanced features:
|
||||
|
||||
- No multi-model serving.
|
||||
- No authentication.
|
||||
- No request batching.
|
||||
- No tensor parallelism across machines.
|
||||
- No expert parallelism.
|
||||
- No dynamic model hot-swapping.
|
||||
- No high-performance CUDA-only quantization kernels as a requirement.
|
||||
- No dependency on every node having HuggingFace access.
|
||||
- No web UI.
|
||||
- No streaming responses in the first milestone.
|
||||
- No fault-tolerant recovery during an active generation.
|
||||
|
||||
These can be added after a correct baseline works.
|
||||
|
||||
## 3. High-Level Architecture
|
||||
|
||||
TrueCluster uses pipeline-parallel inference.
|
||||
|
||||
The coordinator owns:
|
||||
|
||||
- CLI entry point for `cluster`.
|
||||
- HTTP API.
|
||||
- Tokenizer.
|
||||
- Sampling logic.
|
||||
- HuggingFace model metadata/config loading.
|
||||
- Model weight loading from safetensors.
|
||||
- Model split planning.
|
||||
- Embedding layer.
|
||||
- Final normalization.
|
||||
- LM head.
|
||||
- Node registry and orchestration.
|
||||
|
||||
Each node owns:
|
||||
|
||||
- CLI entry point for `node`.
|
||||
- Persistent socket connection to coordinator.
|
||||
- Device detection and selection.
|
||||
- One contiguous range of transformer layers.
|
||||
- KV cache for its assigned layers.
|
||||
- Local forward execution on CUDA, MPS, or CPU.
|
||||
|
||||
Generation path:
|
||||
|
||||
```text
|
||||
HTTP request
|
||||
-> coordinator tokenizes prompt
|
||||
-> coordinator runs embedding
|
||||
-> hidden states sent to node 1
|
||||
-> node 1 runs assigned layers
|
||||
-> hidden states sent to node 2
|
||||
-> ...
|
||||
-> final hidden states returned to coordinator
|
||||
-> coordinator runs final norm + lm_head
|
||||
-> coordinator samples next token
|
||||
-> repeat until completion
|
||||
```
|
||||
|
||||
Weights are transferred once during assignment. During generation, only hidden states and small metadata are passed between coordinator and nodes.
|
||||
|
||||
## 4. Execution Model
|
||||
|
||||
### 4.1 Pipeline Parallelism
|
||||
|
||||
The model is split by complete transformer blocks. Each node receives a contiguous set of layers:
|
||||
|
||||
```text
|
||||
coordinator:
|
||||
embed_tokens
|
||||
final_norm
|
||||
lm_head
|
||||
|
||||
node 1:
|
||||
layers 0-7
|
||||
|
||||
node 2:
|
||||
layers 8-15
|
||||
|
||||
node 3:
|
||||
layers 16-23
|
||||
```
|
||||
|
||||
This is simpler and more reliable than tensor parallelism for heterogeneous machines and normal Ethernet/Wi-Fi networks.
|
||||
|
||||
### 4.2 KV Cache Ownership
|
||||
|
||||
KV cache is stored on the node that owns the relevant layers.
|
||||
|
||||
For example:
|
||||
|
||||
```text
|
||||
node 1 cache: layers 0-7
|
||||
node 2 cache: layers 8-15
|
||||
node 3 cache: layers 16-23
|
||||
```
|
||||
|
||||
During prefill, each node creates cache entries for its layers. During decode, each node appends one token of keys/values to its cache.
|
||||
|
||||
### 4.3 Request Concurrency
|
||||
|
||||
Initial prototype supports one active generation at a time per cluster.
|
||||
|
||||
Reason: distributed KV cache management is much easier with a single active request. Later versions can introduce request IDs, cache slots, and batching.
|
||||
|
||||
## 5. CLI Design
|
||||
|
||||
Use `typer` for the CLI.
|
||||
|
||||
Package command:
|
||||
|
||||
```bash
|
||||
truecluster
|
||||
```
|
||||
|
||||
### 5.1 Cluster Command
|
||||
|
||||
Example:
|
||||
|
||||
```bash
|
||||
truecluster cluster \
|
||||
--model Qwen/Qwen3.5-0.8B \
|
||||
--node-host 0.0.0.0 \
|
||||
--node-port 7001 \
|
||||
--api-host 0.0.0.0 \
|
||||
--api-port 8000 \
|
||||
--max-nodes 4 \
|
||||
--quant fp16
|
||||
```
|
||||
|
||||
Options:
|
||||
|
||||
```text
|
||||
--model TEXT HuggingFace model id or local path.
|
||||
--node-host TEXT Host/IP for worker-node socket server. Default: 0.0.0.0
|
||||
--node-port INT Port for worker-node socket server. Default: 7001
|
||||
--api-host TEXT Host/IP for HTTP API. Default: 0.0.0.0
|
||||
--api-port INT HTTP API port. Default: 8000
|
||||
--max-nodes INT Maximum number of worker nodes to use.
|
||||
--quant [fp16|int8|int4] Weight mode. Default: fp16
|
||||
--dtype [fp16|bf16|fp32] Compute dtype preference. Default: fp16
|
||||
--target-node-memory-gb FLOAT Optional planning hint if node memory is unknown.
|
||||
--trust-remote-code BOOL HuggingFace trust_remote_code. Default: false
|
||||
--hf-cache-dir PATH Optional HuggingFace cache directory.
|
||||
--log-level TEXT Default: info
|
||||
```
|
||||
|
||||
Behavior:
|
||||
|
||||
1. Resolve/download model.
|
||||
2. Load config and tokenizer.
|
||||
3. Load safetensors into coordinator RAM or memory-mapped index.
|
||||
4. Build model tensor index.
|
||||
5. Start node socket server.
|
||||
6. Start HTTP API.
|
||||
7. Wait for enough nodes.
|
||||
8. Assign layer ranges and transmit weights.
|
||||
9. Mark cluster ready.
|
||||
|
||||
### 5.2 Node Command
|
||||
|
||||
Example:
|
||||
|
||||
```bash
|
||||
truecluster node \
|
||||
--cluster-host 192.168.1.50 \
|
||||
--cluster-port 7001 \
|
||||
--device auto
|
||||
```
|
||||
|
||||
Options:
|
||||
|
||||
```text
|
||||
--cluster-host TEXT Coordinator node socket host.
|
||||
--cluster-port INT Coordinator node socket port.
|
||||
--device TEXT auto, cuda, cuda:0, mps, or cpu. Default: auto
|
||||
--node-id TEXT Optional stable node id.
|
||||
--work-dir PATH Temporary local directory for received weights/cache.
|
||||
--max-memory-gb FLOAT Optional memory capability override.
|
||||
--log-level TEXT Default: info
|
||||
```
|
||||
|
||||
Behavior:
|
||||
|
||||
1. Detect hardware and PyTorch backends.
|
||||
2. Connect to coordinator socket.
|
||||
3. Send `HELLO` capability message.
|
||||
4. Wait for assignment.
|
||||
5. Receive config and layer weights.
|
||||
6. Build local model shard.
|
||||
7. Mark itself ready.
|
||||
8. Serve prefill/decode requests over the persistent connection.
|
||||
|
||||
## 6. HTTP API
|
||||
|
||||
Use FastAPI and Uvicorn.
|
||||
|
||||
The API is unauthenticated for the prototype.
|
||||
|
||||
### 6.1 `GET /v1/models`
|
||||
|
||||
Returns the single loaded model.
|
||||
|
||||
Example response:
|
||||
|
||||
```json
|
||||
{
|
||||
"object": "list",
|
||||
"data": [
|
||||
{
|
||||
"id": "Qwen/Qwen3.5-0.8B",
|
||||
"object": "model",
|
||||
"created": 0,
|
||||
"owned_by": "truecluster"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### 6.2 `POST /v1/completions`
|
||||
|
||||
Supported request fields initially:
|
||||
|
||||
```json
|
||||
{
|
||||
"model": "Qwen/Qwen3.5-0.8B",
|
||||
"prompt": "Hello",
|
||||
"max_tokens": 64,
|
||||
"temperature": 0.7,
|
||||
"top_p": 0.95,
|
||||
"stop": null
|
||||
}
|
||||
```
|
||||
|
||||
Response should be OpenAI-compatible enough for common clients:
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "cmpl-...",
|
||||
"object": "text_completion",
|
||||
"created": 0,
|
||||
"model": "Qwen/Qwen3.5-0.8B",
|
||||
"choices": [
|
||||
{
|
||||
"text": " world",
|
||||
"index": 0,
|
||||
"logprobs": null,
|
||||
"finish_reason": "stop"
|
||||
}
|
||||
],
|
||||
"usage": {
|
||||
"prompt_tokens": 1,
|
||||
"completion_tokens": 1,
|
||||
"total_tokens": 2
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 6.3 `POST /v1/chat/completions`
|
||||
|
||||
Supported request fields initially:
|
||||
|
||||
```json
|
||||
{
|
||||
"model": "Qwen/Qwen3.5-0.8B",
|
||||
"messages": [
|
||||
{"role": "user", "content": "Hello"}
|
||||
],
|
||||
"max_tokens": 64,
|
||||
"temperature": 0.7,
|
||||
"top_p": 0.95,
|
||||
"stop": null,
|
||||
"stream": false
|
||||
}
|
||||
```
|
||||
|
||||
The coordinator should use the HuggingFace tokenizer chat template if available.
|
||||
|
||||
`stream: true` may return a clear unsupported error in the initial prototype.
|
||||
|
||||
### 6.4 Not-Ready Error
|
||||
|
||||
If a generation request arrives before enough nodes are connected and loaded:
|
||||
|
||||
```json
|
||||
{
|
||||
"error": {
|
||||
"message": "Model is not ready. Required nodes: 3, connected ready nodes: 1",
|
||||
"type": "cluster_not_ready"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
HTTP status: `503`.
|
||||
|
||||
## 7. Node Socket Protocol
|
||||
|
||||
Use persistent TCP sockets with asyncio streams.
|
||||
|
||||
Encoding:
|
||||
|
||||
```text
|
||||
[8-byte unsigned big-endian payload length][msgpack payload]
|
||||
```
|
||||
|
||||
Large tensor payloads are sent as chunked binary data inside protocol messages or as msgpack metadata followed by raw bytes.
|
||||
|
||||
### 7.1 Message Envelope
|
||||
|
||||
Every message should include:
|
||||
|
||||
```json
|
||||
{
|
||||
"type": "MESSAGE_TYPE",
|
||||
"request_id": "optional-request-id",
|
||||
"seq": 1,
|
||||
"payload": {}
|
||||
}
|
||||
```
|
||||
|
||||
### 7.2 Core Message Types
|
||||
|
||||
Coordinator/node lifecycle:
|
||||
|
||||
```text
|
||||
HELLO
|
||||
HELLO_ACK
|
||||
ASSIGNMENT
|
||||
WEIGHT_CHUNK
|
||||
WEIGHTS_COMPLETE
|
||||
LOAD_COMPLETE
|
||||
LOAD_FAILED
|
||||
PING
|
||||
PONG
|
||||
ERROR
|
||||
```
|
||||
|
||||
Inference:
|
||||
|
||||
```text
|
||||
CLEAR_CACHE
|
||||
RUN_PREFILL
|
||||
RUN_DECODE
|
||||
HIDDEN_STATE
|
||||
INFERENCE_ERROR
|
||||
```
|
||||
|
||||
### 7.3 `HELLO`
|
||||
|
||||
Sent by node immediately after connection.
|
||||
|
||||
Example:
|
||||
|
||||
```json
|
||||
{
|
||||
"type": "HELLO",
|
||||
"payload": {
|
||||
"node_id": "macbook-pro-1",
|
||||
"hostname": "macbook-pro.local",
|
||||
"python_version": "3.10.13",
|
||||
"torch_version": "2.x",
|
||||
"platform": "darwin",
|
||||
"devices": [
|
||||
{
|
||||
"id": "mps",
|
||||
"type": "mps",
|
||||
"name": "Apple Silicon MPS",
|
||||
"total_memory": null,
|
||||
"free_memory": null
|
||||
}
|
||||
],
|
||||
"selected_device": "mps",
|
||||
"max_memory_bytes": null
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
CUDA example device:
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "cuda:0",
|
||||
"type": "cuda",
|
||||
"name": "NVIDIA GeForce RTX 4090",
|
||||
"total_memory": 25757220864,
|
||||
"free_memory": 23000000000
|
||||
}
|
||||
```
|
||||
|
||||
### 7.4 `ASSIGNMENT`
|
||||
|
||||
Sent by coordinator after planning.
|
||||
|
||||
```json
|
||||
{
|
||||
"type": "ASSIGNMENT",
|
||||
"payload": {
|
||||
"model_id": "Qwen/Qwen3.5-0.8B",
|
||||
"architecture": "qwen",
|
||||
"quant": "fp16",
|
||||
"compute_dtype": "fp16",
|
||||
"layer_start": 0,
|
||||
"layer_end_exclusive": 8,
|
||||
"config": {},
|
||||
"tensor_count": 128,
|
||||
"total_weight_bytes": 123456789
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 7.5 Tensor Transfer
|
||||
|
||||
Each tensor chunk message includes metadata:
|
||||
|
||||
```json
|
||||
{
|
||||
"type": "WEIGHT_CHUNK",
|
||||
"payload": {
|
||||
"tensor_name": "model.layers.0.self_attn.q_proj.weight",
|
||||
"dtype": "float16",
|
||||
"shape": [1024, 1024],
|
||||
"chunk_index": 0,
|
||||
"chunk_count": 4,
|
||||
"offset": 0,
|
||||
"data": "binary payload or external raw section"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
For very large tensors, the preferred format is:
|
||||
|
||||
```text
|
||||
[frame length][msgpack metadata][raw bytes referenced by metadata]
|
||||
```
|
||||
|
||||
The implementation should keep this hidden behind `protocol/tensors.py`.
|
||||
|
||||
### 7.6 Inference Messages
|
||||
|
||||
`RUN_PREFILL`:
|
||||
|
||||
```json
|
||||
{
|
||||
"type": "RUN_PREFILL",
|
||||
"request_id": "req-1",
|
||||
"payload": {
|
||||
"position_start": 0,
|
||||
"input_length": 42,
|
||||
"hidden_state": "tensor payload"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`RUN_DECODE`:
|
||||
|
||||
```json
|
||||
{
|
||||
"type": "RUN_DECODE",
|
||||
"request_id": "req-1",
|
||||
"payload": {
|
||||
"position": 42,
|
||||
"hidden_state": "tensor payload"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`HIDDEN_STATE`:
|
||||
|
||||
```json
|
||||
{
|
||||
"type": "HIDDEN_STATE",
|
||||
"request_id": "req-1",
|
||||
"payload": {
|
||||
"hidden_state": "tensor payload"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## 8. Model Support
|
||||
|
||||
### 8.1 Initial Architecture
|
||||
|
||||
Initial implementation should support Qwen2/Qwen2.5-style decoder-only causal LMs. The first implementation does not support Qwen3.5 hybrid models with `linear_attention`/GatedDeltaNet layers.
|
||||
|
||||
Required components:
|
||||
|
||||
- Token embedding.
|
||||
- Stacked transformer decoder blocks.
|
||||
- RMSNorm.
|
||||
- Rotary position embeddings.
|
||||
- Grouped-query attention.
|
||||
- Causal attention mask.
|
||||
- Gated MLP/SwiGLU.
|
||||
- Final RMSNorm.
|
||||
- LM head.
|
||||
- Tied or untied output embeddings.
|
||||
|
||||
### 8.2 HuggingFace Files
|
||||
|
||||
Coordinator should support models with:
|
||||
|
||||
```text
|
||||
config.json
|
||||
tokenizer.json/tokenizer.model/tokenizer_config.json
|
||||
model.safetensors or model-00001-of-000xx.safetensors
|
||||
model.safetensors.index.json, if sharded
|
||||
```
|
||||
|
||||
Use libraries:
|
||||
|
||||
- `huggingface_hub`
|
||||
- `safetensors`
|
||||
- `transformers` for tokenizer/config only where possible
|
||||
- `torch`
|
||||
|
||||
### 8.3 Weight Name Mapping
|
||||
|
||||
For Qwen-style models, expected tensor names include patterns like:
|
||||
|
||||
```text
|
||||
model.embed_tokens.weight
|
||||
model.layers.{i}.input_layernorm.weight
|
||||
model.layers.{i}.self_attn.q_proj.weight
|
||||
model.layers.{i}.self_attn.k_proj.weight
|
||||
model.layers.{i}.self_attn.v_proj.weight
|
||||
model.layers.{i}.self_attn.o_proj.weight
|
||||
model.layers.{i}.post_attention_layernorm.weight
|
||||
model.layers.{i}.mlp.gate_proj.weight
|
||||
model.layers.{i}.mlp.up_proj.weight
|
||||
model.layers.{i}.mlp.down_proj.weight
|
||||
model.norm.weight
|
||||
lm_head.weight
|
||||
```
|
||||
|
||||
The model loader should validate required tensors before accepting a model.
|
||||
|
||||
## 9. Model Planning and Splitting
|
||||
|
||||
The coordinator computes layer assignments from model metadata and node capacity.
|
||||
|
||||
### 9.1 Inputs
|
||||
|
||||
- Number of transformer layers.
|
||||
- Per-layer tensor byte sizes.
|
||||
- Quantization mode.
|
||||
- Max nodes.
|
||||
- Connected node capabilities.
|
||||
- Optional target memory hint.
|
||||
|
||||
### 9.2 Rules
|
||||
|
||||
- Split only on full transformer layer boundaries.
|
||||
- Assign contiguous layer ranges.
|
||||
- Preserve layer order.
|
||||
- Coordinator keeps embeddings, final norm, and lm head.
|
||||
- Required nodes must be connected and loaded before generation.
|
||||
- If insufficient nodes are available, API returns `cluster_not_ready`.
|
||||
|
||||
### 9.3 First Planner Algorithm
|
||||
|
||||
Simple deterministic version:
|
||||
|
||||
1. Compute total transformer layer bytes after quantization.
|
||||
2. Estimate bytes per layer.
|
||||
3. Determine number of partitions as `min(max_nodes, num_layers)`.
|
||||
4. If node memory information is available, reduce/increase partitions so each assignment fits.
|
||||
5. Otherwise split evenly by layer byte size.
|
||||
6. Assign partitions to the first compatible ready nodes.
|
||||
|
||||
### 9.4 Future Planner Improvements
|
||||
|
||||
- Benchmark node speed and assign more layers to faster GPUs.
|
||||
- Prefer CUDA nodes for larger shards.
|
||||
- Consider network latency and bandwidth.
|
||||
- Replicate small layers for resilience.
|
||||
- Rebalance between generations.
|
||||
|
||||
## 10. Quantization
|
||||
|
||||
Quantization must work on CUDA, MPS, and CPU. Therefore the baseline implementation should avoid CUDA-only dependencies such as bitsandbytes.
|
||||
|
||||
Supported modes:
|
||||
|
||||
```text
|
||||
fp16
|
||||
int8
|
||||
int4
|
||||
```
|
||||
|
||||
### 10.1 `fp16`
|
||||
|
||||
- Store weights as `torch.float16`.
|
||||
- Compute in `float16` by default.
|
||||
- Works on CUDA and MPS.
|
||||
- CPU fallback may use `float32` internally if needed.
|
||||
|
||||
### 10.2 Portable `int8`
|
||||
|
||||
Use symmetric per-output-channel weight-only quantization.
|
||||
|
||||
For a linear weight `W` shaped `[out_features, in_features]`:
|
||||
|
||||
```text
|
||||
scale[out_features] = max(abs(W[row])) / 127
|
||||
qweight[row] = round(W[row] / scale[row]).clamp(-127, 127).int8
|
||||
```
|
||||
|
||||
Forward path:
|
||||
|
||||
```text
|
||||
W_dequant = qweight.float() * scale[:, None]
|
||||
y = x @ W_dequant.T
|
||||
```
|
||||
|
||||
This is portable but not maximally fast.
|
||||
|
||||
### 10.3 Portable `int4`
|
||||
|
||||
Use group-wise weight-only quantization.
|
||||
|
||||
Suggested default group size: `128`.
|
||||
|
||||
Store:
|
||||
|
||||
```text
|
||||
packed_qweight: uint8
|
||||
scale: float16/float32 per group
|
||||
zero_point: optional
|
||||
metadata: original shape, group size, packing order
|
||||
```
|
||||
|
||||
Forward path:
|
||||
|
||||
1. Unpack int4 values.
|
||||
2. Dequantize to compute dtype.
|
||||
3. Perform normal PyTorch matmul.
|
||||
|
||||
This is designed for correctness and portability, not peak speed.
|
||||
|
||||
### 10.4 Quantization Timing
|
||||
|
||||
Preferred prototype behavior:
|
||||
|
||||
- Coordinator loads original safetensors.
|
||||
- Coordinator quantizes tensors before sending to nodes if `int8` or `int4` is selected.
|
||||
- Nodes receive already-quantized tensors plus quantization metadata.
|
||||
- Coordinator also quantizes/loads its own embedding/lm_head as needed.
|
||||
|
||||
## 11. Generation Algorithm
|
||||
|
||||
### 11.1 Prefill
|
||||
|
||||
For prompt token IDs of length `N`:
|
||||
|
||||
1. Coordinator computes embeddings: `[1, N, hidden_size]`.
|
||||
2. Coordinator sends hidden state to first node with position start `0`.
|
||||
3. Each node runs its assigned layers across the full sequence.
|
||||
4. Each node initializes KV cache for its layers.
|
||||
5. Final node returns hidden state to coordinator.
|
||||
6. Coordinator applies final norm and lm head to the last token.
|
||||
7. Coordinator samples next token.
|
||||
|
||||
### 11.2 Decode
|
||||
|
||||
For each generated token:
|
||||
|
||||
1. Coordinator embeds last token: `[1, 1, hidden_size]`.
|
||||
2. Coordinator sends hidden state to first node with current position.
|
||||
3. Each node runs one-token decode using local KV cache.
|
||||
4. Each node appends to its KV cache.
|
||||
5. Final node returns hidden state to coordinator.
|
||||
6. Coordinator computes logits and samples next token.
|
||||
7. Stop if EOS, stop sequence, or `max_tokens` reached.
|
||||
|
||||
### 11.3 Sampling
|
||||
|
||||
Initial sampler supports:
|
||||
|
||||
- Greedy when `temperature == 0`.
|
||||
- Temperature scaling.
|
||||
- Top-p nucleus sampling.
|
||||
- EOS handling.
|
||||
- Stop strings after decoding.
|
||||
|
||||
Future additions:
|
||||
|
||||
- Top-k.
|
||||
- Repetition penalty.
|
||||
- Frequency/presence penalties.
|
||||
- Logprobs.
|
||||
|
||||
## 12. Device Support
|
||||
|
||||
### 12.1 Device Auto Detection
|
||||
|
||||
Node device priority when `--device auto`:
|
||||
|
||||
1. CUDA if available.
|
||||
2. MPS if available.
|
||||
3. CPU fallback.
|
||||
|
||||
### 12.2 CUDA
|
||||
|
||||
Use:
|
||||
|
||||
```python
|
||||
torch.cuda.is_available()
|
||||
torch.cuda.get_device_properties(index)
|
||||
torch.cuda.mem_get_info(index)
|
||||
```
|
||||
|
||||
### 12.3 Apple MPS
|
||||
|
||||
Use:
|
||||
|
||||
```python
|
||||
torch.backends.mps.is_available()
|
||||
torch.device("mps")
|
||||
```
|
||||
|
||||
MPS memory reporting is limited, so allow `--max-memory-gb` override.
|
||||
|
||||
### 12.4 CPU
|
||||
|
||||
CPU is allowed for testing and fallback, but may be slow.
|
||||
|
||||
## 13. Package Structure
|
||||
|
||||
Recommended source tree:
|
||||
|
||||
```text
|
||||
truecluster/
|
||||
__init__.py
|
||||
cli.py
|
||||
|
||||
cluster/
|
||||
__init__.py
|
||||
server.py # node TCP server
|
||||
api.py # FastAPI OpenAI-compatible API
|
||||
planner.py # layer splitting
|
||||
scheduler.py # generation orchestration
|
||||
model_store.py # HF/safetensors loading
|
||||
sampler.py
|
||||
state.py
|
||||
|
||||
node/
|
||||
__init__.py
|
||||
client.py # connects to cluster
|
||||
runtime.py # owns assigned layers + cache
|
||||
device.py
|
||||
|
||||
model/
|
||||
__init__.py
|
||||
qwen.py # minimal Qwen implementation
|
||||
layers.py
|
||||
rotary.py
|
||||
kv_cache.py
|
||||
quant.py
|
||||
tensor_names.py
|
||||
|
||||
protocol/
|
||||
__init__.py
|
||||
framing.py
|
||||
messages.py
|
||||
tensors.py
|
||||
|
||||
tests/
|
||||
test_single_node_matches_transformers.py
|
||||
test_protocol.py
|
||||
test_quant.py
|
||||
test_planner.py
|
||||
```
|
||||
|
||||
Project metadata:
|
||||
|
||||
```text
|
||||
pyproject.toml
|
||||
README.md
|
||||
SPEC.md
|
||||
```
|
||||
|
||||
## 14. Suggested Dependencies
|
||||
|
||||
Runtime:
|
||||
|
||||
```text
|
||||
torch
|
||||
transformers
|
||||
huggingface_hub
|
||||
safetensors
|
||||
fastapi
|
||||
uvicorn[standard]
|
||||
typer
|
||||
msgpack
|
||||
pydantic
|
||||
numpy
|
||||
tqdm
|
||||
```
|
||||
|
||||
Development/test:
|
||||
|
||||
```text
|
||||
pytest
|
||||
pytest-asyncio
|
||||
httpx
|
||||
ruff
|
||||
mypy optional
|
||||
```
|
||||
|
||||
Python version:
|
||||
|
||||
```text
|
||||
>=3.10,<3.13
|
||||
```
|
||||
|
||||
## 15. Validation and Testing
|
||||
|
||||
### 15.1 Correctness Test Against Transformers
|
||||
|
||||
Most important validation:
|
||||
|
||||
1. Load the target model with HuggingFace Transformers locally.
|
||||
2. Load the same model through TrueCluster with one local node.
|
||||
3. Run the same prompt.
|
||||
4. Compare final logits before sampling.
|
||||
5. Assert max difference is within tolerance for selected dtype.
|
||||
|
||||
### 15.2 Multi-node Local Test
|
||||
|
||||
Run on one machine:
|
||||
|
||||
```bash
|
||||
truecluster cluster --model ... --max-nodes 2
|
||||
truecluster node --cluster-host 127.0.0.1 --cluster-port 7001 --device cpu
|
||||
truecluster node --cluster-host 127.0.0.1 --cluster-port 7001 --device cpu
|
||||
```
|
||||
|
||||
Verify:
|
||||
|
||||
- Both nodes receive different layer ranges.
|
||||
- Prefill works.
|
||||
- Decode works.
|
||||
- Output matches single-node output within tolerance.
|
||||
|
||||
### 15.3 Mixed Hardware Test
|
||||
|
||||
Example:
|
||||
|
||||
```text
|
||||
coordinator: Mac or Linux host
|
||||
node 1: Nvidia CUDA machine
|
||||
node 2: Apple Silicon Mac MPS machine
|
||||
```
|
||||
|
||||
Verify:
|
||||
|
||||
- Both nodes connect.
|
||||
- Assignments are sent.
|
||||
- Generation completes.
|
||||
|
||||
### 15.4 Quantization Tests
|
||||
|
||||
For `int8` and `int4`:
|
||||
|
||||
- Quantize/dequantize synthetic tensors.
|
||||
- Check shape preservation.
|
||||
- Check error bounds.
|
||||
- Run short generation and verify no crashes.
|
||||
|
||||
## 16. Implementation Phases
|
||||
|
||||
### Phase 1: Local Single-Process Model Proof
|
||||
|
||||
Deliverables:
|
||||
|
||||
- Minimal Qwen model implementation.
|
||||
- Safetensors loading.
|
||||
- Local full-model forward.
|
||||
- Logit comparison against Transformers.
|
||||
|
||||
### Phase 2: One Node Distributed Inference
|
||||
|
||||
Deliverables:
|
||||
|
||||
- Socket protocol.
|
||||
- Coordinator sends all transformer layers to one local node.
|
||||
- Node loads layers and runs them.
|
||||
- Coordinator keeps embedding/final norm/lm head.
|
||||
- `/v1/completions` works.
|
||||
|
||||
### Phase 3: Multi-node Layer Split
|
||||
|
||||
Deliverables:
|
||||
|
||||
- Planner assigns contiguous layer ranges.
|
||||
- Multiple nodes are supported.
|
||||
- Distributed KV cache works.
|
||||
- Not-ready errors work.
|
||||
|
||||
### Phase 4: Mixed CUDA/MPS Support
|
||||
|
||||
Deliverables:
|
||||
|
||||
- Device detection.
|
||||
- CUDA execution.
|
||||
- MPS execution.
|
||||
- CPU fallback.
|
||||
- Mixed Nvidia/Mac cluster generation test.
|
||||
|
||||
### Phase 5: Portable Quantization
|
||||
|
||||
Deliverables:
|
||||
|
||||
- `fp16` baseline.
|
||||
- Portable `int8` linear.
|
||||
- Portable `int4` linear.
|
||||
- Quantized weight transfer.
|
||||
- CLI `--quant` option.
|
||||
|
||||
### Phase 6: API Polish
|
||||
|
||||
Deliverables:
|
||||
|
||||
- `/v1/models`.
|
||||
- `/v1/chat/completions`.
|
||||
- Stop sequence support.
|
||||
- Usage accounting.
|
||||
- Better OpenAI-compatible errors.
|
||||
|
||||
## 17. Initial Acceptance Criteria
|
||||
|
||||
A prototype is considered working when:
|
||||
|
||||
1. A cluster can be started with a HuggingFace safetensors Qwen-style model.
|
||||
2. A node can connect to the cluster with no local model files.
|
||||
3. The cluster sends layer weights to the node over the socket.
|
||||
4. The node loads assigned layers on CUDA, MPS, or CPU.
|
||||
5. `/v1/models` returns the loaded model.
|
||||
6. `/v1/completions` generates text through the distributed pipeline.
|
||||
7. `/v1/chat/completions` works for simple chat prompts.
|
||||
8. If insufficient nodes are ready, API returns a clear `cluster_not_ready` error.
|
||||
9. One local-node output matches Transformers logits within reasonable dtype tolerance.
|
||||
10. Multi-node local CPU test works.
|
||||
|
||||
## 18. Key Design Decisions
|
||||
|
||||
- Use pipeline parallelism, not tensor parallelism.
|
||||
- Use contiguous layer ranges only.
|
||||
- Keep tokenizer, embeddings, final norm, lm head, and sampler on the coordinator.
|
||||
- Send weights once at node assignment time.
|
||||
- Send hidden states during generation.
|
||||
- Store KV cache on worker nodes.
|
||||
- Implement portable quantization instead of relying on CUDA-only libraries.
|
||||
- Start with Qwen2/Qwen2.5-style causal LMs only.
|
||||
- Optimize for correctness and clean architecture before speed.
|
||||
Reference in New Issue
Block a user