Inital commit

This commit is contained in:
2026-06-04 18:11:26 -05:00
commit 9bd7f5838e
24 changed files with 2454 additions and 0 deletions
+979
View File
@@ -0,0 +1,979 @@
# TrueCluster Prototype Specification
## 1. Goal
TrueCluster is a Python 3.10 prototype for distributed LLM inference across a small cluster of heterogeneous machines. It allows a coordinator/cluster process to load a HuggingFace safetensors model, split transformer layers across connected worker nodes, send each node only the weights it needs over a socket connection, and expose a simple OpenAI-compatible HTTP API for generation.
The initial target is small Qwen2/Qwen2.5-style decoder-only models around 0.5B-1.5B parameters, with first-class support for mixed Nvidia CUDA and Apple Silicon Mac MPS nodes in the same cluster. Qwen3.5 hybrid linear-attention models are a later target and are not part of the first fp16 prototype.
Primary prototype goals:
- Python 3.10.
- Easy CLI for running a cluster or node.
- Coordinator loads one model at a time.
- Nodes connect to the coordinator by host/port.
- Coordinator sends assigned model shards to nodes over sockets.
- Nodes do not need local model files.
- Support HuggingFace safetensors models.
- Initial architecture target: Qwen2/Qwen2.5-style causal language models.
- Support `fp16`, portable `int8`, and portable `int4` weight-only quantization modes.
- OpenAI-compatible unauthenticated HTTP API.
- Efficient generation by transferring activations during inference, not weights.
## 2. Non-goals for the Initial Prototype
The first prototype will intentionally avoid several advanced features:
- No multi-model serving.
- No authentication.
- No request batching.
- No tensor parallelism across machines.
- No expert parallelism.
- No dynamic model hot-swapping.
- No high-performance CUDA-only quantization kernels as a requirement.
- No dependency on every node having HuggingFace access.
- No web UI.
- No streaming responses in the first milestone.
- No fault-tolerant recovery during an active generation.
These can be added after a correct baseline works.
## 3. High-Level Architecture
TrueCluster uses pipeline-parallel inference.
The coordinator owns:
- CLI entry point for `cluster`.
- HTTP API.
- Tokenizer.
- Sampling logic.
- HuggingFace model metadata/config loading.
- Model weight loading from safetensors.
- Model split planning.
- Embedding layer.
- Final normalization.
- LM head.
- Node registry and orchestration.
Each node owns:
- CLI entry point for `node`.
- Persistent socket connection to coordinator.
- Device detection and selection.
- One contiguous range of transformer layers.
- KV cache for its assigned layers.
- Local forward execution on CUDA, MPS, or CPU.
Generation path:
```text
HTTP request
-> coordinator tokenizes prompt
-> coordinator runs embedding
-> hidden states sent to node 1
-> node 1 runs assigned layers
-> hidden states sent to node 2
-> ...
-> final hidden states returned to coordinator
-> coordinator runs final norm + lm_head
-> coordinator samples next token
-> repeat until completion
```
Weights are transferred once during assignment. During generation, only hidden states and small metadata are passed between coordinator and nodes.
## 4. Execution Model
### 4.1 Pipeline Parallelism
The model is split by complete transformer blocks. Each node receives a contiguous set of layers:
```text
coordinator:
embed_tokens
final_norm
lm_head
node 1:
layers 0-7
node 2:
layers 8-15
node 3:
layers 16-23
```
This is simpler and more reliable than tensor parallelism for heterogeneous machines and normal Ethernet/Wi-Fi networks.
### 4.2 KV Cache Ownership
KV cache is stored on the node that owns the relevant layers.
For example:
```text
node 1 cache: layers 0-7
node 2 cache: layers 8-15
node 3 cache: layers 16-23
```
During prefill, each node creates cache entries for its layers. During decode, each node appends one token of keys/values to its cache.
### 4.3 Request Concurrency
Initial prototype supports one active generation at a time per cluster.
Reason: distributed KV cache management is much easier with a single active request. Later versions can introduce request IDs, cache slots, and batching.
## 5. CLI Design
Use `typer` for the CLI.
Package command:
```bash
truecluster
```
### 5.1 Cluster Command
Example:
```bash
truecluster cluster \
--model Qwen/Qwen3.5-0.8B \
--node-host 0.0.0.0 \
--node-port 7001 \
--api-host 0.0.0.0 \
--api-port 8000 \
--max-nodes 4 \
--quant fp16
```
Options:
```text
--model TEXT HuggingFace model id or local path.
--node-host TEXT Host/IP for worker-node socket server. Default: 0.0.0.0
--node-port INT Port for worker-node socket server. Default: 7001
--api-host TEXT Host/IP for HTTP API. Default: 0.0.0.0
--api-port INT HTTP API port. Default: 8000
--max-nodes INT Maximum number of worker nodes to use.
--quant [fp16|int8|int4] Weight mode. Default: fp16
--dtype [fp16|bf16|fp32] Compute dtype preference. Default: fp16
--target-node-memory-gb FLOAT Optional planning hint if node memory is unknown.
--trust-remote-code BOOL HuggingFace trust_remote_code. Default: false
--hf-cache-dir PATH Optional HuggingFace cache directory.
--log-level TEXT Default: info
```
Behavior:
1. Resolve/download model.
2. Load config and tokenizer.
3. Load safetensors into coordinator RAM or memory-mapped index.
4. Build model tensor index.
5. Start node socket server.
6. Start HTTP API.
7. Wait for enough nodes.
8. Assign layer ranges and transmit weights.
9. Mark cluster ready.
### 5.2 Node Command
Example:
```bash
truecluster node \
--cluster-host 192.168.1.50 \
--cluster-port 7001 \
--device auto
```
Options:
```text
--cluster-host TEXT Coordinator node socket host.
--cluster-port INT Coordinator node socket port.
--device TEXT auto, cuda, cuda:0, mps, or cpu. Default: auto
--node-id TEXT Optional stable node id.
--work-dir PATH Temporary local directory for received weights/cache.
--max-memory-gb FLOAT Optional memory capability override.
--log-level TEXT Default: info
```
Behavior:
1. Detect hardware and PyTorch backends.
2. Connect to coordinator socket.
3. Send `HELLO` capability message.
4. Wait for assignment.
5. Receive config and layer weights.
6. Build local model shard.
7. Mark itself ready.
8. Serve prefill/decode requests over the persistent connection.
## 6. HTTP API
Use FastAPI and Uvicorn.
The API is unauthenticated for the prototype.
### 6.1 `GET /v1/models`
Returns the single loaded model.
Example response:
```json
{
"object": "list",
"data": [
{
"id": "Qwen/Qwen3.5-0.8B",
"object": "model",
"created": 0,
"owned_by": "truecluster"
}
]
}
```
### 6.2 `POST /v1/completions`
Supported request fields initially:
```json
{
"model": "Qwen/Qwen3.5-0.8B",
"prompt": "Hello",
"max_tokens": 64,
"temperature": 0.7,
"top_p": 0.95,
"stop": null
}
```
Response should be OpenAI-compatible enough for common clients:
```json
{
"id": "cmpl-...",
"object": "text_completion",
"created": 0,
"model": "Qwen/Qwen3.5-0.8B",
"choices": [
{
"text": " world",
"index": 0,
"logprobs": null,
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 1,
"completion_tokens": 1,
"total_tokens": 2
}
}
```
### 6.3 `POST /v1/chat/completions`
Supported request fields initially:
```json
{
"model": "Qwen/Qwen3.5-0.8B",
"messages": [
{"role": "user", "content": "Hello"}
],
"max_tokens": 64,
"temperature": 0.7,
"top_p": 0.95,
"stop": null,
"stream": false
}
```
The coordinator should use the HuggingFace tokenizer chat template if available.
`stream: true` may return a clear unsupported error in the initial prototype.
### 6.4 Not-Ready Error
If a generation request arrives before enough nodes are connected and loaded:
```json
{
"error": {
"message": "Model is not ready. Required nodes: 3, connected ready nodes: 1",
"type": "cluster_not_ready"
}
}
```
HTTP status: `503`.
## 7. Node Socket Protocol
Use persistent TCP sockets with asyncio streams.
Encoding:
```text
[8-byte unsigned big-endian payload length][msgpack payload]
```
Large tensor payloads are sent as chunked binary data inside protocol messages or as msgpack metadata followed by raw bytes.
### 7.1 Message Envelope
Every message should include:
```json
{
"type": "MESSAGE_TYPE",
"request_id": "optional-request-id",
"seq": 1,
"payload": {}
}
```
### 7.2 Core Message Types
Coordinator/node lifecycle:
```text
HELLO
HELLO_ACK
ASSIGNMENT
WEIGHT_CHUNK
WEIGHTS_COMPLETE
LOAD_COMPLETE
LOAD_FAILED
PING
PONG
ERROR
```
Inference:
```text
CLEAR_CACHE
RUN_PREFILL
RUN_DECODE
HIDDEN_STATE
INFERENCE_ERROR
```
### 7.3 `HELLO`
Sent by node immediately after connection.
Example:
```json
{
"type": "HELLO",
"payload": {
"node_id": "macbook-pro-1",
"hostname": "macbook-pro.local",
"python_version": "3.10.13",
"torch_version": "2.x",
"platform": "darwin",
"devices": [
{
"id": "mps",
"type": "mps",
"name": "Apple Silicon MPS",
"total_memory": null,
"free_memory": null
}
],
"selected_device": "mps",
"max_memory_bytes": null
}
}
```
CUDA example device:
```json
{
"id": "cuda:0",
"type": "cuda",
"name": "NVIDIA GeForce RTX 4090",
"total_memory": 25757220864,
"free_memory": 23000000000
}
```
### 7.4 `ASSIGNMENT`
Sent by coordinator after planning.
```json
{
"type": "ASSIGNMENT",
"payload": {
"model_id": "Qwen/Qwen3.5-0.8B",
"architecture": "qwen",
"quant": "fp16",
"compute_dtype": "fp16",
"layer_start": 0,
"layer_end_exclusive": 8,
"config": {},
"tensor_count": 128,
"total_weight_bytes": 123456789
}
}
```
### 7.5 Tensor Transfer
Each tensor chunk message includes metadata:
```json
{
"type": "WEIGHT_CHUNK",
"payload": {
"tensor_name": "model.layers.0.self_attn.q_proj.weight",
"dtype": "float16",
"shape": [1024, 1024],
"chunk_index": 0,
"chunk_count": 4,
"offset": 0,
"data": "binary payload or external raw section"
}
}
```
For very large tensors, the preferred format is:
```text
[frame length][msgpack metadata][raw bytes referenced by metadata]
```
The implementation should keep this hidden behind `protocol/tensors.py`.
### 7.6 Inference Messages
`RUN_PREFILL`:
```json
{
"type": "RUN_PREFILL",
"request_id": "req-1",
"payload": {
"position_start": 0,
"input_length": 42,
"hidden_state": "tensor payload"
}
}
```
`RUN_DECODE`:
```json
{
"type": "RUN_DECODE",
"request_id": "req-1",
"payload": {
"position": 42,
"hidden_state": "tensor payload"
}
}
```
`HIDDEN_STATE`:
```json
{
"type": "HIDDEN_STATE",
"request_id": "req-1",
"payload": {
"hidden_state": "tensor payload"
}
}
```
## 8. Model Support
### 8.1 Initial Architecture
Initial implementation should support Qwen2/Qwen2.5-style decoder-only causal LMs. The first implementation does not support Qwen3.5 hybrid models with `linear_attention`/GatedDeltaNet layers.
Required components:
- Token embedding.
- Stacked transformer decoder blocks.
- RMSNorm.
- Rotary position embeddings.
- Grouped-query attention.
- Causal attention mask.
- Gated MLP/SwiGLU.
- Final RMSNorm.
- LM head.
- Tied or untied output embeddings.
### 8.2 HuggingFace Files
Coordinator should support models with:
```text
config.json
tokenizer.json/tokenizer.model/tokenizer_config.json
model.safetensors or model-00001-of-000xx.safetensors
model.safetensors.index.json, if sharded
```
Use libraries:
- `huggingface_hub`
- `safetensors`
- `transformers` for tokenizer/config only where possible
- `torch`
### 8.3 Weight Name Mapping
For Qwen-style models, expected tensor names include patterns like:
```text
model.embed_tokens.weight
model.layers.{i}.input_layernorm.weight
model.layers.{i}.self_attn.q_proj.weight
model.layers.{i}.self_attn.k_proj.weight
model.layers.{i}.self_attn.v_proj.weight
model.layers.{i}.self_attn.o_proj.weight
model.layers.{i}.post_attention_layernorm.weight
model.layers.{i}.mlp.gate_proj.weight
model.layers.{i}.mlp.up_proj.weight
model.layers.{i}.mlp.down_proj.weight
model.norm.weight
lm_head.weight
```
The model loader should validate required tensors before accepting a model.
## 9. Model Planning and Splitting
The coordinator computes layer assignments from model metadata and node capacity.
### 9.1 Inputs
- Number of transformer layers.
- Per-layer tensor byte sizes.
- Quantization mode.
- Max nodes.
- Connected node capabilities.
- Optional target memory hint.
### 9.2 Rules
- Split only on full transformer layer boundaries.
- Assign contiguous layer ranges.
- Preserve layer order.
- Coordinator keeps embeddings, final norm, and lm head.
- Required nodes must be connected and loaded before generation.
- If insufficient nodes are available, API returns `cluster_not_ready`.
### 9.3 First Planner Algorithm
Simple deterministic version:
1. Compute total transformer layer bytes after quantization.
2. Estimate bytes per layer.
3. Determine number of partitions as `min(max_nodes, num_layers)`.
4. If node memory information is available, reduce/increase partitions so each assignment fits.
5. Otherwise split evenly by layer byte size.
6. Assign partitions to the first compatible ready nodes.
### 9.4 Future Planner Improvements
- Benchmark node speed and assign more layers to faster GPUs.
- Prefer CUDA nodes for larger shards.
- Consider network latency and bandwidth.
- Replicate small layers for resilience.
- Rebalance between generations.
## 10. Quantization
Quantization must work on CUDA, MPS, and CPU. Therefore the baseline implementation should avoid CUDA-only dependencies such as bitsandbytes.
Supported modes:
```text
fp16
int8
int4
```
### 10.1 `fp16`
- Store weights as `torch.float16`.
- Compute in `float16` by default.
- Works on CUDA and MPS.
- CPU fallback may use `float32` internally if needed.
### 10.2 Portable `int8`
Use symmetric per-output-channel weight-only quantization.
For a linear weight `W` shaped `[out_features, in_features]`:
```text
scale[out_features] = max(abs(W[row])) / 127
qweight[row] = round(W[row] / scale[row]).clamp(-127, 127).int8
```
Forward path:
```text
W_dequant = qweight.float() * scale[:, None]
y = x @ W_dequant.T
```
This is portable but not maximally fast.
### 10.3 Portable `int4`
Use group-wise weight-only quantization.
Suggested default group size: `128`.
Store:
```text
packed_qweight: uint8
scale: float16/float32 per group
zero_point: optional
metadata: original shape, group size, packing order
```
Forward path:
1. Unpack int4 values.
2. Dequantize to compute dtype.
3. Perform normal PyTorch matmul.
This is designed for correctness and portability, not peak speed.
### 10.4 Quantization Timing
Preferred prototype behavior:
- Coordinator loads original safetensors.
- Coordinator quantizes tensors before sending to nodes if `int8` or `int4` is selected.
- Nodes receive already-quantized tensors plus quantization metadata.
- Coordinator also quantizes/loads its own embedding/lm_head as needed.
## 11. Generation Algorithm
### 11.1 Prefill
For prompt token IDs of length `N`:
1. Coordinator computes embeddings: `[1, N, hidden_size]`.
2. Coordinator sends hidden state to first node with position start `0`.
3. Each node runs its assigned layers across the full sequence.
4. Each node initializes KV cache for its layers.
5. Final node returns hidden state to coordinator.
6. Coordinator applies final norm and lm head to the last token.
7. Coordinator samples next token.
### 11.2 Decode
For each generated token:
1. Coordinator embeds last token: `[1, 1, hidden_size]`.
2. Coordinator sends hidden state to first node with current position.
3. Each node runs one-token decode using local KV cache.
4. Each node appends to its KV cache.
5. Final node returns hidden state to coordinator.
6. Coordinator computes logits and samples next token.
7. Stop if EOS, stop sequence, or `max_tokens` reached.
### 11.3 Sampling
Initial sampler supports:
- Greedy when `temperature == 0`.
- Temperature scaling.
- Top-p nucleus sampling.
- EOS handling.
- Stop strings after decoding.
Future additions:
- Top-k.
- Repetition penalty.
- Frequency/presence penalties.
- Logprobs.
## 12. Device Support
### 12.1 Device Auto Detection
Node device priority when `--device auto`:
1. CUDA if available.
2. MPS if available.
3. CPU fallback.
### 12.2 CUDA
Use:
```python
torch.cuda.is_available()
torch.cuda.get_device_properties(index)
torch.cuda.mem_get_info(index)
```
### 12.3 Apple MPS
Use:
```python
torch.backends.mps.is_available()
torch.device("mps")
```
MPS memory reporting is limited, so allow `--max-memory-gb` override.
### 12.4 CPU
CPU is allowed for testing and fallback, but may be slow.
## 13. Package Structure
Recommended source tree:
```text
truecluster/
__init__.py
cli.py
cluster/
__init__.py
server.py # node TCP server
api.py # FastAPI OpenAI-compatible API
planner.py # layer splitting
scheduler.py # generation orchestration
model_store.py # HF/safetensors loading
sampler.py
state.py
node/
__init__.py
client.py # connects to cluster
runtime.py # owns assigned layers + cache
device.py
model/
__init__.py
qwen.py # minimal Qwen implementation
layers.py
rotary.py
kv_cache.py
quant.py
tensor_names.py
protocol/
__init__.py
framing.py
messages.py
tensors.py
tests/
test_single_node_matches_transformers.py
test_protocol.py
test_quant.py
test_planner.py
```
Project metadata:
```text
pyproject.toml
README.md
SPEC.md
```
## 14. Suggested Dependencies
Runtime:
```text
torch
transformers
huggingface_hub
safetensors
fastapi
uvicorn[standard]
typer
msgpack
pydantic
numpy
tqdm
```
Development/test:
```text
pytest
pytest-asyncio
httpx
ruff
mypy optional
```
Python version:
```text
>=3.10,<3.13
```
## 15. Validation and Testing
### 15.1 Correctness Test Against Transformers
Most important validation:
1. Load the target model with HuggingFace Transformers locally.
2. Load the same model through TrueCluster with one local node.
3. Run the same prompt.
4. Compare final logits before sampling.
5. Assert max difference is within tolerance for selected dtype.
### 15.2 Multi-node Local Test
Run on one machine:
```bash
truecluster cluster --model ... --max-nodes 2
truecluster node --cluster-host 127.0.0.1 --cluster-port 7001 --device cpu
truecluster node --cluster-host 127.0.0.1 --cluster-port 7001 --device cpu
```
Verify:
- Both nodes receive different layer ranges.
- Prefill works.
- Decode works.
- Output matches single-node output within tolerance.
### 15.3 Mixed Hardware Test
Example:
```text
coordinator: Mac or Linux host
node 1: Nvidia CUDA machine
node 2: Apple Silicon Mac MPS machine
```
Verify:
- Both nodes connect.
- Assignments are sent.
- Generation completes.
### 15.4 Quantization Tests
For `int8` and `int4`:
- Quantize/dequantize synthetic tensors.
- Check shape preservation.
- Check error bounds.
- Run short generation and verify no crashes.
## 16. Implementation Phases
### Phase 1: Local Single-Process Model Proof
Deliverables:
- Minimal Qwen model implementation.
- Safetensors loading.
- Local full-model forward.
- Logit comparison against Transformers.
### Phase 2: One Node Distributed Inference
Deliverables:
- Socket protocol.
- Coordinator sends all transformer layers to one local node.
- Node loads layers and runs them.
- Coordinator keeps embedding/final norm/lm head.
- `/v1/completions` works.
### Phase 3: Multi-node Layer Split
Deliverables:
- Planner assigns contiguous layer ranges.
- Multiple nodes are supported.
- Distributed KV cache works.
- Not-ready errors work.
### Phase 4: Mixed CUDA/MPS Support
Deliverables:
- Device detection.
- CUDA execution.
- MPS execution.
- CPU fallback.
- Mixed Nvidia/Mac cluster generation test.
### Phase 5: Portable Quantization
Deliverables:
- `fp16` baseline.
- Portable `int8` linear.
- Portable `int4` linear.
- Quantized weight transfer.
- CLI `--quant` option.
### Phase 6: API Polish
Deliverables:
- `/v1/models`.
- `/v1/chat/completions`.
- Stop sequence support.
- Usage accounting.
- Better OpenAI-compatible errors.
## 17. Initial Acceptance Criteria
A prototype is considered working when:
1. A cluster can be started with a HuggingFace safetensors Qwen-style model.
2. A node can connect to the cluster with no local model files.
3. The cluster sends layer weights to the node over the socket.
4. The node loads assigned layers on CUDA, MPS, or CPU.
5. `/v1/models` returns the loaded model.
6. `/v1/completions` generates text through the distributed pipeline.
7. `/v1/chat/completions` works for simple chat prompts.
8. If insufficient nodes are ready, API returns a clear `cluster_not_ready` error.
9. One local-node output matches Transformers logits within reasonable dtype tolerance.
10. Multi-node local CPU test works.
## 18. Key Design Decisions
- Use pipeline parallelism, not tensor parallelism.
- Use contiguous layer ranges only.
- Keep tokenizer, embeddings, final norm, lm head, and sampler on the coordinator.
- Send weights once at node assignment time.
- Send hidden states during generation.
- Store KV cache on worker nodes.
- Implement portable quantization instead of relying on CUDA-only libraries.
- Start with Qwen2/Qwen2.5-style causal LMs only.
- Optimize for correctness and clean architecture before speed.