13 Commits
7 changed files with 190 additions and 19 deletions
+23 -3
View File
@@ -14,13 +14,33 @@
{ {
"key": "LLAMA_SERVER_CMD_ARGS", "key": "LLAMA_SERVER_CMD_ARGS",
"input": { "input": {
"name": "Model Name", "name": "Command line arguments for llama-server",
"type": "string", "type": "string",
"description": "Launch command line arguments (argv) for the llama-server binary. Do not define the port.", "description": "Launch command line arguments (argv) for the llama-server binary. Do not define the port. If using caching, do not define -hf or -m here.",
"default": "-hf unsloth/gemma-3-270m-it-GGUF:Q6_K --ctx-size 4096", "default": "-hf unsloth/gemma-3-270m-it-GGUF:Q6_K --ctx-size 4096 -ngl 999",
"advanced": false "advanced": false
} }
}, },
{
"key": "LLAMA_CACHED_MODEL",
"input": {
"name": "Hugging Face Hub model name for cached model",
"type": "string",
"description": "Hugging Face Hub model name to use for the cached GGUF model. Leave empty to disable caching. Example: user/model-name",
"default": "",
"advanced": true
}
},
{
"key": "LLAMA_CACHED_GGUF_PATH",
"input": {
"name": "Path to GGUF file in the Hugging Face Hub model repository",
"type": "string",
"description": "Path to the GGUF file in the Hugging Face Hub model repository to use for caching. Example: model.gguf",
"default": "",
"advanced": true
}
},
{ {
"key": "MAX_CONCURRENCY", "key": "MAX_CONCURRENCY",
"input": { "input": {
+2 -2
View File
@@ -5,7 +5,7 @@
"input": { "input": {
"prompt": "Hi! Who are you?" "prompt": "Hi! Who are you?"
}, },
"timeout": 120000 "timeout": 60000
} }
], ],
"config": { "config": {
@@ -14,7 +14,7 @@
"env": [ "env": [
{ {
"key": "LLAMA_SERVER_CMD_ARGS", "key": "LLAMA_SERVER_CMD_ARGS",
"value": "-hf unsloth/gemma-3-270m-it-GGUF:Q6_K --ctx-size 4096" "value": "-hf unsloth/gemma-3-270m-it-GGUF:IQ2_XXS --ctx-size 512 -ngl 999"
} }
], ],
"allowedCudaVersions": [ "allowedCudaVersions": [
+5 -2
View File
@@ -13,10 +13,13 @@ The following OpenAI API endpoints are supported:
Streaming responses is also supported. Streaming responses is also supported.
**Important!** This project is still relatively new. Please [open a new issue](https://github.com/Jacob-ML/inference-worker/issues/new) if you encounter any problems in order to get help.
**This is a fork of [SvenBrnn's `runpod-worker-ollama`](https://github.com/SvenBrnn/runpod-worker-ollama).**
## Setup ## Setup
For the setup to work best, it is recommended to use a network volume attached to all workers which stores the model GGUFs and then reference those files in the launch arguments. To get the best performance out of this worker, it is recommended to use cached models. Please see the [cached models documentation](./docs/cached.md) for more information, this is **highly recommended and will save many resources**.
Make sure your RunPod worker has access to the network volume, i.e. is located in the correct data center.
## Configuration ## Configuration
+63
View File
@@ -0,0 +1,63 @@
# Using cached models
## Introduction
The classic way of loading a model from the Hugging Face Hub with the `LLAMA_SERVER_CMD_ARGS` is as follows:
```bash
-hf /path/to/model.gguf:Q4_K_M --ctx-size 4096 # etc...
```
However, this will cause every worker to download the model from the Hugging Face Hub every time it is started, which can be slow and inefficient.
A naive way to cache the model would be to store it on a network volume in RunPod and reference the model files this way:
```bash
-m /runpod-volume/model.gguf --ctx-size 4096 # etc...
```
Unfortunately, network volume performance is often not sufficient for loading large models, leading to long load times. RunPod introduced a [caching mechanism](https://docs.runpod.io/serverless/endpoints/model-caching) to solve this problem.
The `inference-worker` for llama.cpp now supports this caching mechanism.
## How to use the new caching mechanism
It ships the `src/find_cached.py` script which can be used to reference any Hugging Face model of your choice and get its cached path on the local worker storage.
Here is how the script can be used independently (which you will likely never need to do):
```bash
python3 src/find_cached.py HF_MODEL_ID GGUF_PATH_IN_REPO
```
Example:
```bash
python3 src/find_cached.py unsloth/gemma-3-270m-it-GGUF gemma-3-270m-it-Q8_0.gguf
```
Or, if your model is in a folder (an edge case nobody seems to be thinking about, driving me absolutely crazy):
```bash
python3 src/find_cached.py jacob-ml/jacob-24b models/jacob-24b-q4_k_m.gguf
```
We will now integrate this into our workflow. Hang tight.
## Step-by-step guide
1. First of all, please enter the Hugging Face URL of the model you want to use in RunPod's `Model` field of your worker settings.
Example: For the model `unsloth/gemma-3-270m-it-GGUF`, you would enter `https://huggingface.co/unsloth/gemma-3-270m-it-GGUF`.
2. Now, in the environment variables, do NOT enter the `-hf` argument as before and also do NOT define `-m` in the `LLAMA_SERVER_CMD_ARGS`. The inference worker will take care of that for you.
Instead, set the `LLAMA_CACHED_MODEL` to the model ID, a.e. `unsloth/gemma-3-270m-it-GGUF`. Then, set the `LLAMA_CACHED_GGUF_PATH` to the path of the GGUF file in the repository, e.g. `gemma-3-270m-it-Q8_0.gguf`.
3. Finally, in the `LLAMA_SERVER_CMD_ARGS`, you can now simply add the other arguments you want to use, e.g.:
```bash
--ctx-size 4096 --temp 0.7 --top-p 0.9
```
4. Done! The rest will be handled by the inference worker automatically. When the worker starts, it will resolve the cached model path and launch `llama-server` with the correct arguments.
-1
View File
@@ -21,7 +21,6 @@ Typical usage:
""" """
import json import json
import os
from dotenv import load_dotenv from dotenv import load_dotenv
from openai import OpenAI from openai import OpenAI
+59
View File
@@ -0,0 +1,59 @@
"""
Finds the full LLM GGUF path from the Hugging Face cache.
"""
import os
import argparse
CACHE_DIR = "/runpod-volume/huggingface-cache/hub"
def find_model_path(model_name, gguf_in_repo="model.gguf"):
"""
Find the path to a cached model.
Args:
model_name: The model name from Hugging Face
Returns:
The full path to the cached model, or None if not found
"""
cache_name = model_name.replace("/", "--")
snapshots_dir = os.path.join(
CACHE_DIR, f"models--{cache_name}", "snapshots"
)
if os.path.exists(snapshots_dir):
snapshots = os.listdir(snapshots_dir)
if snapshots:
return os.path.join(snapshots_dir, snapshots[0], gguf_in_repo)
return None
def main():
"""
Main function to find and print the model path.
"""
parser = argparse.ArgumentParser(
description="Find the full GGUF path from the Hugging Face cache."
)
parser.add_argument(
"model", type=str, help="The model name from Hugging Face"
)
parser.add_argument(
"path",
type=str,
help="The path to the GGUF file within the model repository",
)
args = parser.parse_args()
model_path = find_model_path(args.model, args.path)
print(model_path, end="")
if __name__ == "__main__":
main()
+38 -11
View File
@@ -14,15 +14,26 @@ cleanup() {
exit 0 exit 0
} }
# check if $LLAMA_SERVER_CMD_ARGS is set CACHED_LLAMA_ARGS=""
if [ -z "$LLAMA_SERVER_CMD_ARGS" ]; then
echo "start.sh: Warning: LLAMA_SERVER_CMD_ARGS is not set. Defaulting to -hf unsloth/gemma-3-270m-it-GGUF:Q6_K --ctx-size 4096" find_cached_path() {
LLAMA_SERVER_CMD_ARGS="-hf unsloth/gemma-3-270m-it-GGUF:Q6_K --ctx-size 4096" CACHED_LLAMA_ARGS="-m $(python ./find_cached.py $LLAMA_CACHED_MODEL $LLAMA_CACHED_GGUF_PATH)"
}
# check if $LLAMA_CACHED_MODEL is set and not empty
if [ -n "$LLAMA_CACHED_MODEL" ]; then
echo "start.sh: Caching is enabled. Finding cached model path..."
find_cached_path
echo "start.sh: Using cached model with arguments: $CACHED_LLAMA_ARGS"
else
echo "start.sh: WARNING: Caching is disabled. Please visit the inference-worker README and docs to learn more."
fi fi
# check if the substring /workspace is in LLAMA_SERVER_CMD_ARGS # check if $LLAMA_SERVER_CMD_ARGS is set
if [[ "$LLAMA_SERVER_CMD_ARGS" != *"/workspace"* ]]; then if [ -z "$LLAMA_SERVER_CMD_ARGS" ]; then
echo "start.sh: Tip: For reduced downloads and faster startup times, consider using a model stored in a network volume mounted to /workspace." echo "start.sh: Warning: LLAMA_SERVER_CMD_ARGS is not set. Defaulting to -hf unsloth/gemma-3-270m-it-GGUF:IQ2_XXS --ctx-size 512 -ngl 999"
LLAMA_SERVER_CMD_ARGS="-hf unsloth/gemma-3-270m-it-GGUF:IQ2_XXS --ctx-size 512 -ngl 999"
fi fi
# check if the substring --port is in LLAMA_SERVER_CMD_ARGS and if yes, raise an error: # check if the substring --port is in LLAMA_SERVER_CMD_ARGS and if yes, raise an error:
@@ -43,17 +54,19 @@ echo "start.sh: Stopping existing llama-server instances (if any)..."
} }
# we have a string with all the command line arguments in the env var LLAMA_SERVER_CMD_ARGS; # we have a string with all the command line arguments in the env var LLAMA_SERVER_CMD_ARGS;
# it contains a.e. "-hf modelname --ctx-size 4096". # it contains a.e. "-hf modelname --ctx-size 4096 -ngl 999".
echo "start.sh: Running llama-server $LLAMA_SERVER_CMD_ARGS --port 3098" echo "start.sh: Running /app/llama-server $CACHED_LLAMA_ARGS $LLAMA_SERVER_CMD_ARGS --port 3098"
touch llama.server.log touch llama.server.log
# We need to pass these arguments to llama-server verbatim. # We need to pass these arguments to llama-server verbatim.
LD_LIBRARY_PATH=/app /app/llama-server $LLAMA_SERVER_CMD_ARGS --port 3098 2>&1 | tee llama.server.log & LD_LIBRARY_PATH=/app /app/llama-server $CACHED_LLAMA_ARGS $LLAMA_SERVER_CMD_ARGS --port 3098 2>&1 | tee llama.server.log &
LLAMA_SERVER_PID=$! # store the process ID (PID) of the background command LLAMA_SERVER_PID=$! # store the process ID (PID) of the background command
tries_so_far=0
check_server_is_running() { check_server_is_running() {
echo "start.sh: Checking if llama-server is done initializing..." echo "start.sh: Checking if llama-server is done initializing..."
@@ -62,13 +75,27 @@ check_server_is_running() {
else else
return 1 # failure return 1 # failure
fi fi
tries_so_far=$((tries_so_far + 1))
if [ $tries_so_far -ge 120 ]; then
echo "start.sh: Error: llama-server did not start within 60 seconds."
exit 1
fi
# check if the process is still running
if ! kill -0 $LLAMA_SERVER_PID 2>/dev/null; then
echo "start.sh: Error: llama-server process has exited unexpectedly."
exit 1
fi
} }
echo "start.sh: Waiting for llama-server to start..." echo "start.sh: Waiting for llama-server to start..."
# wait for the server to start # wait for the server to start
while ! check_server_is_running; do while ! check_server_is_running; do
sleep 5 # we don't want to lose too much time, so we check very frequently
sleep 0.5
done done
echo "start.sh: llama-server is up and running, delegating to the handler script." echo "start.sh: llama-server is up and running, delegating to the handler script."