Claude dd2d52b5e3 Fix -m None passed to llama-server when cached model not found
When LLAMA_CACHED_MODEL is set but the model isn't present in the cache,
find_cached.py was printing Python's None as the string "None", causing
start.sh to pass "-m None" to llama-server.

- find_cached.py: print an error to stderr and exit 1 when the model path
  cannot be resolved, instead of printing "None"
- start.sh: capture find_cached.py output into a local variable, check
  the exit code and guard against empty output before constructing
  CACHED_LLAMA_ARGS; also quote the env-var expansions to handle spaces

https://claude.ai/code/session_011ny5CFYnrzPbRneSzio5CR
2026-04-15 19:34:58 +00:00
2026-02-12 08:45:54 +01:00
2026-02-11 02:14:18 +01:00

llama.cpp logo

Serverless llama.cpp inference worker for RunPod

This repository contains a serverless inference worker for running llama.cpp models on RunPod. It uses the llama-server image to provide an API for interacting with the models. The following OpenAI API endpoints are supported:

  • v1/models
  • v1/chat/completions
  • v1/completions

Streaming responses is also supported.

Important! This project is still relatively new. Please open a new issue if you encounter any problems in order to get help.

This is a fork of SvenBrnn's runpod-worker-ollama.

Setup

To get the best performance out of this worker, it is recommended to use cached models. Please see the cached models documentation for more information, this is highly recommended and will save many resources.

Configuration

The worker can be configured via environment variables set in the RunPod hub configuration:

  • LLAMA_SERVER_CMD_ARGS: Command line arguments (argv) for the llama-server binary. Example: -hf /path/to/model.gguf:Q4_K_M --ctx-size 4096. IMPORTANT: Please do not define the port argument here, as the worker will always use port 3098 automatically.
  • MAX_CONCURRENCY: Maximum number of concurrent requests the worker can handle. Default is 8.

License

Please see the LICENSE file for more information.

Runpod badge

S
Description
A serverless worker to run LLMs in the cloud - using llama.cpp!
Readme
108 KiB
Languages
Python 76.9%
Shell 17%
Dockerfile 6.1%