Magnus and GitHub
85c10953d6
add pseudo API key so the client doesn't raise an error
...
this seems to have been introduced in a recent OpenAI SDK update
2026-06-14 18:00:42 +02:00
Magnus and GitHub
ea82e2a901
remove "inference" from title to make it more consistent with the other runpod repos on the hub
2026-06-03 18:22:06 +02:00
Magnus and GitHub
f890c50c8a
Merge pull request #3 from cardinalfan1/claude/fix-runpod-startup-xxJqT
...
Fix -m None passed to llama-server when cached model not found
v1.2.2
2026-06-03 18:12:26 +02:00
Claude
dd2d52b5e3
Fix -m None passed to llama-server when cached model not found
...
When LLAMA_CACHED_MODEL is set but the model isn't present in the cache,
find_cached.py was printing Python's None as the string "None", causing
start.sh to pass "-m None" to llama-server.
- find_cached.py: print an error to stderr and exit 1 when the model path
cannot be resolved, instead of printing "None"
- start.sh: capture find_cached.py output into a local variable, check
the exit code and guard against empty output before constructing
CACHED_LLAMA_ARGS; also quote the env-var expansions to handle spaces
https://claude.ai/code/session_011ny5CFYnrzPbRneSzio5CR
2026-04-15 19:34:58 +00:00
mags0ft and GitHub
1196621d77
Merge pull request #1 from aparmar2000/patch-1
...
automatically convert path to lowercase
2026-03-10 19:11:04 +01:00
Ashan Parmar and GitHub
3bcbee2d93
Update find_cached.py
...
The cache name is lowercase
2026-02-20 18:04:59 -07:00
mags0ft
b24c32c024
re-activate automatic testing
v1.2.1
2026-02-12 08:45:54 +01:00
mags0ft and GitHub
c7b115bec0
fix typo in cached models docs
2026-02-11 02:14:18 +01:00
mags0ft
8a2982faa4
adjust incorrect llama-server start command log
2025-12-27 23:51:00 +01:00
mags0ft
d4ac091735
add timeout for llama-server startup to prevent indefinite waiting
2025-12-27 18:00:58 +01:00
mags0ft
3d442d9123
add error handling for unexpected llama-server process exit
v1.2.0
2025-12-26 14:51:06 +01:00
mags0ft
0c814177fe
add caching support for models in start.sh and update documentation
2025-12-26 14:49:03 +01:00
mags0ft
0a59d57025
rename tests.json temporarily to skip apparently buggy RunPod CI
...
The RunPod support told me to do this until the issues are resolved.
v1.1.2
2025-12-26 11:37:23 +01:00
mags0ft
398b59c0a1
update test timeout and default LLAMA_SERVER_CMD_ARGS for improved performance
v1.1.1
2025-12-25 19:17:31 +01:00
mags0ft
889d3c5c96
move find_cached.py to src
v1.1.0
2025-12-18 10:57:10 +01:00
mags0ft
41b3b3d02b
add helper script to use cached models
2025-12-18 10:50:33 +01:00
mags0ft and GitHub
a13d6c1b9f
add issue note and fork explanation to README
2025-11-24 22:24:46 +01:00
mags0ft
403d318ffc
update default args to include -ngl 99, improve startup sleep duration
v1.0.1
2025-11-19 19:16:18 +01:00
mags0ft
381ed9ffff
prepare for production release
v1.0.0
2025-11-19 18:40:00 +01:00
mags0ft
c42b80ebbd
add default behavior for undefined LLAMA_SERVER_CMD_ARGS
v0.0.6
2025-11-19 18:11:36 +01:00
mags0ft
40d4097799
fix several major bugs in startup script
v0.0.5
2025-11-19 17:47:06 +01:00
mags0ft
36a8b32d1e
fix: ensure script fails on error by setting 'set -e'
v0.0.4
2025-11-15 16:07:54 +01:00
mags0ft
5f6c099504
further fixes for start.sh script
v0.0.3
2025-11-15 15:15:17 +01:00
mags0ft
be3de61c52
fix: correct logic for port validation in start.sh
2025-11-15 15:02:05 +01:00
mags0ft
558da755c3
add empty handler.py file to mark repository as Runpod-compatible endpoint
v0.0.2
2025-11-15 14:11:32 +01:00
mags0ft
9388fae14b
add badge to README
v0.0.1-alpha
2025-11-15 14:07:00 +01:00
mags0ft
3a8d2decfd
change project to work with llama.cpp instead of Ollama
2025-11-15 14:03:10 +01:00
github-actions[bot]
d58239e290
chore: bump OLLAMA_VERSION to 0.12.11
2025-11-14 08:31:45 +00:00
github-actions[bot]
672b233f8d
chore: bump OLLAMA_VERSION to 0.12.10
2025-11-07 08:32:06 +00:00
github-actions[bot]
20c2836294
chore: bump OLLAMA_VERSION to 0.12.9
2025-11-02 08:26:29 +00:00
github-actions[bot]
b33647a947
chore: bump OLLAMA_VERSION to 0.12.8
2025-11-01 08:27:14 +00:00
github-actions[bot]
4d1c590077
chore: bump OLLAMA_VERSION to 0.12.7
2025-10-30 08:30:47 +00:00
github-actions[bot]
21bc109307
chore: bump OLLAMA_VERSION to 0.12.6
2025-10-17 08:31:01 +00:00
github-actions[bot]
aa0b219f17
chore: bump OLLAMA_VERSION to 0.12.5
2025-10-11 08:26:10 +00:00
SvenBrnn
3ef6d0bb62
fix pulling model name
2025-10-06 09:13:08 +02:00
SvenBrnn
4bf93d6680
fix pulling model name
2025-10-06 09:12:39 +02:00
SvenBrnn
3920146e31
change MODEL_NAME to OLLAMA_MODEL_NAME to prevent runpod from blocking deploy
2025-10-06 08:10:12 +02:00
github-actions[bot]
074e069cdd
chore: bump OLLAMA_VERSION to 0.12.3
2025-09-27 08:25:20 +00:00
github-actions[bot]
fddafc3d08
chore: bump OLLAMA_VERSION to 0.12.2
2025-09-25 08:29:56 +00:00
github-actions[bot]
2e49d8370c
chore: bump OLLAMA_VERSION to 0.12.1
2025-09-24 08:30:14 +00:00
github-actions[bot]
93410b19be
chore: bump OLLAMA_VERSION to 0.12.0
2025-09-20 08:26:22 +00:00
github-actions[bot]
a89f6169ab
chore: bump OLLAMA_VERSION to 0.11.11
2025-09-16 08:30:28 +00:00
SvenBrnn and GitHub
10d0941e6f
Merge pull request #3 from seurimas/master
...
Borrow vLLM concurrenccy logic a little.
2025-09-14 09:48:58 +02:00
Nicholas
51d38adcf3
Borrow vLLM concurrenccy logic a little.
2025-09-11 00:31:09 -05:00
github-actions[bot]
0df1db40e3
chore: bump OLLAMA_VERSION to 0.11.10
2025-09-05 08:29:16 +00:00
github-actions[bot]
861093b09a
chore: bump OLLAMA_VERSION to 0.11.9
2025-09-04 08:28:50 +00:00
github-actions[bot]
034257b65f
chore: bump OLLAMA_VERSION to 0.11.8
2025-08-29 08:29:38 +00:00
github-actions[bot]
26a950daa6
chore: bump OLLAMA_VERSION to 0.11.7
2025-08-26 08:32:11 +00:00
github-actions[bot]
4cf7821528
chore: bump OLLAMA_VERSION to 0.11.6
2025-08-21 08:30:22 +00:00
github-actions[bot]
38341b51f6
chore: bump OLLAMA_VERSION to 0.11.5
2025-08-20 08:31:03 +00:00