Here[1] is the offending pull request. The llama.cpp devs are trying to refactor the code, so the "engine" bits are shared between the current server and client applications.
Georgi himself noted it should be possible to replace the transport layer from HTTP to say an internal message queue, and the main developer of the PR said the HTTP bit is already abstracted and intended to include a non-HTTP transport example, so seems clear to me it's not intended to be HTTP-only.
3 comments of 5
[ 0.30 ms ] story [ 6.1 ms ] threadThe CLI now asking for access to load the model?
Here the change that screw everything:
b9927 @github-actions github-actions released this 17 hours ago b9927 c264f65 Details cli : move to HTTP-based implementation (#24948)
cli: move to HTTP-based implementation
wip
working
remote server ok
cli support router mode
Co-authored-by: Piotr Wilkin ilintar@gmail.com
case: router with only one model
Apply suggestions from code review
Co-authored-by: Piotr Wilkin (ilintar) piotr.wilkin@syndatis.com
remove outdated comment
use destructor instead
add ftype
cli-view --> cli-ui
pimpl
no more json in header
nits fixes
also show model aliases
Adios llama.cpp.
Georgi himself noted it should be possible to replace the transport layer from HTTP to say an internal message queue, and the main developer of the PR said the HTTP bit is already abstracted and intended to include a non-HTTP transport example, so seems clear to me it's not intended to be HTTP-only.
[1]: https://github.com/ggml-org/llama.cpp/pull/24948