llm-llama-server
Interact with llama-server models
Installation
Install this plugin in the same environment as LLM.
llm install llm-llama-server
Usage
You'll need to be running a llama-server on port 8080 to use this plugin.
You can brew install llama.cpp to obtain that binary. Then run it like this:
llama-server -hf unsloth/gemma-3-4b-it-GGUF:Q4_K_XL
This loads and serves the unsloth/gemma-3-4b-it-GGUF GGUF version of Gemma 3 4B - a 3.2GB download.
To access regular models from LLM, use the llama-server model:
llm -m llama-server "say hi"
For vision models, use llama-server-vision:
llm -m llama-server-vision describe -a path/to/image.png
For models with tools (which also support vision) use llama-server-tools:
llm -m llama-server-tools -T llm_time 'time?' --td
You'll need to run the llama-server with the --jinja flag in order for this to work:
llama-server --jinja -hf unsloth/gemma-3-4b-it-GGUF:Q4_K_XL
Development
To set up this plugin locally, first checkout the code. Then create a new virtual environment:
cd llm-llama-server
python -m venv venv
source venv/bin/activate
Now install the dependencies and test dependencies:
python -m pip install -e '.[test]'
To run the tests:
python -m pytest
Release files for llm-llama-server 0.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llm_llama_server-0.2.tar.gz | 2.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llm_llama_server-0.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 5.6 kB
Release files / llm_llama_server-0.2.tar.gz
| Download URL | llm_llama_server-0.2.tar.gz |
|---|---|
| Size | 2.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ffb45c48814804f1b88cb136bfb66a39672392138ffdc14615178e8fa37887c9
|
|
BLAKE2b-256 checksum How to use checksums |
3ae886705c9f99737c8e22bf391cfafc9535f5043ab2366a4471bbeed8488887
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.12.9
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on May 28, 2025.
Transparency logRelease files / llm_llama_server-0.2-py3-none-any.whl
| Download URL | llm_llama_server-0.2-py3-none-any.whl |
|---|---|
| Size | 2.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2d129489cd82051608512ae9b1832d36a736d7108b5f4c7ed778df7ada0a8fd3
|
|
BLAKE2b-256 checksum How to use checksums |
8e4e3ecfd864bfabf6b49ca86f4f040b2674093ca777893d83aacbb287d24086
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.12.9
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on May 28, 2025.
Transparency log