llamacpp-cli
Ollama-like CLI wrapper around llama.cpp. Provides a simple command-line interface that mirrors Ollama's subcommands but powered by llama.cpp as the backend inference engine.
Features
- pull - Download GGUF models from Hugging Face
- run - Run models interactively using llama.cpp
- serve - Start the llama.cpp server
- lb-proxy - Multi-backend load balancer proxy (NEW!)
- list - List downloaded models
- ps - Show running llama.cpp processes
- rm - Remove a downloaded model
- search - Search Hugging Face for GGUF models
- install - Install/update llama.cpp binaries
Installation
From PyPI
pip install llamacpp-cli
From Source
pip install -e .
Quick Start
1. Install llama.cpp binaries
llamacpp install
This downloads the latest llama.cpp release to ~/.llamacpp/bin/.
2. Pull a model
llamacpp pull unsloth/gemma-3-270m-it-GGUF:Q4_K_M
Or use a short alias:
llamacpp pull gemma3:270m
3. Run interactively
llamacpp run gemma3:270m
4. Start the server
llamacpp serve -m gemma3:270m
The server runs at http://0.0.0.0:8080 with OpenAI-compatible API.
CPU-Optimized Presets
For CPU-only servers, use presets optimized for different workloads:
# Code tasks (default): 16K context, 2-4 parallel requests
llamacpp serve --preset code
# Chat/conversational: 8K context, 4-6 parallel requests
llamacpp serve --preset chat
# Fast queries: 4K context, 6-8 parallel requests
llamacpp serve --preset fast
# Large codebases: 32K context, 1 parallel request (slower)
llamacpp serve --preset max-context
See CPU_OPTIMIZATION.md for detailed tuning guide.
Commands
llamacpp pull <model> Download GGUF model from Hugging Face
llamacpp run <model> Run a model interactively
llamacpp serve Start the llama.cpp server
llamacpp lb-proxy Start multi-backend load balancer (see LB_PROXY.md)
llamacpp list List downloaded models
llamacpp ps Show running processes
llamacpp rm <model> Remove a model
llamacpp search <query> Search for models on Hugging Face
llamacpp install Install/update llama.cpp binaries
Load Balancer Proxy
For distributing requests across multiple machines, use the load balancer:
# Auto-discover backends on your network
llamacpp lb-proxy --discover-subnet 192.168.1.0/24
# Or specify backends manually
llamacpp lb-proxy -b http://machine1:8000 -b http://machine2:8000
See LB_PROXY.md for detailed documentation on:
- Model-aware routing
- Least-connections load balancing
- Auto-discovery and health checks
- Configuration options
Model Names
Model names can be specified in multiple ways:
- Full Hugging Face path:
unsloth/gemma-3-270m-it-GGUF:Q4_K_M - Short format:
namespace/model:quantization(e.g.,gemma3:270m) - Short name:
gemma3:270m,qwen3,llama3:8b
Alias support is planned for future releases.
Configuration
- Models are stored in
~/.llamacpp/models/ - Binaries are installed to
~/.llamacpp/bin/ - Database (SQLite) is at
~/.llamacpp/llamacpp.db
Environment Variables
| Variable | Description | Default |
|---|---|---|
LLAMACPP_BIN_DIR |
Directory for llama.cpp binaries | ~/.llamacpp/bin |
LLAMACPP_MODEL_DIR |
Directory for models | ~/.llamacpp/models |
Usage with LLM CLI
This package also registers as an LLM plugin for the llm CLI:
# Install the plugin (requires llm and llama-cpp-python)
pip install llm-llama-cpp llama-cpp-python
# Register a model
llm llama-cpp add-model ~/.llamacpp/models/gemma-3-270m-it-Q4_K_M.gguf --alias gemma3:270m
# Use with llm
llm -m gemma3:270m "Your prompt here"
Development
# Install in editable mode with dev dependencies
pip install -e ".[dev]"
# Run tests
pytest
# Run a single test file
pytest tests/test_foo.py
# Lint
ruff check .
# Format
ruff format .
Publishing to PyPI
Prerequisites
- Create a PyPI account at https://pypi.org/
- Install build tools:
pip install build twine
Build and Publish
- Update version in
pyproject.toml:
[project]
version = "0.1.0"
- Build the package:
python -m build
This creates distributable archives in dist/.
- Upload to PyPI:
twine upload dist/*
You'll be prompted for your PyPI username and password.
For Test PyPI (testing first):
twine upload --repository testpypi dist/*
Using uv (Alternative)
# Install uv if not already
pip install uv
# Build
uv build
# Publish to PyPI
uv publish
# Or Test PyPI
uv publish --test
License
MIT
Metadata
Release files for llamacpp-cli 0.1.9
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llamacpp_cli-0.1.9.tar.gz | 127.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llamacpp_cli-0.1.9-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 216.4 kB
Release files / llamacpp_cli-0.1.9.tar.gz
| Download URL | llamacpp_cli-0.1.9.tar.gz |
|---|---|
| Size | 127.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8564d73c881c7a68eb883e3bab921a3713490def8b35a8d287debce60cca585a
|
|
BLAKE2b-256 checksum How to use checksums |
7dd0195070070be1fbdc1cc74b7e42a6643f0c2d72a85f75ba92aa8bc402df27
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jun 4, 2026.
Transparency logRelease files / llamacpp_cli-0.1.9-py3-none-any.whl
| Download URL | llamacpp_cli-0.1.9-py3-none-any.whl |
|---|---|
| Size | 89.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2dc97cd0e78b1f695492ae134ade9f103f52a43c4eb1e96d7e02503e36f9ce5a
|
|
BLAKE2b-256 checksum How to use checksums |
72135db504eeb567a7707f585d443b02b1b13869ab0979734090a2a53c1b58db
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jun 4, 2026.
Transparency log