Skip to main content

llamacpp-cli

Ollama-like CLI wrapper around llama.cpp. Provides a simple command-line interface that mirrors Ollama's subcommands but powered by llama.cpp as the backend inference engine.

Features

  • pull - Download GGUF models from Hugging Face
  • run - Run models interactively using llama.cpp
  • serve - Start the llama.cpp server
  • lb-proxy - Multi-backend load balancer proxy (NEW!)
  • list - List downloaded models
  • ps - Show running llama.cpp processes
  • rm - Remove a downloaded model
  • search - Search Hugging Face for GGUF models
  • install - Install/update llama.cpp binaries

Installation

From PyPI

pip install llamacpp-cli

From Source

pip install -e .

Quick Start

1. Install llama.cpp binaries

llamacpp install

This downloads the latest llama.cpp release to ~/.llamacpp/bin/.

2. Pull a model

llamacpp pull unsloth/gemma-3-270m-it-GGUF:Q4_K_M

Or use a short alias:

llamacpp pull gemma3:270m

3. Run interactively

llamacpp run gemma3:270m

4. Start the server

llamacpp serve -m gemma3:270m

The server runs at http://0.0.0.0:8080 with OpenAI-compatible API.

CPU-Optimized Presets

For CPU-only servers, use presets optimized for different workloads:

# Code tasks (default): 16K context, 2-4 parallel requests
llamacpp serve --preset code

# Chat/conversational: 8K context, 4-6 parallel requests
llamacpp serve --preset chat

# Fast queries: 4K context, 6-8 parallel requests
llamacpp serve --preset fast

# Large codebases: 32K context, 1 parallel request (slower)
llamacpp serve --preset max-context

See CPU_OPTIMIZATION.md for detailed tuning guide.

Commands

llamacpp pull <model>      Download GGUF model from Hugging Face
llamacpp run <model>       Run a model interactively
llamacpp serve             Start the llama.cpp server
llamacpp lb-proxy          Start multi-backend load balancer (see LB_PROXY.md)
llamacpp list              List downloaded models
llamacpp ps                Show running processes
llamacpp rm <model>        Remove a model
llamacpp search <query>    Search for models on Hugging Face
llamacpp install           Install/update llama.cpp binaries

Load Balancer Proxy

For distributing requests across multiple machines, use the load balancer:

# Auto-discover backends on your network
llamacpp lb-proxy --discover-subnet 192.168.1.0/24

# Or specify backends manually
llamacpp lb-proxy -b http://machine1:8000 -b http://machine2:8000

See LB_PROXY.md for detailed documentation on:

  • Model-aware routing
  • Least-connections load balancing
  • Auto-discovery and health checks
  • Configuration options

Model Names

Model names can be specified in multiple ways:

  • Full Hugging Face path: unsloth/gemma-3-270m-it-GGUF:Q4_K_M
  • Short format: namespace/model:quantization (e.g., gemma3:270m)
  • Short name: gemma3:270m, qwen3, llama3:8b

Alias support is planned for future releases.

Configuration

  • Models are stored in ~/.llamacpp/models/
  • Binaries are installed to ~/.llamacpp/bin/
  • Database (SQLite) is at ~/.llamacpp/llamacpp.db

Environment Variables

Variable Description Default
LLAMACPP_BIN_DIR Directory for llama.cpp binaries ~/.llamacpp/bin
LLAMACPP_MODEL_DIR Directory for models ~/.llamacpp/models

Usage with LLM CLI

This package also registers as an LLM plugin for the llm CLI:

# Install the plugin (requires llm and llama-cpp-python)
pip install llm-llama-cpp llama-cpp-python

# Register a model
llm llama-cpp add-model ~/.llamacpp/models/gemma-3-270m-it-Q4_K_M.gguf --alias gemma3:270m

# Use with llm
llm -m gemma3:270m "Your prompt here"

Development

# Install in editable mode with dev dependencies
pip install -e ".[dev]"

# Run tests
pytest

# Run a single test file
pytest tests/test_foo.py

# Lint
ruff check .

# Format
ruff format .

Publishing to PyPI

Prerequisites

  1. Create a PyPI account at https://pypi.org/
  2. Install build tools:
pip install build twine

Build and Publish

  1. Update version in pyproject.toml:
[project]
version = "0.1.0"
  1. Build the package:
python -m build

This creates distributable archives in dist/.

  1. Upload to PyPI:
twine upload dist/*

You'll be prompted for your PyPI username and password.

For Test PyPI (testing first):

twine upload --repository testpypi dist/*

Using uv (Alternative)

# Install uv if not already
pip install uv

# Build
uv build

# Publish to PyPI
uv publish

# Or Test PyPI
uv publish --test

License

MIT

Metadata

Release files for llamacpp-cli 0.1.9

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llamacpp-cli 0.1.9
File Size Uploaded
llamacpp_cli-0.1.9.tar.gz 127.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llamacpp-cli 0.1.9
File Interpreter ABI Platform
llamacpp_cli-0.1.9-py3-none-any.whl Python 3 none any Details

Total release size: 216.4 kB

Release files / llamacpp_cli-0.1.9.tar.gz

Download URL llamacpp_cli-0.1.9.tar.gz
Size 127.2 kB
Tags Source
SHA-256 checksum
How to use checksums
8564d73c881c7a68eb883e3bab921a3713490def8b35a8d287debce60cca585a
BLAKE2b-256 checksum
How to use checksums
7dd0195070070be1fbdc1cc74b7e42a6643f0c2d72a85f75ba92aa8bc402df27
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jun 4, 2026.

Transparency log

Release files / llamacpp_cli-0.1.9-py3-none-any.whl

Download URL llamacpp_cli-0.1.9-py3-none-any.whl
Size 89.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2dc97cd0e78b1f695492ae134ade9f103f52a43c4eb1e96d7e02503e36f9ce5a
BLAKE2b-256 checksum
How to use checksums
72135db504eeb567a7707f585d443b02b1b13869ab0979734090a2a53c1b58db
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jun 4, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.9 This release

2 release files

0.1.8

2 release files

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page