Skip to main content

Run AI locally. Beautifully. A drop-in replacement for Ollama.

Project description

Kapri

Run AI locally. Beautifully.

PyPI License Python GitHub release


Why Kapri?

Kapri is a complete drop-in replacement for Ollama. Built on llama.cpp's native router mode, it gives you full control with zero cloud dependency.

Key Features

  • Full GPU Support — Vulkan (AMD), CUDA (NVIDIA), ROCm (AMD Linux), SYCL (Intel), Metal (Apple Silicon), CPU
  • Any GGUF Model — Pull directly from HuggingFace, not limited to Ollama's registry
  • Transparent Config — Plain INI you can read and edit, no hidden internals
  • Multi-Model Router — native llama.cpp router with model auto-switching
  • Universal Installpip install kapri-ai, works everywhere Python exists

Kapri vs Ollama

Feature Kapri Ollama
Vulkan Support Native Blocked
Any HuggingFace Model Yes Registry only
Full llama-server Flags Complete Abstracted
Transparent Config INI Hidden
Multi-model Router Yes Partial
Universal Install pip Platform-specific

Quick Start

Step 1: Install

pip install kapri-ai

Step 2: Setup

kapri install

This downloads the llama-server binary (auto-detects your GPU backend).

Or specify a backend:

kapri install --backend vulkan   # AMD GPUs
kapri install --backend cuda     # NVIDIA GPUs
kapri install --backend cpu      # CPU only

Using an existing llama-server: If you have a working llama-server build (e.g., custom Vulkan build):

kapri install --llama-server /path/to/llama-server.exe

Step 3: Start Server

kapri serve

Step 4: Chat

# Opens web UI (default)
kapri run qwen3.5-0.8b

# Or terminal chat
kapri run qwen3.5-0.8b --tui

The server is also available at http://localhost:11434 for any OpenAI-compatible client.

Then use with any OpenAI-compatible client:

curl http://localhost:11434/v1/models
curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "qwen2.5-coder", "messages": [{"role": "user", "content": "Hello!"}]}'

Pulling Models

Kapri supports multiple ways to pull models:

1. From Built-in Registry

# Default Q4_K_M quantization
kapri pull llama3.2-3b

# Specify quantization
kapri pull llama3.2-3b:Q5_K_M

# List available models
kapri search llama

2. Direct from HuggingFace

Pull any GGUF model directly from HuggingFace:

# Unsloth models
kapri pull unsloth/Qwen3.5-0.8B-GGUF
kapri pull unsloth/Qwen3.5-0.8B-GGUF:Q4_K_M

# Bartowski quantized models
kapri pull bartowski/Llama-3.2-3B-Instruct-GGUF
kapri pull bartowski/Qwen2.5-Coder-7B-Instruct-GGUF:Q4_K_M

# ggml-org TinyLlama
kapri pull ggml-org/TinyLlama-1.1B-Chat-v1.0-GGUF

# Any other GGUF repo
kapri pull <hf-repo>/<model-name>-GGUF
kapri pull <hf-repo>/<model-name>-GGUF:Q4_K_M

Examples of GGUF repos on HuggingFace:


Supported Backends

Backend Description GPU Brands
CUDA NVIDIA GPUs NVIDIA
Vulkan AMD GPUs (full speed) AMD
ROCm AMD GPUs (Linux) AMD
SYCL Intel GPUs Intel
Metal Apple Silicon Apple
CPU Fallback Any

Forcing a Backend

kapri install --backend cuda    # Force CUDA
kapri install --backend vulkan  # Force Vulkan
kapri install --backend cpu    # Force CPU

CLI Reference

Command Description
kapri install Install binaries (interactive wizard)
kapri install --backend vulkan Install specific backend
kapri backend vulkan Switch GPU backend
kapri update Update binaries
kapri update --all Update kapri and binaries
kapri pull <model> Download a model
kapri serve Start the server
kapri stop Stop the server
kapri status Show server status
kapri list List downloaded models
kapri search <query> Search registry
kapri remove <model> Remove a model
kapri run <model> Open web UI chat (default)
kapri run <model> --tui Open terminal chat
kapri config Show full config INI
kapri config show-all Show full config INI
kapri config show <model> Show specific model config
kapri config edit <model> Edit model config in editor
kapri config search <query> Search models in config
kapri config show --path Show config file path
kapri config show --reset Regenerate config
kapri logs View logs

Backend Options

kapri backend auto      # Auto-detect (recommended)
kapri backend vulkan   # AMD GPUs
kapri backend cuda     # NVIDIA GPUs
kapri backend rocm     # AMD Linux
kapri backend sycl     # Intel GPUs
kapri backend metal   # Apple Silicon
kapri backend cpu     # CPU only

API Endpoints

The server exposes OpenAI-compatible endpoints:

  • GET /v1/models — List models
  • POST /v1/chat/completions — Chat completions
  • POST /v1/completions — Text completions

Python Example

import httpx

client = httpx.Client(base_url="http://localhost:11434")
response = client.post("/v1/chat/completions", json={
    "model": "qwen2.5-coder",
    "messages": [{"role": "user", "content": "Write a Python hello world"}]
})
print(response.json())

OpenAI Python SDK

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="not-needed"
)

chat = client.chat.completions.create(
    model="qwen2.5-coder",
    messages=[{"role": "user", "content": "Hello!"}]
)
print(chat.choices[0].message.content)

Architecture

~/.kapri/                    # Base directory
├── bin/                    # llama-server binary
├── models/                # Downloaded GGUF files
├── models.ini             # llama.cpp router preset
├── server.pid            # Server process ID
└── server.log            # Server logs

Troubleshooting

No GPU detected

kapri install --backend cuda    # Force CUDA
kapri install --backend vulkan  # Force Vulkan (AMD)

Port in use

kapri serve --port 11435  # Different port

Model not found

Make sure the GGUF file exists on HuggingFace. Try:

# Use full repo path
kapri pull unsloth/Qwen3.5-0.8B-GGUF

Requirements

  • Python 3.10+
  • No root/sudo required
  • GPU optional (CPU fallback works)

Links


Built on llama.cpp

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

kapri_ai-0.3.0.tar.gz (26.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

kapri_ai-0.3.0-py3-none-any.whl (26.6 kB view details)

Uploaded Python 3

File details

Details for the file kapri_ai-0.3.0.tar.gz.

File metadata

  • Download URL: kapri_ai-0.3.0.tar.gz
  • Upload date:
  • Size: 26.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.9

File hashes

Hashes for kapri_ai-0.3.0.tar.gz
Algorithm Hash digest
SHA256 92aa55c006db50461bb42f085ad1680f2d245403127cdd7c16d6df749b6ab7c7
MD5 0d5166fca9f942eb374278e35516cbf1
BLAKE2b-256 aed099857dc64cfb45d9262195f73aa3ae3e0085fa4663b602c67eecf949163b

See more details on using hashes here.

File details

Details for the file kapri_ai-0.3.0-py3-none-any.whl.

File metadata

  • Download URL: kapri_ai-0.3.0-py3-none-any.whl
  • Upload date:
  • Size: 26.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.9

File hashes

Hashes for kapri_ai-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 c42de722aefe21e7c4e682f374898f593317e6a0229061b7ce212ff051c3033b
MD5 049a358598ab97cb4092abf62e90aafa
BLAKE2b-256 87f71d98bc5e6ebdbfc747004344f6624c3a1e3d67a9ee11e6ccfba7bb9af25c

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page