Skip to main content

Run AI locally. Beautifully. A drop-in replacement for Ollama.

Project description

Kapri

Run AI locally. Beautifully.

PyPI License Python GitHub release


Why Kapri?

Kapri is a complete drop-in replacement for Ollama. Built on llama.cpp + llama-swap, it gives you full control with zero cloud dependency.

Key Features

  • Full GPU Support — Vulkan (AMD), CUDA (NVIDIA), ROCm (AMD Linux), SYCL (Intel), Metal (Apple Silicon), CPU
  • Any GGUF Model — Pull directly from HuggingFace, not limited to Ollama's registry
  • Transparent Config — Plain YAML you can read and edit, no hidden internals
  • Multi-Model Hot-Swap — llama-swap with TTL-based auto-unload
  • Universal Installpip install kapri-ai, works everywhere Python exists

Kapri vs Ollama

Feature Kapri Ollama
Vulkan Support Native Blocked
Any HuggingFace Model Yes Registry only
Full llama-server Flags Complete Abstracted
Transparent Config YAML Hidden
Multi-model Hot-swap Yes Partial
Universal Install pip Platform-specific

Quick Start

Step 1: Install

pip install kapri-ai

Step 2: Setup

kapri install

This downloads llama-server and llama-swap binaries (auto-detects your GPU backend).

Or specify a backend:

kapri install --backend vulkan   # AMD GPUs
kapri install --backend cuda     # NVIDIA GPUs
kapri install --backend cpu      # CPU only

Using an existing llama-server: If you have a working llama-server build (e.g., custom Vulkan build):

kapri install --llama-server /path/to/llama-server.exe

Step 3: Start Server

kapri serve

Step 4: Chat

# Opens web UI (default)
kapri run qwen3.5-0.8b

# Or terminal chat
kapri run qwen3.5-0.8b --tui

The server is also available at http://localhost:11434 for any OpenAI-compatible client.

Then use with any OpenAI-compatible client:

curl http://localhost:11434/v1/models
curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "qwen2.5-coder", "messages": [{"role": "user", "content": "Hello!"}]}'

Pulling Models

Kapri supports multiple ways to pull models:

1. From Built-in Registry

# Default Q4_K_M quantization
kapri pull llama3.2-3b

# Specify quantization
kapri pull llama3.2-3b:Q5_K_M

# List available models
kapri search llama

2. Direct from HuggingFace

Pull any GGUF model directly from HuggingFace:

# Unsloth models
kapri pull unsloth/Qwen3.5-0.8B-GGUF
kapri pull unsloth/Qwen3.5-0.8B-GGUF:Q4_K_M

# Bartowski quantized models
kapri pull bartowski/Llama-3.2-3B-Instruct-GGUF
kapri pull bartowski/Qwen2.5-Coder-7B-Instruct-GGUF:Q4_K_M

# ggml-org TinyLlama
kapri pull ggml-org/TinyLlama-1.1B-Chat-v1.0-GGUF

# Any other GGUF repo
kapri pull <hf-repo>/<model-name>-GGUF
kapri pull <hf-repo>/<model-name>-GGUF:Q4_K_M

Examples of GGUF repos on HuggingFace:


Supported Backends

Backend Description GPU Brands
CUDA NVIDIA GPUs NVIDIA
Vulkan AMD GPUs (full speed) AMD
ROCm AMD GPUs (Linux) AMD
SYCL Intel GPUs Intel
Metal Apple Silicon Apple
CPU Fallback Any

Forcing a Backend

kapri install --backend cuda    # Force CUDA
kapri install --backend vulkan  # Force Vulkan
kapri install --backend cpu    # Force CPU

CLI Reference

Command Description
kapri install Install binaries (interactive wizard)
kapri install --backend vulkan Install specific backend
kapri backend vulkan Switch GPU backend
kapri update Update binaries
kapri update --all Update kapri and binaries
kapri pull <model> Download a model
kapri serve Start the server
kapri stop Stop the server
kapri status Show server status
kapri list List downloaded models
kapri search <query> Search registry
kapri remove <model> Remove a model
kapri run <model> Open web UI chat (default)
kapri run <model> --tui Open terminal chat
kapri config show Show config
kapri logs View logs

Backend Options

kapri backend auto      # Auto-detect (recommended)
kapri backend vulkan   # AMD GPUs
kapri backend cuda     # NVIDIA GPUs
kapri backend rocm     # AMD Linux
kapri backend sycl     # Intel GPUs
kapri backend metal   # Apple Silicon
kapri backend cpu     # CPU only

API Endpoints

The server exposes OpenAI-compatible endpoints:

  • GET /v1/models — List models
  • POST /v1/chat/completions — Chat completions
  • POST /v1/completions — Text completions

Python Example

import httpx

client = httpx.Client(base_url="http://localhost:11434")
response = client.post("/v1/chat/completions", json={
    "model": "qwen2.5-coder",
    "messages": [{"role": "user", "content": "Write a Python hello world"}]
})
print(response.json())

OpenAI Python SDK

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="not-needed"
)

chat = client.chat.completions.create(
    model="qwen2.5-coder",
    messages=[{"role": "user", "content": "Hello!"}]
)
print(chat.choices[0].message.content)

Architecture

~/.kapri/                    # Base directory
├── bin/                    # llama-server, llama-swap binaries
├── models/                # Downloaded GGUF files
├── config.yaml            # llama-swap configuration
├── server.pid            # Server process ID
└── server.log            # Server logs

Troubleshooting

No GPU detected

kapri install --backend cuda    # Force CUDA
kapri install --backend vulkan  # Force Vulkan (AMD)

Port in use

kapri serve --port 11435  # Different port

Model not found

Make sure the GGUF file exists on HuggingFace. Try:

# Use full repo path
kapri pull unsloth/Qwen3.5-0.8B-GGUF

Requirements

  • Python 3.10+
  • No root/sudo required
  • GPU optional (CPU fallback works)

Links


Built on llama.cpp + llama-swap

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

kapri_ai-0.1.1.tar.gz (31.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

kapri_ai-0.1.1-py3-none-any.whl (31.7 kB view details)

Uploaded Python 3

File details

Details for the file kapri_ai-0.1.1.tar.gz.

File metadata

  • Download URL: kapri_ai-0.1.1.tar.gz
  • Upload date:
  • Size: 31.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.9

File hashes

Hashes for kapri_ai-0.1.1.tar.gz
Algorithm Hash digest
SHA256 6fb7d0e273c5eff29cf0742f927a450c6ddc3bb15f66c51af6512253b873e1f0
MD5 e56bf1fb9d4681106c9ef0b881c3b697
BLAKE2b-256 3f1940c3fbd5f811e5082bc6e0273907c193d685989846db7623f381d7d5d3be

See more details on using hashes here.

File details

Details for the file kapri_ai-0.1.1-py3-none-any.whl.

File metadata

  • Download URL: kapri_ai-0.1.1-py3-none-any.whl
  • Upload date:
  • Size: 31.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.9

File hashes

Hashes for kapri_ai-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 248647299932e53ac2c505fe785eebb2179fbe81cb5950a7d3d051f1519040b9
MD5 68b00b90eb69315591fd4b3bcf548511
BLAKE2b-256 eda96cc49a52e98a56c8fe06a67bf9f71f47a99e4239a7eee54a5ce571df656e

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page