Skip to main content

gk-server

Local OpenAI-compatible LLM server GUI for GGUF models, packaged for Python — the server half of the unified ggk package. The GUI runs in your browser against a local backend; inference is served by the gguf-server C/C++ engine, compiled during pip install and bundled with the package as a single binary. The engine evaluates its graphs on gk, an independent tensor library — there is no llama.cpp checkout and no ggml anywhere in the tree. Model and template files are referenced by filesystem path through a built-in file browser — nothing is uploaded or copied to temp storage.

Install

pip install gk-server

Building the bundled engine requires a C/C++ toolchain and CMake ≥ 3.15 (on Windows: MSVC Build Tools). The engine source is resolved from the vendored vendor/server copy (see scripts/vendor_engine.py) or GK_SERVER_ENGINE_DIR. That tree is self-contained — it carries the gk kernels, the GGUF runtime, the common layer and the HTTP server — so the build compiles the server binary and nothing else.

The vendor/server/gk kernels are shared verbatim with the gk-diffuser package; both are vendored from the same unified ggk engine tree, so the two packages always compute with the same gk.

GPU and accelerator backends

The default build is CPU-only (on macOS, Metal is on by default — no flag needed). Backends are opt-in and can be requested with an environment variable or a CMake define — the env var is usually easier to pass through pip:

GK_SERVER_CUDA=1   pip install gk-server    # NVIDIA (needs the CUDA toolkit)
GK_SERVER_HIP=1    pip install gk-server    # AMD (needs ROCm/HIP)
GK_SERVER_VULKAN=1 pip install gk-server    # cross-vendor (needs the Vulkan SDK)

CMAKE_ARGS="-DGK_SERVER_CUDA=ON" pip install gk-server   # equivalent

Available: CUDA, HIP, VULKAN, METAL. Each option maps to the gk backend of the same name, and the finer-grained GK_* knobs (GK_NATIVE, GK_CUDA_ARCHITECTURES, …) can still be passed straight through as -DGK_<NAME>=….

A CUDA build works its own architecture list out from nvcc and the installed GPUs, and embeds PTX for the newest, so an unlisted card JITs rather than failing. A wheel built on one machine for another should still say what it targets, e.g. CMAKE_ARGS="-DGK_SERVER_CUDA=ON -DGK_CUDA_ARCHITECTURES=89".

HTTPS / OpenSSL

Off by default. OpenSSL is only needed to download models over HTTPS (-hf / URL arguments); this package always hands the engine local file paths. Turn it on with GK_SERVER_OPENSSL=1 if you script the engine directly and want HTTPS downloads.

Run

gk-server              # GUI on http://127.0.0.1:8642, opens the browser
python -m gk_server    # same thing
gk-server --port 0     # pick a free port; --no-browser to stay headless

The LLM server the GUI manages defaults to port 8888 and exposes the usual OpenAI-compatible endpoints (/v1/chat/completions, /v1/completions, /v1/embeddings, /v1/rerank, /v1/messages, /props, /health, …).

The engine is directly scriptable from the CLI too:

gk-server engine -- --model model.gguf --port 8888
gk-server engine -- --help

Metadata

Release files for gk-server 0.0.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for gk-server 0.0.2
File Size Uploaded
gk_server-0.0.2.tar.gz 2.8 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for gk-server 0.0.2
File Interpreter ABI Platform
gk_server-0.0.2-py3-none-win_amd64.whl Python 3 none Windows x86-64 Details

Total release size: 18.7 MB

Release files / gk_server-0.0.2.tar.gz

Download URL gk_server-0.0.2.tar.gz
Size 2.8 MB
Tags Source
SHA-256 checksum
How to use checksums
184b9c606a9cf5d6e72b5e4f2efd5fb284dd123a40eb9dec21193cae7e411ac4
BLAKE2b-256 checksum
How to use checksums
2cd7bb1323450441432a5cf7a7f8c30bd7ea86be4885604b8eb96ec356e6b5f0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.9

Release files / gk_server-0.0.2-py3-none-win_amd64.whl

Download URL gk_server-0.0.2-py3-none-win_amd64.whl
Size 15.9 MB
Tags Python 3 Windows x86-64
SHA-256 checksum
How to use checksums
2eeffd26c32f83f9557e5603767c88dcf64b3e96c5ddf782b890e427a5c7e5cf
BLAKE2b-256 checksum
How to use checksums
56d20773919f2068fa8db3ecf72b8bbbb36afff5d21b8f1fe3772a7a7c120387
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.9

Release history Release notifications | RSS feed

This release

0.0.2 This release

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page