Skip to main content

ggk

One package for working with GGUF models locally: an OpenAI-compatible LLM server, a diffusion image/video/audio generator and a GGUF metadata/tensor editor with a built-in quantizer — three panels on one GUI, powered by one unified engine compiled in a single build on top of gk, an independent tensor library. There is no ggml anywhere in the tree.

Install

pip install ggk

The build compiles the bundled engine (CPU by default, Metal on macOS). GPU backends are opt-in at install time:

GGK_CUDA=1 pip install ggk     # NVIDIA
GGK_HIP=1 pip install ggk      # AMD ROCm
GGK_VULKAN=1 pip install ggk   # Vulkan

Each switch drives the whole engine — the server, the diffusion runtime and the multimodal projectors all evaluate their graphs on the one gk build.

Building a wheel

python -m build --wheel

--wheel is not optional for a GPU build. Plain python -m build builds an sdist first and then compiles the wheel from it in a temporary directory, so every run starts from scratch and a multi-hour CUDA build spends those hours inside a directory the OS is free to sweep — on Windows that surfaces at the very end as FileNotFoundError: ...\wheel\scripts from scikit-build-core's packaging step, long after the compile succeeded. With --wheel the CMake build directory stays at build/{wheel_tag} in the tree and rebuilds are incremental.

On Windows a CUDA build needs MSVC, because that is the only host compiler nvcc accepts there — run it from a vcvars64 shell. Use Ninja rather than -G "Visual Studio 17 2022": scikit-build-core passes no -j on the pyproject path and CMake's Visual Studio generator does not set /MP, so an MSBuild-driven build compiles one file at a time.

call "C:\Program Files\Microsoft Visual Studio\2022\Community\VC\Auxiliary\Build\vcvars64.bat"
set CMAKE_ARGS=-G Ninja -DGGK_CUDA=ON
python -m build --wheel

The build works its own architecture list out from nvcc and the installed GPUs. A wheel built on one machine for another should say what it targets: -DGK_CUDA_ARCHITECTURES="75;86;89;120".

Run

ggk                 # unified GUI — Server / Diffuser / Editor panels
python -m ggk       # same thing

Each panel also runs on its own, exactly like the standalone gguf-server / gguf-diffusion / gguf-editor packages did:

ggk server          # LLM server GUI
ggk diffuser        # image generation GUI
ggk editor          # GGUF editor GUI

And the engines are directly scriptable from the CLI:

ggk server engine -- --model model.gguf --port 8888
ggk diffuser engine -- -m sd.gguf -p "a lighthouse at dusk" -o out.png
ggk editor quantize -m in.gguf -o out-q4_k.gguf --type q4_k
ggk editor devices

Which hardware it picked

Every engine prints the device list to stderr before it does anything else:

gk: found 2 devices
  CUDA0: NVIDIA GeForce RTX 4050 Laptop GPU, 6140 MiB | compute capability = 8.9 | SMs = 20 | shared memory = 99 KiB | built for = 89
  CPU: gk CPU backend, 32014 MiB | SIMD = AVX2 | AVX2 = 1 | FMA = 1 | F16C = 1 | F16_VEC = 1

If a GPU you expected is missing, the install was a CPU-only one (the GPU switches above are opt-in at install time) or its driver was not found — either way the run works, on the CPU, at CPU speed, which is otherwise indistinguishable from a slow GPU. GK_QUIET=1 suppresses the banner.

Layout

vendor/engine/           the unified ggk engine (one CMake build)
  gk/                    the gk compute kernels (CPU + optional GPU backends)
  gk/compat/             the historical ggml C API, implemented on gk
  src/ common/ mtmd/     GGUF LLM runtime
  app/                   the gguf-server HTTP server
  diffusion/             diffusion runtime + CLI
  quantizer/             quantizer shared library (its own quant kernels)
src/ggk/                 the Python package
  server/ diffuser/ editor/   the three panels (backend + web frontend each)
  gui.py static/         the unified 3-panel GUI shell

Nothing above gk/compat/ knows gk exists: the runtimes include the same ggml.h / ggml-backend.h / gguf.h headers and call the same functions they always did, while graph building, allocation, scheduling and the kernels themselves are gk's. See vendor/engine/README.md for the engine's own build options.

The editor's quantizer stays independent — its qz_* codec is compiled both into the quantizer library the editor drives and into gk itself, so the encoder and the runtimes' decoder can never disagree about a GGUF block.

Metadata

Release files for ggk 0.3.6

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ggk 0.3.6
File Size Uploaded
ggk-0.3.6.tar.gz 36.0 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for ggk 0.3.6
File Interpreter ABI Platform
ggk-0.3.6-py3-none-win_amd64.whl Python 3 none Windows x86-64 Details

Total release size: 100.2 MB

Release files / ggk-0.3.6.tar.gz

Download URL ggk-0.3.6.tar.gz
Size 36.0 MB
Tags Source
SHA-256 checksum
How to use checksums
69be67ac83d54e01d0c2f7a0bc956275b62ce8a4b09181fdbe478afe5cc43366
BLAKE2b-256 checksum
How to use checksums
f5f2b8807a77aec9abf92b616662813e987d28450f87bef8a59dc57689b5f369
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.9

Release files / ggk-0.3.6-py3-none-win_amd64.whl

Download URL ggk-0.3.6-py3-none-win_amd64.whl
Size 64.3 MB
Tags Python 3 Windows x86-64
SHA-256 checksum
How to use checksums
2796848f170ce6df54cb225c5c87ce1238ae1b1a298599768338ef679f42d683
BLAKE2b-256 checksum
How to use checksums
b04a6590fc16867782676ab28a7cc0a4fd69d81b6fcc1b507cf986e1a7e0d6aa
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.9

Release history Release notifications | RSS feed

0.7.9

2 release files

0.7.8

2 release files

0.7.7

2 release files

0.7.6

2 release files

0.7.5

2 release files

0.7.4

2 release files

0.7.3

2 release files

0.7.2

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.9

2 release files

0.6.8

2 release files

0.6.7

2 release files

0.6.6

2 release files

0.6.5

2 release files

0.6.4

2 release files

0.6.3

2 release files

0.6.2

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.9

2 release files

0.5.8

2 release files

0.5.7

2 release files

0.5.6

2 release files

0.5.5

2 release files

0.5.4

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.9

2 release files

0.4.8

2 release files

0.4.7

2 release files

0.4.6

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.9

2 release files

0.3.8

2 release files

0.3.7

2 release files

This release

0.3.6 This release

2 release files

0.3.5

2 release files

0.3.4

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.9

2 release files

0.2.8

2 release files

0.2.7

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.9

2 release files

0.1.8

1 release file

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page