Skip to main content

ggk

One package for working with GGUF models locally: an OpenAI-compatible LLM server, a diffusion image/video/audio generator and a GGUF metadata/tensor editor with a built-in quantizer — three panels on one GUI, powered by one unified engine compiled in a single build on top of gk, an independent tensor library. There is no ggml anywhere in the tree.

Install

pip install ggk

The build compiles the bundled engine (CPU by default, Metal on macOS). GPU backends are opt-in at install time:

GGK_CUDA=1 pip install ggk     # NVIDIA
GGK_HIP=1 pip install ggk      # AMD ROCm
GGK_VULKAN=1 pip install ggk   # Vulkan

Each switch drives the whole engine — the server, the diffusion runtime and the multimodal projectors all evaluate their graphs on the one gk build.

Run

ggk                 # unified GUI — Server / Diffuser / Editor panels
python -m ggk       # same thing

Each panel also runs on its own, exactly like the standalone gguf-server / gguf-diffusion / gguf-editor packages did:

ggk server          # LLM server GUI
ggk diffuser        # image generation GUI
ggk editor          # GGUF editor GUI

And the engines are directly scriptable from the CLI:

ggk server engine -- --model model.gguf --port 8888
ggk diffuser engine -- -m sd.gguf -p "a lighthouse at dusk" -o out.png
ggk editor quantize -m in.gguf -o out-q4_k.gguf --type q4_k
ggk editor devices

Which hardware it picked

Every engine prints the device list to stderr before it does anything else:

gk: found 2 devices
  CUDA0: NVIDIA GeForce RTX 4050 Laptop GPU, 6140 MiB | compute capability = 8.9 | SMs = 20 | shared memory = 99 KiB | built for = 89
  CPU: gk CPU backend, 32014 MiB | SIMD = AVX2 | AVX2 = 1 | FMA = 1 | F16C = 1 | F16_VEC = 1

If a GPU you expected is missing, the install was a CPU-only one (the GPU switches above are opt-in at install time) or its driver was not found — either way the run works, on the CPU, at CPU speed, which is otherwise indistinguishable from a slow GPU. GK_QUIET=1 suppresses the banner.

Layout

vendor/engine/           the unified ggk engine (one CMake build)
  gk/                    the gk compute kernels (CPU + optional GPU backends)
  gk/compat/             the historical ggml C API, implemented on gk
  src/ common/ mtmd/     GGUF LLM runtime
  app/                   the gguf-server HTTP server
  diffusion/             diffusion runtime + CLI
  quantizer/             quantizer shared library (its own quant kernels)
src/ggk/                 the Python package
  server/ diffuser/ editor/   the three panels (backend + web frontend each)
  gui.py static/         the unified 3-panel GUI shell

Nothing above gk/compat/ knows gk exists: the runtimes include the same ggml.h / ggml-backend.h / gguf.h headers and call the same functions they always did, while graph building, allocation, scheduling and the kernels themselves are gk's. See vendor/engine/README.md for the engine's own build options.

The editor's quantizer stays independent — its qz_* codec is compiled both into the quantizer library the editor drives and into gk itself, so the encoder and the runtimes' decoder can never disagree about a GGUF block.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

ggk-0.1.7.tar.gz (35.8 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

ggk-0.1.7-py3-none-win_amd64.whl (47.5 MB view details)

Uploaded Python 3Windows x86-64

File details

Details for the file ggk-0.1.7.tar.gz.

File metadata

  • Download URL: ggk-0.1.7.tar.gz
  • Upload date:
  • Size: 35.8 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.9

File hashes

Hashes for ggk-0.1.7.tar.gz
Algorithm Hash digest
SHA256 c24d6e2553b8f4a06a2b6d8c503d510c083ffdc26148ebfcc12ba7ca2c1caabd
MD5 52ed2cac2216b56a9757b61736cf9c81
BLAKE2b-256 bb0c6a82bf954a16d16e98f6177671f118a79aa6bf071eaa039fcdf4c29333fa

See more details on using hashes here.

File details

Details for the file ggk-0.1.7-py3-none-win_amd64.whl.

File metadata

  • Download URL: ggk-0.1.7-py3-none-win_amd64.whl
  • Upload date:
  • Size: 47.5 MB
  • Tags: Python 3, Windows x86-64
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.9

File hashes

Hashes for ggk-0.1.7-py3-none-win_amd64.whl
Algorithm Hash digest
SHA256 0c6cb86c0ccbb4fbade7cb7a7eadbd6ba5d160be09fe44d794d2a4dadb4cdb83
MD5 c3f5e32cdf3c40aa7ca2595fb7bcda80
BLAKE2b-256 2356d0c44867aa26fb4bafa3b823837dd0d73de160bcabbba301678b74e70e46

See more details on using hashes here.

Release history Release notifications | RSS feed

0.3.5

2 files

0.3.4

2 files

0.3.3

2 files

0.3.2

2 files

0.3.1

2 files

0.3.0

2 files

0.2.9

2 files

0.2.8

2 files

0.2.7

2 files

0.2.6

2 files

0.2.5

2 files

0.2.4

2 files

0.2.3

2 files

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

0.1.9

2 files

0.1.8

1 file

This release

0.1.7 This release

2 files

0.1.6

2 files

0.1.5

2 files

0.1.4

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

0.0.9

2 files

0.0.8

2 files

0.0.7

2 files

0.0.6

2 files

0.0.5

2 files

0.0.4

2 files

0.0.3

2 files

0.0.2

2 files

0.0.1

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page