Skip to main content

A platform to optimize AND run PyTorch models: license-gated compiler (enhanced planner, persistent cache, multi-accelerator routing), a quantum circuit simulator with batched parameter sweeps, plus a model registry, inference server, and local LLM runtime (int4 quantization, partial offload, hf:/gguf: sources, streaming generation), on top of open-core g2n.

Project description

g2n — optimize and run PyTorch models

g2n is a PyTorch platform with two halves:

  • Optimize — a torch.compile backend: custom fusion passes, Triton kernels, a persistent compile cache, and (on Enterprise) full autotuning. Honest framing: steady-state latency is ~torch.compile parity on most blocks — the wins are cold-start time (cached kernels load in <1 ms per restart), VRAM, and the serving layer.
  • Run — a serving layer that turns those models into a production inference node: model registry, HTTP server, dynamic batching, quantization (int8w/int4w), CUDA-graph replay, and VRAM management for small GPUs. This includes local LLMs: hf: and gguf: models with packed int4 quantization (fused int4 GEMM on CUDA — faster, lower-VRAM decode), partial CPU offload for models bigger than VRAM, and streaming generation over HTTP.

It also ships a quantum circuit simulator (g2n.quantum) — a classical statevector simulator built on torch, for developing and testing quantum algorithms. It is a simulator, not quantum hardware, and never claims otherwise.

import torch, g2n_enterprise as g2n

g2n.activate("G2N-XXXX-XXXX-XXXX")        # once per machine; cached offline

# Optimize: drop-in
model = g2n.compile(my_model)              # or torch.compile(m, backend="g2n")

# Run: registry -> HTTP inference node on :8900
g2n.register_model("clf", "torchscript:/models/clf.pt",
                   precision="int8w",      # weight-only int8: ~4x less weight memory
                   cuda_graph=True, max_batch=16)
g2n.serve()                                # GET/POST http://host:8900/v1/...

# Run a local LLM (Pro+): int4w fits an 8B model on a 6 GB card
g2n.register_model("qwen", "hf:Qwen/Qwen2.5-7B-Instruct",
                   framework="llm", precision="int4w")
g2n.generate("qwen", "Hello!", on_token=print)

Install

pip install g2n-enterprise   # full stack: torch, triton, open-core g2n,
                             # transformers (hf: LLMs), llama.cpp (gguf: LLMs)
pip install g2n              # open-core only (free, Community tier)

The full runtime ships in the base install — no extras to remember. (One caveat: llama-cpp-python compiles from source, so the install needs a C/C++ toolchain, or a prebuilt wheel from the llama-cpp-python docs.)

Python ≥ 3.10. Everything degrades gracefully: no GPU → CPU paths; no license → Community tier; a compile failure → your unmodified model. Your code never breaks because of g2n.

Tiers

Community (free) Pro ($49/mo) Enterprise ($499/mo)
Hybrid fusion + JIT codegen
Enhanced buffer planner (memory fusion)
Persistent compile cache
Model registry + inference server
Quantum simulator: unlimited qubits + fusion ≤24 qubits
Dynamic request batching
Batched quantum parameter sweeps
Multi-accelerator routing + max-autotune
Validated model-zoo configs, priority support

Buy at g2n.dev · seats: Pro 5, Enterprise 25 · 14-day Pro trial on request (sales@g2n.dev).

Documentation

Full documentation lives at https://g2n.dev/docs.

Doc What it answers
Overview What g2n is, how the pieces fit, what it is NOT
Getting started Install → activate → first compile → first serve
Optimize The compiler: what each tier unlocks, cache, autotune
Serving The inference node: precision/quantization, batching, HTTP API
LLM Local LLMs: hf:/gguf: engines, int4w/fp8w, SDPA, offload, streaming /generate
Quantum The circuit simulator
Licensing Keys, activation, seats, offline use, renewal

Honest numbers, always

g2n never ships fabricated benchmarks. Every published number is measured by a script in benchmarks/ on named hardware, and the tooling to measure on your hardware is built in:

import g2n_enterprise as g2n
from g2n_enterprise.serve.reference import example_inputs
g2n.benchmark("mlp", example_inputs(batch=32), rounds=200)
# -> eager vs optimized median latency + peak VRAM, on THIS box

Speedups depend on your model and GPU. Measure before you trust — that includes our numbers.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

g2n_enterprise-3.4.1.tar.gz (183.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

g2n_enterprise-3.4.1-py3-none-any.whl (144.0 kB view details)

Uploaded Python 3

File details

Details for the file g2n_enterprise-3.4.1.tar.gz.

File metadata

  • Download URL: g2n_enterprise-3.4.1.tar.gz
  • Upload date:
  • Size: 183.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.3

File hashes

Hashes for g2n_enterprise-3.4.1.tar.gz
Algorithm Hash digest
SHA256 4cdd5ef3c480d97c934fbec44e1ebb26b733ca867979227d1793a2bbe624ffcf
MD5 76dea90e90a2878f41a4ff81c3d09059
BLAKE2b-256 653b19d0f9eff62949ade31e1a12dd202fabd0ac7573ef3c7f179ec719b3eb80

See more details on using hashes here.

File details

Details for the file g2n_enterprise-3.4.1-py3-none-any.whl.

File metadata

  • Download URL: g2n_enterprise-3.4.1-py3-none-any.whl
  • Upload date:
  • Size: 144.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.3

File hashes

Hashes for g2n_enterprise-3.4.1-py3-none-any.whl
Algorithm Hash digest
SHA256 3059434ae4c54b06dc1811000a491fc594fc8fefa80d878dde7c6b4d45450ace
MD5 3c6ad4abe8ae289c88e14909900073e8
BLAKE2b-256 f20e4b581e9fcf714cb61effc7f3610ee8dfa6731857adb27ca118ff0c6e2c5b

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page