Skip to main content

A platform to optimize AND run PyTorch models: license-gated compiler (enhanced planner, persistent cache, multi-accelerator routing), a quantum circuit simulator with batched parameter sweeps, plus a model registry, inference server, and local LLM runtime (int4 quantization, partial offload, hf:/gguf: sources, streaming generation), on top of open-core g2n.

Project description

g2n — optimize and run PyTorch models

g2n is a PyTorch platform with two halves:

  • Optimize — a torch.compile backend: custom fusion passes, Triton kernels, a persistent compile cache, and (on Enterprise) full autotuning. Honest framing: steady-state latency is ~torch.compile parity on most blocks — the wins are cold-start time (cached kernels load in <1 ms per restart), VRAM, and the serving layer.
  • Run — a serving layer that turns those models into a production inference node: model registry, HTTP server, dynamic batching, quantization (int8w/int4w), CUDA-graph replay, and VRAM management for small GPUs. This includes local LLMs: hf: and gguf: models with packed int4 quantization (fused int4 GEMM on CUDA — faster, lower-VRAM decode), partial CPU offload for models bigger than VRAM, and streaming generation over HTTP.

It also ships a quantum circuit simulator (g2n.quantum) — a classical statevector simulator built on torch, for developing and testing quantum algorithms. It is a simulator, not quantum hardware, and never claims otherwise.

import torch, g2n_enterprise as g2n

g2n.activate("G2N-XXXX-XXXX-XXXX")        # once per machine; cached offline

# Optimize: drop-in
model = g2n.compile(my_model)              # or torch.compile(m, backend="g2n")

# Run: registry -> HTTP inference node on :8900
g2n.register_model("clf", "torchscript:/models/clf.pt",
                   precision="int8w",      # weight-only int8: ~4x less weight memory
                   cuda_graph=True, max_batch=16)
g2n.serve()                                # GET/POST http://host:8900/v1/...

# Run a local LLM (Pro+): int4w fits an 8B model on a 6 GB card
g2n.register_model("qwen", "hf:Qwen/Qwen2.5-7B-Instruct",
                   framework="llm", precision="int4w")
g2n.generate("qwen", "Hello!", on_token=print)

Install

pip install g2n-enterprise   # full stack: torch, triton, open-core g2n,
                             # transformers (hf: LLMs), llama.cpp (gguf: LLMs)
pip install g2n              # open-core only (free, Community tier)

The full runtime ships in the base install — no extras to remember. (One caveat: llama-cpp-python compiles from source, so the install needs a C/C++ toolchain, or a prebuilt wheel from the llama-cpp-python docs.)

Python ≥ 3.10. Everything degrades gracefully: no GPU → CPU paths; no license → Community tier; a compile failure → your unmodified model. Your code never breaks because of g2n.

Tiers

Community (free) Pro ($49/mo) Enterprise ($499/mo)
Hybrid fusion + JIT codegen
Enhanced buffer planner (memory fusion)
Persistent compile cache
Model registry + inference server
Quantum simulator: unlimited qubits + fusion ≤24 qubits
Dynamic request batching
Batched quantum parameter sweeps
Multi-accelerator routing + max-autotune
Validated model-zoo configs, priority support

Buy at g2n.dev · seats: Pro 5, Enterprise 25 · 14-day Pro trial on request (sales@g2n.dev).

Documentation

Full documentation lives at https://g2n.dev/docs.

Doc What it answers
Overview What g2n is, how the pieces fit, what it is NOT
Getting started Install → activate → first compile → first serve
Optimize The compiler: what each tier unlocks, cache, autotune
Serving The inference node: precision/quantization, batching, HTTP API
LLM Local LLMs: hf:/gguf: engines, int4w/fp8w, SDPA, offload, streaming /generate
Quantum The circuit simulator
Licensing Keys, activation, seats, offline use, renewal

Honest numbers, always

g2n never ships fabricated benchmarks. Every published number is measured by a script in benchmarks/ on named hardware, and the tooling to measure on your hardware is built in:

import g2n_enterprise as g2n
from g2n_enterprise.serve.reference import example_inputs
g2n.benchmark("mlp", example_inputs(batch=32), rounds=200)
# -> eager vs optimized median latency + peak VRAM, on THIS box

Speedups depend on your model and GPU. Measure before you trust — that includes our numbers.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

g2n_enterprise-3.6.0.tar.gz (211.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

g2n_enterprise-3.6.0-py3-none-any.whl (160.2 kB view details)

Uploaded Python 3

File details

Details for the file g2n_enterprise-3.6.0.tar.gz.

File metadata

  • Download URL: g2n_enterprise-3.6.0.tar.gz
  • Upload date:
  • Size: 211.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.3

File hashes

Hashes for g2n_enterprise-3.6.0.tar.gz
Algorithm Hash digest
SHA256 1ca74bcd05609dc790d514c7e7cb144c9a62802b08a76b6309b0c13590f59253
MD5 6a7ca4940e6a88e6eff1aec286d93d95
BLAKE2b-256 aa56b0f40b7f719b2d44c467ef63110e6bddcf1461ba3a0c4d1e81da430d2137

See more details on using hashes here.

File details

Details for the file g2n_enterprise-3.6.0-py3-none-any.whl.

File metadata

  • Download URL: g2n_enterprise-3.6.0-py3-none-any.whl
  • Upload date:
  • Size: 160.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.3

File hashes

Hashes for g2n_enterprise-3.6.0-py3-none-any.whl
Algorithm Hash digest
SHA256 563bf58c875996d4f20670f00f0fdd39e03c2f5213787e1f7fbf46c10328fe2e
MD5 ce9de1dea051a18c3a05e933a47ecc98
BLAKE2b-256 e613980cc2ba79c0ea4c201dfbf516b3b1cc926bae19040bf53ba04a203071c8

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page