gemmadecision
Small, local decisions in one Python call. Powered by GemmaDecision-270M.
Install
pip install gemmadecision
Python 3.11+. No API key or server needed. The first call downloads the model (about 0.5 GB); later calls reuse it. CPU, NVIDIA CUDA and Apple Silicon are selected automatically. After downloading, inference runs locally.
Use it in your code
from gemmadecision import decide
team = decide("I was charged twice", choices=["billing", "technical"])
print(team) # selected choice, as a string
For more specific decisions, give each choice a description:
team = decide(
"The same payment appears twice on my statement.",
choices={
"billing": "Handle charges, payments and refunds",
"technical": "Handle crashes, login errors and app problems",
},
question="Which support team should handle this request?",
)
Need the full ranking? Use rank(...) with the same arguments. It returns
ordered candidates, scores and derived probabilities.
PydanticAI
from typing import Literal
from pydantic_ai import Agent
from gemmadecision import GemmaDecisionModel
agent = Agent(
GemmaDecisionModel.local(),
output_type=Literal["billing", "technical"],
)
result = agent.run_sync("I was charged twice")
print(result.output)
This uses PydanticAI's native decision-model interface. Boolean, enum, rubric and finite Pydantic fields are supported. More PydanticAI examples.
Serve it
gemmadecision serve
That starts the Rust-powered Granian HTTP server at http://127.0.0.1:8700.
Interactive API documentation is available at /docs.
from gemmadecision import DecisionClient
with DecisionClient() as client:
result = client.decide(
"I was charged twice",
candidates={"billing": "Payments and refunds", "technical": "App problems"},
)
print(result.choice)
AsyncDecisionClient works with await. To connect PydanticAI to the server,
use GemmaDecisionModel() instead of .local().
More control when you need it
- Hardware settings, offline use, HTTP API and batching
- vLLM backend for Linux/CUDA
- Performance measurements and reproduction
- Docker and releases
The default uses batched PyTorch for model computation and Rust for HTTP serving. vLLM is optional; no Rust compiler is needed to install the package.
This model chooses among supplied options; it does not generate open-ended text. Its scores are rankings, and the derived probabilities are not a guarantee of correctness. Maximums are 2,048 state/question tokens, 768 tokens per choice and 64 choices. See the model card for evaluation and limitations.
Code: Apache-2.0. Model weights: separate Gemma terms. See NOTICE.
Metadata
Release files for gemmadecision 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| gemmadecision-0.1.0.tar.gz | 111.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| gemmadecision-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 150.7 kB
Release files / gemmadecision-0.1.0.tar.gz
| Download URL | gemmadecision-0.1.0.tar.gz |
|---|---|
| Size | 111.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
34dc9777b9702f38fb785bb3dd8167bea95565c63dac4b61d89e0f1f475c206e
|
|
BLAKE2b-256 checksum How to use checksums |
6ec458596f4d72bf337cb7f2d5774dd6bf8981d5f86447b044aa04023afce0f1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.
Transparency logRelease files / gemmadecision-0.1.0-py3-none-any.whl
| Download URL | gemmadecision-0.1.0-py3-none-any.whl |
|---|---|
| Size | 39.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5897050c6e54deca3a5a683fa8a7a30ae2e2d1b496ae60e3362edca17d3cef52
|
|
BLAKE2b-256 checksum How to use checksums |
129f359c665e94b65ff9d7c3c69f6c9761d02f8de59497c10a4831748af6087b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.
Transparency log