graphql-codegen
This library turns a GraphQL schema and your operations into a typed Python client with:
- no imposed transport;
- no validation overhead;
- no runtime dependencies1.
Installation
pip install graphql-codegen
Or, with uv:
uv add --dev graphql-codegen
Quick start
Every example of this README is a file of bookshop, which uses schema.graphqls:
type Query {
# …
"The order with this identifier."
order(id: ID!): Order
# …
}
enum OrderStatus {
PENDING
SHIPPED
DELIVERED
CANCELED
}
type Order {
id: ID!
status: OrderStatus!
# …
}
Write your operations in .graphql files:
query GetOrder($id: ID!) {
order(id: $id) {
id
status
}
}
The generator reads graphql-config, the file GraphQL editor extensions and linters already use2:
# An SDL file here, but globs, introspection results, and server URLs work too.
schema: schema.graphqls
documents: "*.graphql"
# This library's config, under its extension name.
extensions:
pythonCodegen:
# Where imports start, relative to this file's directory.
moduleRoot: ..
# Dotted name from the module root, so written to `bookshop/client`.
package: bookshop.client
Generate the client3:
graphql-codegen bookshop/graphql.config.yml
The generated client:
- depends on nothing but the standard library1;
- holds only the
enumandinputtypes your operations reach, so that it grows with your documents rather than with the schema.
Each GraphQL document also gets its own Python module holding its operations (get_order_graphql.py here).
Your type checker verifies every variable and every field of the response.
The code block below uses assert_never() and assert_type() so you can see what the type checker can prove.
from typing import Literal, assert_never, assert_type
from bookshop.client.runtime import Client
from bookshop.client.schema import OrderStatus
from bookshop.get_order_graphql import GetOrder
from bookshop.transport import transport
client = Client(transport)
data = client(GetOrder({"id": "o1"}))
order = data["order"]
# `Query.order`'s type is nullable, so the type checker requires this test.
if order is None:
print("No such order.")
else:
# `Order.id: ID!` is a `str` on the wire.
assert_type(order["id"], str)
# An `enum` gets a generated alias of the `Literal` of its values.
assert_type(order["status"], OrderStatus)
assert_type(order["status"], Literal["PENDING", "SHIPPED", "DELIVERED", "CANCELED"])
if "total" in order:
# A field not selected in the GraphQL operation can never be there.
assert_never(order)
print(f"Order {order['id']} is {order['status']}.")
Typing
Type checking, not runtime validation
GraphQL is strongly typed, and the server:
- validates each operation against its schema before running it;
- responds with exactly the operation's shape.
Client-side validation of responses thus mostly adds overhead4.
What a Python client still lacks is knowing, while you write order["status"], that the key exists and holds an OrderStatus.
That is a type checker's job, done once, before the code runs.
This library therefore generates exact types for your type checker and leaves each response as decoded from JSON. Only custom scalars with a codec are converted, and only fields asserted non-null are checked.
Operation types
Each operation gets a type for its variables and one for its data, both keyed by the names the GraphQL document uses. Wherever GraphQL lets a value be one of several things, its Python type is a union:
- a selection on a
unionor aninterfaceis one of several types, and becomes one type per concrete type, told apart by__typename(which the generator selects for you); - an
enumis one of several values, and becomes aLiteral; - a
@oneOfinputis one of several fields, and becomes a union of single-key types, so that a value with two keys fails type checking.
These unions are closed for the type checker, but nothing enforces them at runtime, since responses are not validated.
A server may add a type to a union, an implementation to an interface, a member to an enum, or a field to a struct's input type without it being considered a breaking change.
A validating client would have raised an error before your code even ran, but this library lets the new value reach your code as sent, so you can choose what to do with it.
One way is to accept it in a case _: arm:
def status_label(status: OrderStatus, /) -> str:
match status:
case "PENDING":
return "Being prepared"
case "SHIPPED":
return "On its way"
case "DELIVERED":
return "Delivered"
case "CANCELED":
return "Canceled"
case _:
# A member added after this client was generated.
return status.replace("_", " ").capitalize()
The other way is to reject it with case _ as never: assert_never(never), which:
- raises an
AssertionErrorat runtime; - makes the type checker point at every missing
caseonce the client is regenerated.
For instance, app.graphql selects a publication's length, in pages when it is printed and in minutes when it is an audiobook:
query GetPublication($id: ID!) {
publication(id: $id) {
title
... on Printed {
pages: pageCount
# …
}
... on Audiobook {
duration
}
}
}
publication is then a union of one TypedDict per concrete type: Audiobook, and each type that implements Printed.
Testing for a key narrows it to the TypedDicts with that key, and also covers any new type that implements Printed.
length() then narrows it to one TypedDict with a match on __typename, rejecting any other with assert_never():
def length(publication_id: str, /, *, client: Client) -> str:
data = client(GetPublication({"id": publication_id}))
publication = data["publication"]
if publication is None:
return "No such publication."
if "pages" in publication:
return f"{publication['title']} has {publication['pages']} pages."
match publication["__typename"]:
case "Audiobook":
return f"{publication['title']} lasts {publication['duration']} minutes."
case _ as never:
assert_never(never)
No name clashes
Nothing prevents a schema or a document from using names that clash with Python keywords (class, from), standard library names the generated code uses (list, Sequence), or the generator's own helpers.
The names the generator adds, such as _builtins or _GetBookData_book, avoid every other name in their module, so none can shadow another.
Client
Sans-IO
The small sans-IO runtime is copied into the generated package, and its public API is limited to:
from .client import (
AsyncClient as AsyncClient,
AsyncSubscriptionClient as AsyncSubscriptionClient,
Client as Client,
SubscriptionClient as SubscriptionClient,
)
from .error import (
ClientError as ClientError,
Error as Error,
ExecutionError as ExecutionError,
Location as Location,
ProtocolError as ProtocolError,
RequestError as RequestError,
ResponseError as ResponseError,
UnexpectedNullError as UnexpectedNullError,
)
from .injection import OMITTED as OMITTED
from .operation import Operation as Operation, Request as Request
A transport is a function from a request body to a response body, so any HTTP client (synchronous or asynchronous) works, and so does anything else that carries bytes.
Client, AsyncClient, SubscriptionClient, and AsyncSubscriptionClient take the same generated operations, so one generation serves both synchronous and asynchronous code.
Each client forwards every argument but the first (the request) to its transport, type checked against the transport's signature.
As an example, the bookshop's asynchronous transport uses httpx2 and accepts a timeout (and nothing else):
http = httpx2.AsyncClient(base_url="https://bookshop.example")
HEADERS = {"Accept": mime_type.GRAPHQL_RESPONSE, "Content-Type": mime_type.JSON}
async def transport(body: bytes, /, *, timeout: float | None = None) -> bytes:
response = await http.post(
"/graphql", content=body, headers=HEADERS, timeout=timeout
)
# GraphQL over HTTP sends a request error as a response with a 4xx status.
if not response.headers.get("Content-Type", "").startswith(
mime_type.GRAPHQL_RESPONSE
):
response.raise_for_status()
return response.content
A call through a client over this transport can thus pass a timeout:
from bookshop.app_graphql import GetBook
from bookshop.async_transport import AsyncClient
from bookshop.scalar import ISBN
async def title(isbn: ISBN, /, *, client: AsyncClient) -> str:
data = await client(GetBook({"lookup": {"isbn": isbn}}), timeout=5.0)
return data["book"]["title"]
async_transport.py also streams a subscription's Server-Sent Events, and transport.py does both with the standard library alone.
Subscriptions
A subscription client works over any transport yielding one body per event, Server-Sent Events, graphql-transport-ws, or multipart HTTP alike:
def watch(order_id: str, /, *, client: SubscriptionClient) -> list[OrderStatus]:
"""Follow the order until it is delivered, and return its statuses."""
statuses: list[OrderStatus] = []
events = client(OnOrderStatusChanged({"orderId": order_id}))
# Closing the stream, however the loop ends, unsubscribes.
with closing(events):
for event in events:
statuses.append(event["orderStatusChanged"]["status"])
if statuses[-1] == "DELIVERED":
break
return statuses
Merging
Sometimes you only know at runtime which operations to send together, or you want to group the same queries in many combinations, but writing each as its own operation in a .graphql file is impractical.
Combining several operations into one request can also let the server answer faster, seeing the whole picture instead of independent requests asking for overlapping data.
Some other codegen libraries let you build operations at runtime for this, giving up type safety.
This library instead merges operations written ahead of time into one request, so that each result keeps its exact type.
A tuple of queries, or of mutations, runs in one call to the transport, each result typed by its own operation:
def book_and_similar(
isbn: ISBN, text: str, /, *, client: Client
) -> tuple[str, list[str]]:
# Two queries in one call to the transport.
book_data, search_data = client(
(
GetBook({"lookup": {"isbn": isbn}}),
Search({"text": text}),
)
)
assert_type(book_data, GetBookData)
assert_type(search_data, SearchData)
A list built at runtime also runs in one call to the transport, whatever its length. Its results then share one type, the union of its operations' data types:
def look_up(
isbns: Sequence[ISBN], publication_ids: Sequence[str], /, *, client: Client
) -> tuple[GetBookData | GetPublicationData, ...]:
# Any number of queries in one call to the transport.
return client( # ty: ignore[unsound-return-statement] # Pyright and Pyrefly already infer this.
[
*(GetBook({"lookup": {"isbn": isbn}}) for isbn in isbns),
*(GetPublication({"id": id_}) for id_ in publication_ids),
]
)
Errors
A response with errors raises a RequestError when the request failed before execution, and an ExecutionError when it carries partial data:
class ExecutionError(ResponseError, Generic[_Data_co]):
"""The server raised errors executing the request, but sent the rest of the data.
A field that raised is `null`, as is its nearest nullable parent if it is non-null.
"""
data: Final[Mapping[str, object] | None]
def parse_data(self) -> _Data_co | None:
"""Return the data converted as in a response without errors, in a fresh copy.
When several operations are merged, an ExceptionGroup holds one for each operation that fails.
You can also have the client return the error instead of raising it, by calling returning_error() on the request:
- the result is then typed as either the data or an
ExecutionError, so you can tell them apart withisinstance(); - in a merge, you choose for each request whether its error is returned or raised;
- a subscription carries on past an event with errors.
def cancel(order_ids: list[str], /, *, client: Client) -> list[str]:
"""Cancel the orders in one call to the transport, and explain each failure."""
results = client(
[
CancelOrder({"input": {"order": order_id}}).returning_error()
for order_id in order_ids
]
)
# Each error has a note naming its operation and the variables sent.
return [
f"{error.__notes__[0]} {error!s}"
for error in results
if isinstance(error, ExecutionError)
]
def track(order_id: str, /, *, client: Client) -> str:
result = client(GetOrder({"id": order_id}).returning_error())
if isinstance(result, ExecutionError):
data = result.parse_data()
assert_type(data, GetOrderData | None)
return f"Partially loaded: {data} ({result!s})."
Config
Custom scalars
Custom scalars travel as JSON values in a format decided by the server, such as a date as an ISO 8601 string. This library lets you give each one a Python type, with a codec converting its values when they differ from their JSON form:
extensions:
pythonCodegen:
# …
scalars:
DateTime:
type: datetime.datetime
codec:
decode: ..scalar.decode_datetime
encode: ..scalar.encode_datetime
ISBN:
type: ..scalar.ISBN
Money:
type: decimal.Decimal
codec:
decode: decimal.Decimal
encode: str
UUID:
type: uuid.UUID
codec:
decode: uuid.UUID
encode: str
Dotted names in the config resolve as follows:
- a path starting with
..is relative to the directory holding the package; - a bare name, such as
str, is a builtin; - an unconfigured custom scalar is typed
object.
A dotted name must name an attribute of a module, so a method such as datetime.fromisoformat needs a function of its own:
from datetime import datetime
# …
def decode_datetime(value: str, /) -> datetime:
return datetime.fromisoformat(value)
def encode_datetime(value: datetime, /) -> str:
return value.isoformat()
The client then converts each scalar's values on the way in and out, so your code only ever handles their Python types:
def cheaper_than(limit: Decimal, /, *, client: Client) -> list[str]:
data = client(ListBooks({"filter": {"priceBelow": limit}}))
labels: list[str] = []
for book in data["books"]:
price = book["price"]
assert_type(price, Decimal)
labels.append(f"{book['title']}: {price:.2f}")
return labels
Non-null fields
Schemas often make fields nullable, such as a lookup that may find nothing.
Yet you may know more than the schema, such as that a lookup will succeed, or want your code to fail fast on a null without writing if value is None: raise … at every use.
This library brings the idea of Client Controlled Nullability to any server through a client directive:
extensions:
pythonCodegen:
# …
nonNullDirectiveName: nonNull
Asserted on book, the field's type is then not optional:
query GetBook(
"An identifier or an ISBN."
$lookup: BookLookup!
# …
) {
book(lookup: $lookup) @nonNull {
...BookCard
# …
}
}
fragment BookCard on Book {
# …
author {
name
}
}
def describe(isbn: ISBN, /, *, client: Client) -> str:
variables: GetBookVariables = {"lookup": {"isbn": isbn}}
try:
data = client(GetBook(variables))
except UnexpectedNullError as error:
assert error.path == ["book"]
assert error.__notes__ == [f"Raised by `GetBook` with variables {variables!r}."]
return "No such book."
# No `None` check: `@nonNull` took `| None` out of the type.
book = data["book"]
# `@nonNull` only covers `book`, so `author` may still be `None`.
author = book["author"]
by = "an anthology" if author is None else f"by {author['name']}"
Structs
A selection set has a fixed depth, so data of unbounded depth, such as a tree, can only come back as a JSON scalar.
The Struct RFC proposes a new struct keyword: selected without a selection set, a field of a struct type returns its value whole, like a scalar.
This library brings that idea to today's servers with a convention, without waiting for the new keyword:
- each
typethatimplementsa designatedinterfacecarries that scalar in the interface's single field;- that field's payload is typed by the
inputthetypeis named after, even a recursive one.
For instance:
extensions:
pythonCodegen:
# …
structInterfaceName: Struct
type Query {
# …
"The filter of the saved search with this name, exactly as it was saved."
savedSearch(name: String!): BookFilterStruct
}
"A JSON payload with the shape of the input type the implementation is named after."
interface Struct {
value: JSON
}
type BookFilterStruct implements Struct {
value: JSON
}
"A recursive filter on books."
input BookFilter @oneOf {
genre: Genre
author: ID
priceBelow: Money
and: [BookFilter!]
or: [BookFilter!]
not: BookFilter
}
One query receives a BookFilter as data:
query GetSavedSearch($name: String!) {
savedSearch(name: $name) {
value
}
}
Another takes a BookFilter as a variable:
query ListBooks($filter: BookFilter, $first: Int) {
books(filter: $filter, first: $first) {
...BookCard
genre
}
}
So the same BookFilter goes from one to the other as is:
def run_saved_search(name: str, /, *, client: Client) -> list[str]:
data = client(GetSavedSearch({"name": name}))
search = data["savedSearch"]
if search is None:
return []
book_filter = search["value"]
assert_type(book_filter, BookFilter | None)
# Sent back as is.
books = client(ListBooks({"filter": book_filter}))
return [book["title"] for book in books["books"]]
Injectors
Some input values are client's business rather than each client() call's.
For instance, an idempotency key makes retrying a mutation safe:
- the connection drops after the server placed an order;
- the transport sends the order again;
- the key, unique to the order, tells the server it already placed it, so it does not charge the customer twice.
GraphQL has no built-in idempotency, so implementing it usually means adding the key as an argument or an input field of each mutation that needs it.
When many different mutation operations require idempotency, it becomes the concern of all their client() calls, each having to get hold of a key.
This applies to other concepts too, such as database transaction IDs.
This library handles such a value in one place instead, with an injector given when client is constructed.
The client then passes it to every variable or input field with the name given in the config.
No variables' type accepts it, so that no client() call can pass one by mistake:
extensions:
pythonCodegen:
# …
injectorNames: [idempotencyKey]
PlaceOrderInput holds one, for instance:
input PlaceOrderInput {
"Makes placing the same order twice harmless: the client sends a new one per order."
idempotencyKey: UUID
The generated package's injection.py module types the injectors you must supply:
class InjectorFunctions(_compat.TypedDict, closed=True):
"""The functions supplying each injected value, by name.
Where the value may be null, one returning `OMITTED` leaves it out, and one returning `None` sends `null`."""
idempotencyKey: _typing.NotRequired[_abc.Callable[[], _scalar.UUID | None | _OMITTED]]
def injectors(functions: InjectorFunctions, /) -> _injection._Injectors:
"""Return what a client calls to supply the injected values, each serialized as its type says."""
return _injection._Injectors(functions, injector_functions_type=InjectorFunctions)
You give the client its injectors once when building it:
client = Client(
transport,
injectors=injectors({"idempotencyKey": uuid4}),
)
And no call can pass a key:
def order(book_id: str, address: Address, /, *, client: Client) -> str:
data = client(
PlaceOrder(
{"input": {"lines": [{"book": book_id}], "shippingAddress": address}}
)
)
The client injects new values on each call, so a retry belongs in your transport, which sends the same body, key included, again:
def transport(body: bytes, /, *, timeout: float | None = None) -> bytes:
retries = 2
while True:
try:
response = _post(
"/graphql", body, accept=mime_type.GRAPHQL_RESPONSE, timeout=timeout
)
except HTTPError as error:
# GraphQL over HTTP sends a request error as a response with a 4xx status.
if error.headers.get_content_type() != mime_type.GRAPHQL_RESPONSE:
raise
response = error
except ConnectionError:
# The response was lost.
if not retries:
raise
retries -= 1
continue
with response:
return response.read()
Colocation
Most other codegen libraries gather every operation in one generated module, often inside a single class.
A feature's operations then live away from its .graphql files and the code calling them.
A project split into several packages cannot have each package own its operations either.
This library generates each operation as a module-level constant, so it can be colocated with both the document it comes from and the code calling it:
extensions:
pythonCodegen:
# …
documentSiblingModule: "{document}_graphql"
This pattern puts get_order.graphql's operations and fragments in get_order_graphql.py.
Code calling these operations imports them from there:
from bookshop.client.schema import OrderStatus
from bookshop.get_order_graphql import GetOrder
Python API
generate() does the same as the command:
class _Params(TypedDict, closed=True):
document: DocumentNode
schema: GraphQLSchema
config: Config
def generate(**args: Unpack[_Params]) -> dict[PurePosixPath, bytes]:
"""Pure function returning the content of each file of the generated package by its path."""
The public API is limited to:
from graphql_codegen.config import Config as Config
from graphql_codegen.document_sibling_module import (
DocumentSiblingModule as DocumentSiblingModule,
)
from graphql_codegen.generate import generate as generate
from graphql_codegen.package_location import PackageLocation as PackageLocation
from graphql_codegen.scalar import Codec as Codec, Scalar as Scalar
-
Before Python 3.15, the client also needs
typing_extensions, for features the standard library'stypingdoes not have yet, so add it to your project's own dependencies. ↩ ↩2 -
This library is partly bootstrapped: to fetch a schema from a URL, it uses a client it generated itself from
_introspection.graphql. ↩ -
A YAML config, like the quick start's, needs the
yamlextra: installgraphql-codegen[yaml]instead. JSON and TOML configs need nothing more. ↩ -
Well-behaved GraphQL APIs avoid breaking changes, adding fields and deprecating old ones instead. Most breaking changes made after generation never reach the code anyway: removing or renaming a field, changing its arguments, or turning its type from an object into a leaf or the reverse invalidates the operation, so the server rejects it, and the
client()call raises aRequestError. The few left would fail client-side validation as soon as the response arrives, but otherwise propagate deeper in your code, possibly unnoticed: a field becoming nullable, its leaf type changing, or it switching between a list and a single value. The more often you regenerate against the deployed schema, the sooner your type checker catches such a change. ↩
Metadata
Release files for graphql-codegen 0.1.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| graphql_codegen-0.1.3-py3-none-any.whl | Python 3 | none | any | Details |
Release files / graphql_codegen-0.1.3-py3-none-any.whl
| Download URL | graphql_codegen-0.1.3-py3-none-any.whl |
|---|---|
| Size | 94.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5a2354a8f4e370e3b481c152d2c7da6164fa635325462422bc0d8455cadbba75
|
|
BLAKE2b-256 checksum How to use checksums |
8a8380d1fae43dfeb4afbc8aee37d5860a6a27a5167ff6ef3b5931921f28b839
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.24 {"installer":{"name":"uv","version":"0.12.24","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"26.04","id":"resolute","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|