Skip to main content

graphql-codegen

PyPI Python Coverage License

This library turns a GraphQL schema and your operations into a typed Python client with:

  • no imposed transport;
  • no validation overhead;
  • no runtime dependencies1.

Installation

pip install graphql-codegen

Or, with uv:

uv add --dev graphql-codegen

Quick start

Every example of this README is a file of bookshop, which uses schema.graphqls:

type Query {
  # …
  "The order with this identifier."
  order(id: ID!): Order
  # …
}

enum OrderStatus {
  PENDING
  SHIPPED
  DELIVERED
  CANCELED
}

type Order {
  id: ID!
  status: OrderStatus!
  # …
}

Write your operations in .graphql files:

query GetOrder($id: ID!) {
  order(id: $id) {
    id
    status
  }
}

The generator reads graphql-config, the file GraphQL editor extensions and linters already use2:

# An SDL file here, but globs, introspection results, and server URLs work too.
schema: schema.graphqls
documents: "*.graphql"
# This library's config, under its extension name.
extensions:
  pythonCodegen:
    # Where imports start, relative to this file's directory.
    moduleRoot: ..
    # Dotted name from the module root, so written to `bookshop/client`.
    package: bookshop.client

Generate the client3:

graphql-codegen bookshop/graphql.config.yml

The generated client:

  • depends on nothing but the standard library1;
  • holds only the enum and input types your operations reach, so that it grows with your documents rather than with the schema.

Each GraphQL document also gets its own Python module holding its operations (get_order_graphql.py here).

Your type checker verifies every variable and every field of the response. The code block below uses assert_never() and assert_type() so you can see what the type checker can prove.

from typing import Literal, assert_never, assert_type

from bookshop.client.runtime import Client
from bookshop.client.schema import OrderStatus
from bookshop.get_order_graphql import GetOrder
from bookshop.transport import transport

client = Client(transport)
data = client(GetOrder({"id": "o1"}))
order = data["order"]

# `Query.order`'s type is nullable, so the type checker requires this test.
if order is None:
    print("No such order.")
else:
    # `Order.id: ID!` is a `str` on the wire.
    assert_type(order["id"], str)

    # An `enum` gets a generated alias of the `Literal` of its values.
    assert_type(order["status"], OrderStatus)
    assert_type(order["status"], Literal["PENDING", "SHIPPED", "DELIVERED", "CANCELED"])

    if "total" in order:
        # A field not selected in the GraphQL operation can never be there.
        assert_never(order)

    print(f"Order {order['id']} is {order['status']}.")

Typing

Type checking, not runtime validation

GraphQL is strongly typed, and the server:

  • validates each operation against its schema before running it;
  • responds with exactly the operation's shape.

Client-side validation of responses thus mostly adds overhead4. What a Python client still lacks is knowing, while you write order["status"], that the key exists and holds an OrderStatus. That is a type checker's job, done once, before the code runs.

This library therefore generates exact types for your type checker and leaves each response as decoded from JSON. Only custom scalars with a codec are converted, and only fields asserted non-null are checked.

Operation types

Each operation gets a type for its variables and one for its data, both keyed by the names the GraphQL document uses. Wherever GraphQL lets a value be one of several things, its Python type is a union:

  • a selection on a union or an interface is one of several types, and becomes one type per concrete type, told apart by __typename (which the generator selects for you);
  • an enum is one of several values, and becomes a Literal;
  • a @oneOf input is one of several fields, and becomes a union of single-key types, so that a value with two keys fails type checking.

These unions are closed for the type checker, but nothing enforces them at runtime, since responses are not validated.

A server may add a type to a union, an implementation to an interface, a member to an enum, or a field to a struct's input type without it being considered a breaking change. A validating client would have raised an error before your code even ran, but this library lets the new value reach your code as sent, so you can choose what to do with it.

One way is to accept it in a case _: arm:

def status_label(status: OrderStatus, /) -> str:
    match status:
        case "PENDING":
            return "Being prepared"
        case "SHIPPED":
            return "On its way"
        case "DELIVERED":
            return "Delivered"
        case "CANCELED":
            return "Canceled"
        case _:
            # A member added after this client was generated.
            return status.replace("_", " ").capitalize()

The other way is to reject it with case _ as never: assert_never(never), which:

  • raises an AssertionError at runtime;
  • makes the type checker point at every missing case once the client is regenerated.

For instance, app.graphql selects a publication's length, in pages when it is printed and in minutes when it is an audiobook:

query GetPublication($id: ID!) {
  publication(id: $id) {
    title
    ... on Printed {
      pages: pageCount
      # …
    }
    ... on Audiobook {
      duration
    }
  }
}

publication is then a union of one TypedDict per concrete type: Audiobook, and each type that implements Printed. Testing for a key narrows it to the TypedDicts with that key, and also covers any new type that implements Printed.

length() then narrows it to one TypedDict with a match on __typename, rejecting any other with assert_never():

def length(publication_id: str, /, *, client: Client) -> str:
    data = client(GetPublication({"id": publication_id}))
    publication = data["publication"]

    if publication is None:
        return "No such publication."

    if "pages" in publication:
        return f"{publication['title']} has {publication['pages']} pages."

    match publication["__typename"]:
        case "Audiobook":
            return f"{publication['title']} lasts {publication['duration']} minutes."
        case _ as never:
            assert_never(never)

No name clashes

Nothing prevents a schema or a document from using names that clash with Python keywords (class, from), standard library names the generated code uses (list, Sequence), or the generator's own helpers.

The names the generator adds, such as _builtins or _GetBookData_book, avoid every other name in their module, so none can shadow another.

Client

Sans-IO

The small sans-IO runtime is copied into the generated package, and its public API is limited to:

from .client import (
    AsyncClient as AsyncClient,
    AsyncSubscriptionClient as AsyncSubscriptionClient,
    Client as Client,
    SubscriptionClient as SubscriptionClient,
)
from .error import (
    ClientError as ClientError,
    Error as Error,
    ExecutionError as ExecutionError,
    Location as Location,
    ProtocolError as ProtocolError,
    RequestError as RequestError,
    ResponseError as ResponseError,
    UnexpectedNullError as UnexpectedNullError,
)
from .injection import OMITTED as OMITTED
from .operation import Operation as Operation, Request as Request

A transport is a function from a request body to a response body, so any HTTP client (synchronous or asynchronous) works, and so does anything else that carries bytes. Client, AsyncClient, SubscriptionClient, and AsyncSubscriptionClient take the same generated operations, so one generation serves both synchronous and asynchronous code. Each client forwards every argument but the first (the request) to its transport, type checked against the transport's signature.

As an example, the bookshop's asynchronous transport uses httpx2 and accepts a timeout (and nothing else):

http = httpx2.AsyncClient(base_url="https://bookshop.example")
HEADERS = {"Accept": mime_type.GRAPHQL_RESPONSE, "Content-Type": mime_type.JSON}


async def transport(body: bytes, /, *, timeout: float | None = None) -> bytes:
    response = await http.post(
        "/graphql", content=body, headers=HEADERS, timeout=timeout
    )

    # GraphQL over HTTP sends a request error as a response with a 4xx status.
    if not response.headers.get("Content-Type", "").startswith(
        mime_type.GRAPHQL_RESPONSE
    ):
        response.raise_for_status()

    return response.content

A call through a client over this transport can thus pass a timeout:

from bookshop.app_graphql import GetBook
from bookshop.async_transport import AsyncClient
from bookshop.scalar import ISBN


async def title(isbn: ISBN, /, *, client: AsyncClient) -> str:
    data = await client(GetBook({"lookup": {"isbn": isbn}}), timeout=5.0)
    return data["book"]["title"]

async_transport.py also streams a subscription's Server-Sent Events, and transport.py does both with the standard library alone.

Subscriptions

A subscription client works over any transport yielding one body per event, Server-Sent Events, graphql-transport-ws, or multipart HTTP alike:

def watch(order_id: str, /, *, client: SubscriptionClient) -> list[OrderStatus]:
    """Follow the order until it is delivered, and return its statuses."""
    statuses: list[OrderStatus] = []
    events = client(OnOrderStatusChanged({"orderId": order_id}))

    # Closing the stream, however the loop ends, unsubscribes.
    with closing(events):
        for event in events:
            statuses.append(event["orderStatusChanged"]["status"])

            if statuses[-1] == "DELIVERED":
                break

    return statuses

Merging

Sometimes you only know at runtime which operations to send together, or you want to group the same queries in many combinations, but writing each as its own operation in a .graphql file is impractical. Combining several operations into one request can also let the server answer faster, seeing the whole picture instead of independent requests asking for overlapping data. Some other codegen libraries let you build operations at runtime for this, giving up type safety. This library instead merges operations written ahead of time into one request, so that each result keeps its exact type.

A tuple of queries, or of mutations, runs in one call to the transport, each result typed by its own operation:

def book_and_similar(
    isbn: ISBN, text: str, /, *, client: Client
) -> tuple[str, list[str]]:
    # Two queries in one call to the transport.
    book_data, search_data = client(
        (
            GetBook({"lookup": {"isbn": isbn}}),
            Search({"text": text}),
        )
    )
    assert_type(book_data, GetBookData)
    assert_type(search_data, SearchData)

A list built at runtime also runs in one call to the transport, whatever its length. Its results then share one type, the union of its operations' data types:

def look_up(
    isbns: Sequence[ISBN], publication_ids: Sequence[str], /, *, client: Client
) -> tuple[GetBookData | GetPublicationData, ...]:
    # Any number of queries in one call to the transport.
    return client(  # ty: ignore[unsound-return-statement]  # Pyright and Pyrefly already infer this.
        [
            *(GetBook({"lookup": {"isbn": isbn}}) for isbn in isbns),
            *(GetPublication({"id": id_}) for id_ in publication_ids),
        ]
    )

Errors

A response with errors raises a RequestError when the request failed before execution, and an ExecutionError when it carries partial data:

class ExecutionError(ResponseError, Generic[_Data_co]):
    """The server raised errors executing the request, but sent the rest of the data.

    A field that raised is `null`, as is its nearest nullable parent if it is non-null.
    """

    data: Final[Mapping[str, object] | None]

    def parse_data(self) -> _Data_co | None:
        """Return the data converted as in a response without errors, in a fresh copy.

When several operations are merged, an ExceptionGroup holds one for each operation that fails.

You can also have the client return the error instead of raising it, by calling returning_error() on the request:

  • the result is then typed as either the data or an ExecutionError, so you can tell them apart with isinstance();
  • in a merge, you choose for each request whether its error is returned or raised;
  • a subscription carries on past an event with errors.
def cancel(order_ids: list[str], /, *, client: Client) -> list[str]:
    """Cancel the orders in one call to the transport, and explain each failure."""
    results = client(
        [
            CancelOrder({"input": {"order": order_id}}).returning_error()
            for order_id in order_ids
        ]
    )
    # Each error has a note naming its operation and the variables sent.
    return [
        f"{error.__notes__[0]} {error!s}"
        for error in results
        if isinstance(error, ExecutionError)
    ]


def track(order_id: str, /, *, client: Client) -> str:
    result = client(GetOrder({"id": order_id}).returning_error())

    if isinstance(result, ExecutionError):
        data = result.parse_data()
        assert_type(data, GetOrderData | None)
        return f"Partially loaded: {data} ({result!s})."

Config

Custom scalars

Custom scalars travel as JSON values in a format decided by the server, such as a date as an ISO 8601 string. This library lets you give each one a Python type, with a codec converting its values when they differ from their JSON form:

extensions:
  pythonCodegen:
    # …
    scalars:
      DateTime:
        type: datetime.datetime
        codec:
          decode: ..scalar.decode_datetime
          encode: ..scalar.encode_datetime
      ISBN:
        type: ..scalar.ISBN
      Money:
        type: decimal.Decimal
        codec:
          decode: decimal.Decimal
          encode: str
      UUID:
        type: uuid.UUID
        codec:
          decode: uuid.UUID
          encode: str

Dotted names in the config resolve as follows:

  • a path starting with .. is relative to the directory holding the package;
  • a bare name, such as str, is a builtin;
  • an unconfigured custom scalar is typed object.

A dotted name must name an attribute of a module, so a method such as datetime.fromisoformat needs a function of its own:

from datetime import datetime
# …


def decode_datetime(value: str, /) -> datetime:
    return datetime.fromisoformat(value)


def encode_datetime(value: datetime, /) -> str:
    return value.isoformat()

The client then converts each scalar's values on the way in and out, so your code only ever handles their Python types:

def cheaper_than(limit: Decimal, /, *, client: Client) -> list[str]:
    data = client(ListBooks({"filter": {"priceBelow": limit}}))
    labels: list[str] = []

    for book in data["books"]:
        price = book["price"]
        assert_type(price, Decimal)
        labels.append(f"{book['title']}: {price:.2f}")

    return labels

Non-null fields

Schemas often make fields nullable, such as a lookup that may find nothing. Yet you may know more than the schema, such as that a lookup will succeed, or want your code to fail fast on a null without writing if value is None: raise … at every use. This library brings the idea of Client Controlled Nullability to any server through a client directive:

extensions:
  pythonCodegen:
    # …
    nonNullDirectiveName: nonNull

Asserted on book, the field's type is then not optional:

query GetBook(
  "An identifier or an ISBN."
  $lookup: BookLookup!
  # …
) {
  book(lookup: $lookup) @nonNull {
    ...BookCard
    # …
  }
}

fragment BookCard on Book {
  # …
  author {
    name
  }
}
def describe(isbn: ISBN, /, *, client: Client) -> str:
    variables: GetBookVariables = {"lookup": {"isbn": isbn}}

    try:
        data = client(GetBook(variables))
    except UnexpectedNullError as error:
        assert error.path == ["book"]
        assert error.__notes__ == [f"Raised by `GetBook` with variables {variables!r}."]
        return "No such book."

    # No `None` check: `@nonNull` took `| None` out of the type.
    book = data["book"]

    # `@nonNull` only covers `book`, so `author` may still be `None`.
    author = book["author"]
    by = "an anthology" if author is None else f"by {author['name']}"

Structs

A selection set has a fixed depth, so data of unbounded depth, such as a tree, can only come back as a JSON scalar. The Struct RFC proposes a new struct keyword: selected without a selection set, a field of a struct type returns its value whole, like a scalar. This library brings that idea to today's servers with a convention, without waiting for the new keyword:

  • each type that implements a designated interface carries that scalar in the interface's single field;
  • that field's payload is typed by the input the type is named after, even a recursive one.

For instance:

extensions:
  pythonCodegen:
    # …
    structInterfaceName: Struct
type Query {
  # …
  "The filter of the saved search with this name, exactly as it was saved."
  savedSearch(name: String!): BookFilterStruct
}

"A JSON payload with the shape of the input type the implementation is named after."
interface Struct {
  value: JSON
}

type BookFilterStruct implements Struct {
  value: JSON
}

"A recursive filter on books."
input BookFilter @oneOf {
  genre: Genre
  author: ID
  priceBelow: Money
  and: [BookFilter!]
  or: [BookFilter!]
  not: BookFilter
}

One query receives a BookFilter as data:

query GetSavedSearch($name: String!) {
  savedSearch(name: $name) {
    value
  }
}

Another takes a BookFilter as a variable:

query ListBooks($filter: BookFilter, $first: Int) {
  books(filter: $filter, first: $first) {
    ...BookCard
    genre
  }
}

So the same BookFilter goes from one to the other as is:

def run_saved_search(name: str, /, *, client: Client) -> list[str]:
    data = client(GetSavedSearch({"name": name}))
    search = data["savedSearch"]

    if search is None:
        return []

    book_filter = search["value"]
    assert_type(book_filter, BookFilter | None)
    # Sent back as is.
    books = client(ListBooks({"filter": book_filter}))
    return [book["title"] for book in books["books"]]

Injectors

Some input values are client's business rather than each client() call's.

For instance, an idempotency key makes retrying a mutation safe:

  1. the connection drops after the server placed an order;
  2. the transport sends the order again;
  3. the key, unique to the order, tells the server it already placed it, so it does not charge the customer twice.

GraphQL has no built-in idempotency, so implementing it usually means adding the key as an argument or an input field of each mutation that needs it. When many different mutation operations require idempotency, it becomes the concern of all their client() calls, each having to get hold of a key. This applies to other concepts too, such as database transaction IDs.

This library handles such a value in one place instead, with an injector given when client is constructed. The client then passes it to every variable or input field with the name given in the config. No variables' type accepts it, so that no client() call can pass one by mistake:

extensions:
  pythonCodegen:
    # …
    injectorNames: [idempotencyKey]

PlaceOrderInput holds one, for instance:

input PlaceOrderInput {
  "Makes placing the same order twice harmless: the client sends a new one per order."
  idempotencyKey: UUID

The generated package's injection.py module types the injectors you must supply:

class InjectorFunctions(_compat.TypedDict, closed=True):
    """The functions supplying each injected value, by name.

    Where the value may be null, one returning `OMITTED` leaves it out, and one returning `None` sends `null`."""
    idempotencyKey: _typing.NotRequired[_abc.Callable[[], _scalar.UUID | None | _OMITTED]]

def injectors(functions: InjectorFunctions, /) -> _injection._Injectors:
    """Return what a client calls to supply the injected values, each serialized as its type says."""
    return _injection._Injectors(functions, injector_functions_type=InjectorFunctions)

You give the client its injectors once when building it:

client = Client(
    transport,
    injectors=injectors({"idempotencyKey": uuid4}),
)

And no call can pass a key:

def order(book_id: str, address: Address, /, *, client: Client) -> str:
    data = client(
        PlaceOrder(
            {"input": {"lines": [{"book": book_id}], "shippingAddress": address}}
        )
    )

The client injects new values on each call, so a retry belongs in your transport, which sends the same body, key included, again:

def transport(body: bytes, /, *, timeout: float | None = None) -> bytes:
    retries = 2

    while True:
        try:
            response = _post(
                "/graphql", body, accept=mime_type.GRAPHQL_RESPONSE, timeout=timeout
            )
        except HTTPError as error:
            # GraphQL over HTTP sends a request error as a response with a 4xx status.
            if error.headers.get_content_type() != mime_type.GRAPHQL_RESPONSE:
                raise

            response = error
        except ConnectionError:
            # The response was lost.
            if not retries:
                raise

            retries -= 1
            continue

        with response:
            return response.read()

Colocation

Most other codegen libraries gather every operation in one generated module, often inside a single class. A feature's operations then live away from its .graphql files and the code calling them. A project split into several packages cannot have each package own its operations either. This library generates each operation as a module-level constant, so it can be colocated with both the document it comes from and the code calling it:

extensions:
  pythonCodegen:
    # …
    documentSiblingModule: "{document}_graphql"

This pattern puts get_order.graphql's operations and fragments in get_order_graphql.py. Code calling these operations imports them from there:

from bookshop.client.schema import OrderStatus
from bookshop.get_order_graphql import GetOrder

Python API

generate() does the same as the command:

class _Params(TypedDict, closed=True):
    document: DocumentNode
    schema: GraphQLSchema
    config: Config


def generate(**args: Unpack[_Params]) -> dict[PurePosixPath, bytes]:
    """Pure function returning the content of each file of the generated package by its path."""

The public API is limited to:

from graphql_codegen.config import Config as Config
from graphql_codegen.document_sibling_module import (
    DocumentSiblingModule as DocumentSiblingModule,
)
from graphql_codegen.generate import generate as generate
from graphql_codegen.package_location import PackageLocation as PackageLocation
from graphql_codegen.scalar import Codec as Codec, Scalar as Scalar
  1. Before Python 3.15, the client also needs typing_extensions, for features the standard library's typing does not have yet, so add it to your project's own dependencies. ↩ ↩2

  2. This library is partly bootstrapped: to fetch a schema from a URL, it uses a client it generated itself from _introspection.graphql. ↩

  3. A YAML config, like the quick start's, needs the yaml extra: install graphql-codegen[yaml] instead. JSON and TOML configs need nothing more. ↩

  4. Well-behaved GraphQL APIs avoid breaking changes, adding fields and deprecating old ones instead. Most breaking changes made after generation never reach the code anyway: removing or renaming a field, changing its arguments, or turning its type from an object into a leaf or the reverse invalidates the operation, so the server rejects it, and the client() call raises a RequestError. The few left would fail client-side validation as soon as the response arrives, but otherwise propagate deeper in your code, possibly unnoticed: a field becoming nullable, its leaf type changing, or it switching between a list and a single value. The more often you regenerate against the deployed schema, the sooner your type checker catches such a change. ↩

Metadata

Release files for graphql-codegen 0.1.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distribution (wheel)

Table of built distributions (wheels) for graphql-codegen 0.1.3
File Interpreter ABI Platform
graphql_codegen-0.1.3-py3-none-any.whl Python 3 none any Details

Release files / graphql_codegen-0.1.3-py3-none-any.whl

Download URL graphql_codegen-0.1.3-py3-none-any.whl
Size 94.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5a2354a8f4e370e3b481c152d2c7da6164fa635325462422bc0d8455cadbba75
BLAKE2b-256 checksum
How to use checksums
8a8380d1fae43dfeb4afbc8aee37d5860a6a27a5167ff6ef3b5931921f28b839
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.12.24 {"installer":{"name":"uv","version":"0.12.24","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"26.04","id":"resolute","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.1.3 This release

1 release file

0.1.2

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page