langgraph-checkpoint-objectstorage
A LangGraph BaseCheckpointSaver
that persists checkpoints to local filesystem, Google Cloud Storage, or AWS
S3. One class, backend picked by connection string, nothing to run beyond a
bucket (or a directory).
Table of contents
- Features
- Requirements
- Install
- Quickstart
- Examples
- Choosing a backend
- Architecture
- Runtime type checking
- Logging
- Known limitations
- Development
- Contributing
- License
Features
ObjectStorageSaver.from_conn_string(...)picks local disk, GCS, or S3 from the URI scheme. No per-backend subclasses.- Full sync and async support: every
BaseCheckpointSavermethod, both flavors (get_tuple/aget_tuple,put/aput,list/alist,put_writes/aput_writes,delete_thread/adelete_thread). - Tested against the official contract with
langgraph-checkpoint-conformanceon all three backends, not just hand-written assertions. - Runtime type checking on the public API via typeguard catches wrong-argument-type mistakes at the call site.
- Ships
py.typedfor full static type coverage under mypy/pyright. - No database or extra service required in production, just object storage.
Requirements
Python 3.11+.
Install
pip install langgraph-checkpoint-objectstorage # local filesystem only
pip install "langgraph-checkpoint-objectstorage[s3]" # + AWS S3
pip install "langgraph-checkpoint-objectstorage[gcs]" # + Google Cloud Storage
Quickstart
from langgraph.graph import END, START, StateGraph
from langgraph_checkpoint_objectstorage import ObjectStorageSaver
def increment(state: dict) -> dict:
return {"count": state["count"] + 1}
builder = StateGraph(dict)
builder.add_node("increment", increment)
builder.add_edge(START, "increment")
builder.add_edge("increment", END)
saver = ObjectStorageSaver.from_conn_string("file:///tmp/checkpoints")
graph = builder.compile(checkpointer=saver)
config = {"configurable": {"thread_id": "1"}}
result = graph.invoke({"count": 0}, config)
print(result) # {"count": 1}
# Checkpoints persisted under the thread survive process restarts --
# inspect or resume from the same thread_id at any later point:
history = list(graph.get_state_history(config))
Examples
This same quickstart, runnable under three build tools:
examples/pip: venv +pip install -r requirements.txtexamples/uv/filesystem:uv run main.pyexamples/poetry/filesystem:poetry install && poetry run python main.py
Each installs the package from this repo via a local path dependency (swap for a normal PyPI dependency once the package is published).
Further along, against object storage instead of local disk (real bucket or a local emulator, no cloud account needed): multiple independent sessions run sequentially and one is resumed later, plus the same pattern run concurrently via the async API, for both S3 and GCS:
Sequential and concurrent are separate examples rather than one combined script, see Known limitations for why.
Choosing a backend
Swap the connection string; everything else stays the same.
from langgraph_checkpoint_objectstorage import ObjectStorageSaver
# Local filesystem -- handy for development, or single-node deployments
saver = ObjectStorageSaver.from_conn_string("file:///var/lib/my-app/checkpoints")
# Google Cloud Storage
saver = ObjectStorageSaver.from_conn_string("gcs://my-bucket/checkpoints")
# AWS S3
saver = ObjectStorageSaver.from_conn_string("s3://my-bucket/checkpoints")
from_conn_string forwards extra keyword arguments to the underlying
fsspec filesystem constructor.
Useful for explicit credentials, non-default regions, or S3-compatible
endpoints (MinIO, Cloudflare R2, etc.):
saver = ObjectStorageSaver.from_conn_string(
"s3://my-bucket/checkpoints",
key="...",
secret="...",
client_kwargs={"endpoint_url": "https://minio.internal:9000"},
)
Credentials otherwise follow each backend's normal resolution: AWS's usual
chain (env vars, ~/.aws/credentials, instance/task role) for S3,
Application Default Credentials for GCS. Nothing library-specific to
configure beyond the connection string.
Architecture
Business logic (key layout, filtering, ordering, idempotency) is written
once, as async methods. The public sync API is a thin asyncio.run(...)
wrapper around that same async core, not a second implementation, so
there's a single source of truth per operation instead of sync and async
code drifting apart. An I/O bridge picks native async calls when the
backend supports them (s3fs, gcsfs) and falls back to
asyncio.to_thread when it doesn't (local disk):
flowchart TD
App["Your application<br/>(graph.invoke / ainvoke)"]
subgraph PublicAPI["Public API — BaseCheckpointSaver contract"]
Sync["put / get_tuple / list /<br/>put_writes / delete_thread"]
Async["aput / aget_tuple / alist /<br/>aput_writes / adelete_thread"]
end
Core["Async core<br/>(business logic, written once)"]
Bridge["I/O bridge<br/>_cat / _pipe / _find / _exists / _rm"]
Native["fsspec async-native<br/>(s3fs, gcsfs)"]
Threaded["asyncio.to_thread<br/>(LocalFileSystem)"]
Backend[("Local disk / S3 / GCS")]
App --> Sync
App --> Async
Sync -->|"asyncio.run(...)<br/>thin wrapper, not a<br/>second implementation"| Core
Async --> Core
Core --> Bridge
Bridge -->|backend supports async| Native --> Backend
Bridge -->|no native async| Threaded --> Backend
Each checkpoint and each write becomes its own object: no read-modify-write
on existing keys, so concurrent writers on different threads never race,
and a put or put_writes call is always a single write:
flowchart TD
Root["{root}"] --> Thread["{thread_id}/"]
Thread --> NS["{checkpoint_ns}/"]
NS --> CkptDir["checkpoints/"]
NS --> WriteDir["writes/"]
CkptDir --> Ckpt["{checkpoint_id}.msgpack<br/>checkpoint + metadata + parent_checkpoint_id"]
WriteDir --> WCkpt["{checkpoint_id}/"]
WCkpt --> WTask["{task_id}/"]
WTask --> WIdx["{idx}.msgpack<br/>task_id + idx + channel + value"]
Key technical decisions this reflects:
checkpoint_idis LangGraph's own uuid6, already time-sortable, solist()'s newest-first ordering falls out of a plain key sort, with no secondary index to keep in sync.- Checkpoints and metadata are serialized through LangGraph's own
JsonPlusSerializerand treated as opaque: never hand-extracted field by field, since real LangGraph objects carry fields their own TypedDicts don't declare (see Runtime type checking). put_writes' overwrite-vs-ignore split ("first write wins" for regular channels, always-replace for the control channelsERROR/SCHEDULED/INTERRUPT/RESUME) mirrors the official sqlite/postgres savers exactly.list(filter=...)is deliberately client-side, not an oversight: object storage has no query engine to push a filter into. See Known limitations.
Runtime type checking
Public methods are decorated with typeguard
and raise typeguard.TypeCheckError on a call with the wrong argument
types (e.g. a non-string thread_id, a writes argument that isn't a
sequence of (channel, value) pairs). This catches integration mistakes
at the call site instead of letting them corrupt stored data silently.
config, checkpoint, and metadata arguments are intentionally not
strictly checked against LangGraph's RunnableConfig/Checkpoint/
CheckpointMetadata TypedDicts: real LangGraph objects don't match those
TypedDicts exactly (a real RunnableConfig's metadata field is a
collections.ChainMap, not a plain dict; real checkpoints carry fields
like the legacy pending_sends key that isn't declared at all), and strict
checking would reject every real invocation. Same reasoning applies to
get_tuple/list's return value, which isn't runtime-checked for the same
reason. Other arguments (thread_id, task_id, writes, limit, ...)
are checked normally.
Logging
Uses standard logging under the logger name
langgraph_checkpoint_objectstorage. No handlers are configured, so it
stays silent until your application's logging config says otherwise.
For quick debugging, set LANGGRAPH_CHECKPOINT_OBJECTSTORAGE_LOG_LEVEL=DEBUG
before constructing an ObjectStorageSaver. It emits one DEBUG line per
storage read/write/list with the key or prefix touched.
Known limitations
list(filter=...)is client-side: every checkpoint in the thread/namespace is fetched and filtered in Python, since object storage has no query engine to push the filter into. Fine for typical thread histories (dozens to low hundreds of checkpoints); a very long-running thread'slistcalls will get proportionally slower.- No garbage collection or retention policy. Old checkpoints accumulate
until you call
delete_thread, or you set up bucket lifecycle rules yourself. - Two writers on the same
thread_idwriting concurrently can race, at the same guarantee level as the official sqlite saver (last write "wins" by whichever checkpoint_id sorts last, not by wall-clock order under clock skew). Concurrent writers on different threads never race. - The sync
list()eagerly collects all matching checkpoints before yielding the first one (it wraps the async implementation viaasyncio.run, which can't stream lazily). Usealist()from async code if you need true streaming. - Don't mix sync and async calls on the same
ObjectStorageSaverinstance against S3 or GCS. Sync calls run on a persistent background loop the underlying filesystem maintains; async calls run on whichever loop the caller provides. The aiohttp session those backends use can only belong to one loop at a time, so alternating between the two on one instance breaks withRuntimeError: ... attached to a different loop. Build a separate saver instance per usage style instead (seeexamples/uv/s3vsexamples/uv/s3-async, orexamples/poetry/gcsvsexamples/poetry/gcs-async). Local filesystem isn't affected: it has no persistent session to misalign.
Development
make install # sync deps into an isolated .venv (installs uv if missing)
make lint # black --check, pyflakes, bandit, vulture -- runs before test/build
make format # apply black formatting in place
make test # full suite -- docker-compose integration tests auto-skip if not up
make test-unit # tests/unit only -- no external services, fast
make test-integration # tests/integration -- starts docker-compose emulators first
test/test-unit/test-integration/build all run lint first, so a
formatting or static-analysis failure blocks the run rather than
surfacing only after tests pass. bandit/vulture are scoped to src/
only: bandit flags every pytest assert and the tests' intentionally
fake credentials, and vulture can't see that ObjectStorageSaver's
public methods are called by library consumers rather than this codebase
(see vulture_whitelist.py).
Or without make, directly via uv:
uv sync --all-extras --group dev
uv run pytest
Test layout: tests/unit/ exercises internal modules (key layout,
serialization, the saver's core logic, logging, type checking) with no
external service. tests/integration/ validates the full
BaseCheckpointSaver contract via
langgraph-checkpoint-conformance
against real backends: local disk, an in-process moto server for S3, and
(via docker-compose.yaml) fake-gcs-server/moto-server containers for a
full local GCS/S3 round-trip with no cloud account needed. A gated test
against a real GCS bucket runs only when GCS_TEST_BUCKET is set.
make compose-up # start local S3/GCS emulators
make test-integration
make compose-down # stop them when done
Contributing
Issues and PRs welcome. Before opening one, make test should pass
(make test-integration too, if your change touches backend I/O). CI runs
lint, unit tests (Python 3.11/3.12/3.13), and integration tests against
docker-compose emulators on every push and PR.
Add an entry under ## [Unreleased] in CHANGELOG.md for
any user-facing change. On merge to main, CI moves that section into a
new dated version automatically (see the release job in
.github/workflows/ci.yml). An empty [Unreleased] just gets a generic
placeholder line instead, so it's worth taking the extra minute.
Releasing
The release job only bumps pyproject.toml and CHANGELOG.md and
pushes that commit to main -- it doesn't tag or publish anything.
Publishing to PyPI is a manual step, since it's the one part of this
pipeline that isn't reversible:
git pull origin main
git tag -a vX.Y.Z -m vX.Y.Z
git push origin vX.Y.Z
Pushing that tag triggers .github/workflows/publish.yml, which builds
the package and publishes it to PyPI via trusted publishing (no token
needed). Create the GitHub release from the same tag, e.g.:
gh release create vX.Y.Z --title vX.Y.Z --generate-notes
License
MIT.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file langgraph_checkpoint_objectstorage-0.1.7.tar.gz.
File metadata
- Download URL: langgraph_checkpoint_objectstorage-0.1.7.tar.gz
- Upload date:
- Size: 854.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7ea0a85c4e0648e15c364aa2be59ea14482455b8a760edb2cea691359ee5589b
|
|
| MD5 |
8beb5258152ee84e11073b0fc76188ee
|
|
| BLAKE2b-256 |
b25379283e1a65569139c6ff1c96a241f794851620669ef1978da063d214230b
|
Provenance
The following attestation bundles were made for langgraph_checkpoint_objectstorage-0.1.7.tar.gz:
Publisher:
publish.yml on sergiommarcial/langgraph-objectstorage-checkpoint
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
langgraph_checkpoint_objectstorage-0.1.7.tar.gz -
Subject digest:
7ea0a85c4e0648e15c364aa2be59ea14482455b8a760edb2cea691359ee5589b - Sigstore transparency entry: 2655472069
- Sigstore integration time:
-
Permalink:
sergiommarcial/langgraph-objectstorage-checkpoint@5095de898833a5c732d4a06b0787c699e38b3254 -
Branch / Tag:
refs/tags/v0.1.7 - Owner: https://github.com/sergiommarcial
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@5095de898833a5c732d4a06b0787c699e38b3254 -
Trigger Event:
push
-
Statement type:
File details
Details for the file langgraph_checkpoint_objectstorage-0.1.7-py3-none-any.whl.
File metadata
- Download URL: langgraph_checkpoint_objectstorage-0.1.7-py3-none-any.whl
- Upload date:
- Size: 14.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1b463a2295b98989001881b4c16a619650e57e7fb10ed07b78ac0ecc8c280db7
|
|
| MD5 |
a88c5071ae5f7b20e08d005afe2b3c3b
|
|
| BLAKE2b-256 |
9c88042a857de1174baf43686600cbea8ae0b156c5606dfe2d8fd6861f6452a2
|
Provenance
The following attestation bundles were made for langgraph_checkpoint_objectstorage-0.1.7-py3-none-any.whl:
Publisher:
publish.yml on sergiommarcial/langgraph-objectstorage-checkpoint
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
langgraph_checkpoint_objectstorage-0.1.7-py3-none-any.whl -
Subject digest:
1b463a2295b98989001881b4c16a619650e57e7fb10ed07b78ac0ecc8c280db7 - Sigstore transparency entry: 2655472080
- Sigstore integration time:
-
Permalink:
sergiommarcial/langgraph-objectstorage-checkpoint@5095de898833a5c732d4a06b0787c699e38b3254 -
Branch / Tag:
refs/tags/v0.1.7 - Owner: https://github.com/sergiommarcial
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@5095de898833a5c732d4a06b0787c699e38b3254 -
Trigger Event:
push
-
Statement type: