gc-batch
A command-line interface and Python client for managing Google Cloud Batch jobs, with support for job creation, monitoring, logging, GCS bucket mounts, local SSD storage, and filtering by labels.
Installation
pip install gc-batch
Or, for development, using uv:
git clone https://github.com/empirico-oss/gc-batch.git
cd gc-batch
uv sync
Requires Python 3.10 or newer, and credentials for a Google Cloud project with the
Batch API enabled (gcloud auth application-default login).
Quick start
# Create a job
gc-batch --project-id my-project create \
--job-name my-first-job \
--docker-image gcr.io/my-project/my-image:latest \
--command "python /app/main.py"
# List jobs, or just your own
gc-batch --project-id my-project list-jobs
gc-batch --project-id my-project list-my-jobs
# Check status and read logs
gc-batch --project-id my-project status --job-name my-first-job-1234567890
gc-batch --project-id my-project logs --job-name my-first-job-1234567890
The project is resolved from --project-id, then $GOOGLE_PROJECT, then the
default_project_id setting. If none is set, the command exits with an error
rather than guessing.
Use as a Python library
The same functionality is importable, so scripts, notebooks, and orchestration code can submit and monitor jobs without shelling out to the CLI:
import time
from gc_batch import BatchClientConfig, BatchJobConfig, GCBatchClient, JobRequest
from gc_batch.utils import is_job_finished
client = GCBatchClient(BatchClientConfig(project_id="my-project", location="us-central1"))
job = client.create_job(
JobRequest(
job_name="my-first-job",
docker_image="gcr.io/my-project/my-image:latest",
command="python /app/main.py",
config=BatchJobConfig(
machine_type="n2-standard-4",
boot_disk_type="pd-balanced",
input_bucket="my-bucket/datasets",
output_bucket="my-bucket/results",
),
labels={"team": "data-science"},
)
)
# create_job timestamps the name, so read the submitted one off the job
short_name = job.name.split("/")[-1]
while not is_job_finished(job := client.get_job(short_name)):
time.sleep(30)
print(job.status.state.name)
if job.status.state.name == "FAILED":
print(client.get_failure_message(job))
print(client.get_cloud_logging_url(job, severity="ERROR"))
Listing, cancelling, and log retrieval are all available too:
jobs = client.list_jobs(labels={"team": "data-science"}, status="RUNNING")
client.cancel_job(job.name) # full resource path, not the short name
# Logs, from Cloud Logging or from the bucket when the job used logs_bucket
client.batch_logging.print_logs_for_job(job)
client.gcs_logging.print_logs_for_job(job)
See the Python API guide for job profiles, mounts, local SSD, settings, fan-out, and error handling.
Configuration
Values that differ between deployments — the job-name prefix, the
created-using label, and named job profiles — are resolved in this order,
highest precedence first:
- explicit arguments
GC_BATCH_*environment variables- a TOML config file (
--config-file,$GC_BATCH_CONFIG_FILE,./gc-batch.toml, or$XDG_CONFIG_HOME/gc-batch/config.toml) - neutral defaults
Restricted VPC environments
gc-batch ships a built-in all-of-us job profile for the
All of Us Researcher Workbench, which runs jobs inside
a restricted VPC:
gc-batch --project-id my-project create \
--job-profile all-of-us \
--job-name my-job \
--docker-image gcr.io/my-project/my-image:latest \
--command "python /app/main.py"
You can define additional profiles in the config file. See the Configuration guide.
Documentation
Full documentation: https://empirico-oss.github.io/gc-batch/stable/
- Quick Start
- Commands Reference
- Python API
- Input/Output Mounts
- Local SSD Storage
- Label Management
- Configuration
- Troubleshooting
Contributing
See CONTRIBUTING.md. Security issues should be reported per SECURITY.md.
License
MIT — see LICENSE.
Release files for gc-batch 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| gc_batch-0.2.0.tar.gz | 65.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| gc_batch-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 115.4 kB
Release files / gc_batch-0.2.0.tar.gz
| Download URL | gc_batch-0.2.0.tar.gz |
|---|---|
| Size | 65.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
eb9b9f2cb4417479554113a3c6ccb781f8c0460d6b842e2841b01da9d91ffe64
|
|
BLAKE2b-256 checksum How to use checksums |
c2397ddcc265386fa1a840d3b8294ba1e4ae6aed422d9d4a2d87deea8eb2f821
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency logRelease files / gc_batch-0.2.0-py3-none-any.whl
| Download URL | gc_batch-0.2.0-py3-none-any.whl |
|---|---|
| Size | 50.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7cdcea51c6f5f127671102cba83191e623d3ea4ace00245d2668753b5ae87bd8
|
|
BLAKE2b-256 checksum How to use checksums |
bd9c6200e8059214cf7faf742826abd405eb71144738953029511ddc352dc6ee
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency log