Reusable Google GenAI/Vertex AI Client Module
Project description
GreyCloud
A comprehensive, configurable Python package for interacting with Google's Vertex AI and GenAI services (Gemini), including authentication, content generation, batch processing, token counting, and file management.
GreyCloud wraps the lower-level google-genai client with:
- Unified authentication (API key or OAuth + optional service account impersonation)
- Resilient content generation with automatic retry and re-authentication
- Config-driven client setup via a single
GreyCloudConfigdataclass - Context caching for 75-90% cost savings on repeated queries
- Optional Vertex AI Search tools for retrieval-augmented generation
- Batch helpers for large offline jobs and GCS integration
1. What GreyCloud Does
GreyCloud provides four main building blocks:
GreyCloudConfig– configuration object populated from environment variables or codeGreyCloudClient– high-level client for content generation, streaming, token counting, and retriesGreyCloudCache– context caching for cost-efficient repeated queries on the same contentGreyCloudBatch– helper for batch jobs and GCS-backed workflows
High-level capabilities:
- Content generation (streaming and non-streaming) with per-request overrides
- Automatic retry with exponential backoff and authentication-aware recovery
- Context caching with 75-90% cost savings on cached input tokens
- Token counting with graceful approximation fallback
- Vertex AI Search integration via a simple flag and datastore string
- Batch processing to upload files, create jobs, monitor, and download results
2. Why Use GreyCloud Instead of google-genai Directly?
Using google-genai directly is flexible but verbose. GreyCloud focuses on developer ergonomics and resilience:
-
Unified auth helper
- One function (
create_client/GreyCloudClient) that:- Uses Application Default Credentials when available
- Optionally impersonates a service account when
sa_emailis set - Falls back to
gcloud auth print-access-tokenwhen needed - Supports API key authentication via a simple config flag
- Clear error messages that point to:
gcloud auth application-default login- IAM role requirements for impersonation
- One function (
-
Config normalization
- A single dataclass (
GreyCloudConfig) encapsulates:- Project, location, endpoint, model
- Auth choices (API key vs OAuth + SA impersonation)
- Generation parameters (temperature, top_p, max_output_tokens, seed)
- Safety settings
- Thinking configuration
- Vertex AI Search datastore
- Batch/GCS bucket settings
- A single dataclass (
-
Resilient generation
GreyCloudClient.generate_with_retry(...):- Detects auth-related vs transient errors
- Performs exponential backoff with jitter
- Attempts re-authentication when appropriate (for OAuth-based flows)
- Re-creates the underlying
genai.Clientas needed
-
Tools & Search wiring
- Vertex AI Search is turned on with:
use_vertex_ai_search=Truevertex_ai_search_datastore="projects/.../dataStores/...".
- GreyCloud constructs the appropriate
types.Tooland wires it into calls.
- Vertex AI Search is turned on with:
-
Batch utilities
GreyCloudBatchwraps the more verbose raw batch APIs:- Handles JSONL creation
- Manages GCS paths and result locations
- Tries multiple model naming formats (
publishers/google/models/...vs short name)
-
Sync vs async
-
Same config (
GreyCloudConfig) and same method names for sync and async. -
Use
GreyCloudClientfor synchronous code; useGreyCloudAsyncClientfor async/rate-limited usage. -
The async client applies RPM, TPM, and concurrency limits via
VertexRateLimiter; use it when you need to stay within quotas (e.g. in web backends). -
API mapping:
Sync ( GreyCloudClient)Async ( GreyCloudAsyncClient)generate_content(...)await generate_content(...)generate_content_stream(...)async for x in generate_content_stream(...)generate_with_retry(..., streaming=False)await generate_with_retry(...)generate_with_retry(..., streaming=True)async for x in (await generate_with_retry(..., streaming=True))count_tokens(...)await count_tokens(...) -
For advanced use the underlying
genai.Clientis available as.clienton both clients; rate-limited generation should go through the client’s methods, not rawclient.aio.models.*.
-
3. Installation
Basic Installation
pip install greycloud
Development Installation
git clone https://github.com/jbff/greycloud.git
cd greycloud
pip install -e ".[dev]"
4. Quick Start: Basic Client and Single Call
from greycloud import GreyCloudConfig, GreyCloudClient
from google.genai import types
# Create configuration (override defaults as needed)
config = GreyCloudConfig(
project_id="your-project-id",
location="us-central1",
# Default model is a Gemini 3 flash model; you can override if desired.
model="gemini-3-flash-preview",
)
# Create client
client = GreyCloudClient(config)
# Generate content
contents = [
types.Content(
role="user",
parts=[types.Part.from_text(text="Hello, how are you?")]
)
]
response = client.generate_content(contents)
print(response.text)
5. Detailed Examples
5.1 Creating a Client from Environment Only
Environment:
export PROJECT_ID="your-project-id"
export LOCATION="us-central1"
Code:
from greycloud import GreyCloudClient
from google.genai import types
client = GreyCloudClient() # GreyCloudConfig is created from env
contents = [
types.Content(
role="user",
parts=[types.Part.from_text(text="Summarize the benefits of Vertex AI.")]
)
]
response = client.generate_content(contents)
print(response.text)
5.2 Per-Request Overrides
response = client.generate_content(
contents,
temperature=0.7,
max_output_tokens=1024,
system_instruction="You are a concise technical assistant.",
)
5.3 Streaming Generation
By default, streaming yields plain text strings representing the generated content chunks:
for chunk in client.generate_content_stream(contents):
print(chunk, end="", flush=True)
If you need access to candidate metadata, usage metrics, or safety ratings, you can pass return_chunks=True to yield the raw GenerateContentResponse chunk objects instead of strings:
for chunk in client.generate_content_stream(contents, return_chunks=True):
# chunk is a google.genai.types.GenerateContentResponse object
print(chunk.text, end="", flush=True)
5.4 Automatic Retry & Auth Recovery
from google.genai import types
contents = [
types.Content(
role="user",
parts=[types.Part.from_text(text="Give me a short creative story about a robot therapist.")]
)
]
response = client.generate_with_retry(
contents,
max_retries=5,
streaming=False,
)
print(response.text)
For streaming with retry:
for chunk in client.generate_with_retry(
contents,
max_retries=5,
streaming=True,
):
print(chunk, end="", flush=True)
To stream raw response chunk objects with retry logic, pass return_chunks=True:
for chunk in client.generate_with_retry(
contents,
max_retries=5,
streaming=True,
return_chunks=True,
):
print(chunk.text, end="", flush=True)
5.5 Token Counting with Fallback
from google.genai import types
contents = [
types.Content(
role="user",
parts=[types.Part.from_text(text="Count the tokens in this example message.")]
)
]
token_count = client.count_tokens(
contents,
system_instruction="You are a helpful assistant.",
)
print(f"Total tokens: {token_count}")
If the underlying API is unavailable, GreyCloud falls back to an approximate character-based count.
5.6 Context Caching for Cost Savings
Context caching allows you to cache large content (documents, code, media) and reuse it across multiple requests without re-sending tokens each time. This provides significant cost savings:
- Cached token discount: 75-90% off input token costs (depending on model)
- Storage cost: $1.00 per million tokens per hour (prorated by minute)
from greycloud import GreyCloudConfig, GreyCloudCache
config = GreyCloudConfig(project_id="your-project-id")
cache_client = GreyCloudCache(config)
# Cache a large document (must meet minimum token threshold: 1,024-4,096 tokens)
large_document = "..." # Your large content here
cache = cache_client.create_cache_from_text(
text=large_document,
display_name="my-document-cache",
system_instruction="You are a helpful document analyst.",
ttl_seconds=3600, # 1 hour
)
print(f"Cache created: {cache.name}")
print(f"Cached tokens: {cache.usage_metadata.total_token_count}")
# Query the cache multiple times (each query uses cached tokens at discounted rate)
questions = [
"Summarize the main points",
"What are the key findings?",
"List any recommendations",
]
for question in questions:
response = cache_client.generate_with_cache(
cache_name=cache.name,
prompt=question,
)
print(f"Q: {question}")
print(f"A: {response.text}\n")
# IMPORTANT: Delete cache when done to stop storage charges
cache_client.delete_cache(cache.name)
You can also cache GCS files:
cache = cache_client.create_cache_from_files(
file_uris=[
"gs://your-bucket/document1.pdf",
"gs://your-bucket/document2.txt",
],
display_name="multi-file-cache",
ttl_seconds=7200, # 2 hours
)
Cache management:
# List all caches
for cached_content in cache_client.list_caches():
info = cache_client.get_cache_info(cached_content)
print(f"{info['name']}: {info.get('total_token_count', 'N/A')} tokens")
# Extend cache TTL before it expires
cache_client.update_cache_ttl(cache.name, ttl_seconds=7200)
# Delete all caches with a specific display name
cache_client.delete_all_caches(display_name_filter="my-document-cache")
Note: Context caching is a paid feature and not available in the free tier.
5.7 Using Cached Content with GreyCloudClient
You can also use cached content directly with GreyCloudClient by passing the cached_content parameter:
from greycloud import GreyCloudConfig, GreyCloudClient, GreyCloudCache
from google.genai import types
config = GreyCloudConfig(project_id="your-project-id")
# Create cache
cache_client = GreyCloudCache(config)
cache = cache_client.create_cache_from_text(
text=large_document,
display_name="my-cache",
ttl_seconds=3600,
)
# Use with GreyCloudClient
client = GreyCloudClient(config)
contents = [
types.Content(
role="user",
parts=[types.Part.from_text(text="Summarize the document")]
)
]
response = client.generate_content(
contents,
cached_content=cache.name, # Use the cache
)
# Streaming also works with cached content
for chunk in client.generate_content_stream(
contents,
cached_content=cache.name,
):
print(chunk, end="", flush=True)
# Clean up
cache_client.delete_cache(cache.name)
5.8 Vertex AI Search as a Tool
from greycloud import GreyCloudConfig, GreyCloudClient
from google.genai import types
config = GreyCloudConfig(
project_id="your-project-id",
location="us-central1",
use_vertex_ai_search=True,
vertex_ai_search_datastore=(
"projects/PROJECT_ID/locations/LOCATION/"
"collections/default_collection/dataStores/DATASTORE_ID"
),
)
client = GreyCloudClient(config)
contents = [
types.Content(
role="user",
parts=[types.Part.from_text(text="Using the knowledge base, explain the diagnostic steps for adult ASD.")]
)
]
response = client.generate_content(contents)
print(response.text)
5.9 Batch Processing with GCS
Batch jobs use a GCS bucket for request input and result output. Set batch_gcs_bucket (and optionally gcs_bucket for general uploads). The batch API expects JSONL input following the Vertex AI REST GenerateContentRequest schema: one line per request, each line a JSON object with a request key containing model, contents, and optional generationConfig, systemInstruction, and safetySettings (all camelCase). Note: InlinedRequest.metadata is not forwarded to batch JSONL because Vertex rejects numeric string values in proto label fields — use prompt-embedded ID tags (e.g. [SLICE_ID:...]) for request matching. Results are written by Vertex to predictions.jsonl under the job's destination prefix; download_batch_results finds and downloads that file.
from greycloud import GreyCloudConfig, GreyCloudBatch
from google.genai import types
import json
config = GreyCloudConfig(
project_id="your-project-id",
batch_gcs_bucket="your-project-batch-jobs", # Must exist; used for batch I/O
)
batch = GreyCloudBatch(config)
# Upload a couple of JSON docs (use same bucket via bucket_name)
files = [
{"name": "data1.json", "content": json.dumps({"key": "value"})},
{"name": "data2.json", "content": json.dumps({"key2": "value2"})},
]
file_uris = batch.upload_files_to_gcs(files, bucket_name=config.batch_gcs_bucket)
batch_requests = []
for filename, gcs_uri in file_uris.items():
batch_requests.append(
types.InlinedRequest(
model=config.model,
contents=[
{
"role": "user",
"parts": [
{"text": f"Analyze {filename}: "},
{"file_data": {"file_uri": gcs_uri, "mime_type": "application/json"}},
],
}
],
config=types.GenerateContentConfig(
temperature=0.2,
max_output_tokens=65535,
),
)
)
batch_job = batch.create_batch_job(batch_requests)
batch_job = batch.monitor_batch_job(batch_job)
output_file = batch.download_batch_results(batch_job, "results.jsonl")
print(f"Batch results saved to: {output_file}")
5.10 Custom Auth (Advanced)
from greycloud.auth import create_client
client = create_client(
project_id="your-project-id",
location="us-central1",
sa_email="service-account@project.iam.gserviceaccount.com", # Optional
use_api_key=False,
)
Documentation
All usage and configuration details are documented in this README.md. For additional examples, see:
examples/simple.py– minimal content-generation script.examples/caching.py– context caching for cost-efficient repeated queries.
Requirements
- Python 3.10+
- Google Cloud Project with Vertex AI enabled
google-genaipackage (installed withgreycloud)google-authpackage (installed withgreycloud, for OAuth)google-cloud-storagepackage (installed withgreycloud; only needed if you use batch/GCS helpers)
Testing
Run the test suite:
pytest
Run with coverage:
pytest --cov=greycloud --cov-report=html
License
MIT License (see LICENSE file).
Contributing
Contributions are welcome! Please feel free to submit a Pull Request.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file greycloud-0.3.8.tar.gz.
File metadata
- Download URL: greycloud-0.3.8.tar.gz
- Upload date:
- Size: 50.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.11.19 {"installer":{"name":"uv","version":"0.11.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Fedora Linux","version":"44","id":"","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
683f7428ad99075a77dd6c6635155df62241af55bc84f87e8771deabad4c748c
|
|
| MD5 |
e77994c2402f15983d105085123c6c00
|
|
| BLAKE2b-256 |
bbc9ec51ccab1b7f0d5b11a1dbeba7507fcb1f4d378a0ca0ac022519ad221bfc
|
File details
Details for the file greycloud-0.3.8-py3-none-any.whl.
File metadata
- Download URL: greycloud-0.3.8-py3-none-any.whl
- Upload date:
- Size: 32.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.11.19 {"installer":{"name":"uv","version":"0.11.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Fedora Linux","version":"44","id":"","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
02ea451b785759b9dffd0a927fab7ba2c834c86b4d157a9303b283ff9ac051d3
|
|
| MD5 |
3a0cffff7a6fe9b08cfc25517d6d5847
|
|
| BLAKE2b-256 |
0fd18dcbf660fcf208c13493f3c1e169b6c4d636c051d6f76c34108ae39d5ee8
|