Skip to main content

lmux-aws-bedrock

AWS Bedrock provider for lmux. Talks to the Bedrock Converse and InvokeModel REST endpoints directly over httpx, using boto3 only to resolve AWS credentials for request signing.

Supports chat completions, streaming, and embeddings.

Part of the lmux ecosystem: standardized interface, cost tracking on every response, and registry-based routing across providers.

Optional Extras

  • lmux-aws-bedrock[async]: aiobotocore for async AWS credential resolution in the auth providers. Not required for achat/aembed/achat_stream themselves.

Auth

Two authentication modes are supported, resolved on the first request:

  1. Bedrock API key (simplest): set AWS_BEARER_TOKEN_BEDROCK and each request is sent with an Authorization: Bearer <token> header — no signing required.
  2. SigV4 (fallback): otherwise AWS credentials are resolved through boto3's default credential chain (env vars, AWS config, instance metadata) and each request is SigV4-signed. No extra setup needed if your AWS credentials are already configured.
from lmux_aws_bedrock import BedrockProvider

provider = BedrockProvider()

# Or specify a region
provider = BedrockProvider(region="us-east-1")

For explicit session configuration:

from lmux_aws_bedrock import BedrockSessionAuthProvider

provider = BedrockProvider(auth=BedrockSessionAuthProvider(profile_name="my-profile"))

Usage

Chat

from lmux import UserMessage

response = provider.chat("anthropic.claude-sonnet-4-20250514-v1:0", [UserMessage(content="Hello")])
print(response.content)
print(response.cost)

Streaming

for chunk in provider.chat_stream("anthropic.claude-sonnet-4-20250514-v1:0", [UserMessage(content="Hello")]):
    if chunk.delta:
        print(chunk.delta, end="")

Tool continuations

Bedrock reasoning blocks carry signatures that must be returned unmodified when a tool result continues the assistant turn. lmux-aws-bedrock preserves the native ordered Converse blocks in response.continuation; use to_assistant_message() to keep them with the normalized response:

from lmux import ToolMessage

response = provider.chat(model, messages, tools=tools, reasoning_effort="high")
messages.append(response.to_assistant_message())
messages.append(ToolMessage(content=tool_result, tool_call_id=response.tool_calls[0].id))

A matching Converse continuation is replayed exactly. Other providers ignore it and use the normalized content and tool calls.

Embeddings

response = provider.embed("amazon.titan-embed-text-v2:0", "Hello")
print(response.embeddings)

Async

All methods have async variants: achat, achat_stream, aembed. These run over httpx's async client; credentials are resolved synchronously (via boto3) even on the async path.

Bedrock also supports lmux response_format, mapped to Converse outputConfig.textFormat.

Registry

Use with the lmux registry to route across multiple providers:

from lmux import Registry

registry = Registry()
registry.register("bedrock", provider)
response = registry.chat("bedrock/anthropic.claude-sonnet-4-20250514-v1:0", messages)

Provider Params

from lmux_aws_bedrock import BedrockParams, GuardrailConfig

response = provider.chat(
    "anthropic.claude-sonnet-4-20250514-v1:0",
    messages,
    provider_params=BedrockParams(
        guardrail_config=GuardrailConfig(
            guardrail_identifier="my-guardrail",
            guardrail_version="1",
        ),
    ),
)
Parameter Type Description
guardrail_config GuardrailConfig Bedrock guardrail to apply
additional_model_request_fields dict Extra fields passed to the model
additional_model_response_field_paths list[str] Extra response fields to return
pricing_as_of datetime.date Override the date used for dated pricing; defaults to the current date

For Claude 4.5 and older, reasoning_effort maps to a manual thinking budget capped below maxTokens. When max_tokens is omitted, lmux sends maxTokens: 4096, so medium and high effort both use a 4095-token thinking budget. Pass a larger max_tokens to use their full mapped budgets.

An integer manual-thinking budget in additional_model_request_fields raises the default maxTokens when needed. An explicit max_tokens is preserved instead, along with the provider-specific fields. Ensure those explicit values are compatible with the deployed model; manual thinking normally requires budget_tokens < maxTokens, except when interleaved thinking applies.

Prompt Caching

Place CachePointContent parts in UserMessage content to emit Converse cachePoint blocks marking the end of a stable prompt prefix. A cache point with no preceding block in its message is placed after whatever came before it (the prior message, or the system blocks). Markers with nothing cacheable before them are dropped, and adjacent duplicates are coalesced — the first marker wins.

from lmux import CachePointContent, TextContent, UserMessage

messages = [
    UserMessage(content=[TextContent(text=big_stable_context), CachePointContent()]),
    UserMessage(content="What changed since yesterday?"),
]

Cache points are emitted for whatever model the request targets; models without prompt-caching support reject them at request validation. Cache reads/writes are reported on response.usage (cache_read_tokens, cache_creation_tokens, and the per-TTL cache_creation_tokens_by_ttl breakdown from cacheDetails) and priced into response.cost, including per-TTL write rates where the pricing data carries them.

Constructor Options

BedrockProvider(
    auth=...,          # AuthProvider, default: BedrockEnvAuthProvider()
    region=...,        # AWS region
    endpoint_url=...,  # Custom endpoint URL (overrides region/FIPS host selection)
    use_fips=...,      # bool, default False: target the FIPS 140-3 endpoint (bedrock-runtime-fips.<region>.amazonaws.com)
    timeout=...,       # request timeout in seconds
    max_retries=...,   # retry count for transient failures
    default_headers=...,  # Optional headers included with every request
    transport=...,        # Optional httpx.BaseTransport for the sync client (proxies, testing)
    async_transport=...,  # Optional httpx.AsyncBaseTransport for the async client
)

default_headers is useful for gateway authentication, tracing, and routing. Bedrock-managed authentication and content-type headers take precedence over caller values, case-insensitively. With SigV4 authentication, custom headers are included in the request signature.

Metadata

Release files for lmux-aws-bedrock 0.13.6

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for lmux-aws-bedrock 0.13.6
File Size Uploaded
lmux_aws_bedrock-0.13.6.tar.gz 28.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for lmux-aws-bedrock 0.13.6
File Interpreter ABI Platform
lmux_aws_bedrock-0.13.6-py3-none-any.whl Python 3 none any Details

Total release size: 59.2 kB

Release files / lmux_aws_bedrock-0.13.6.tar.gz

Download URL lmux_aws_bedrock-0.13.6.tar.gz
Size 28.1 kB
Tags Source
SHA-256 checksum
How to use checksums
8fb7d8e1f9497bb1a410d2dab002eae7554520138d47cfef953950d919902545
BLAKE2b-256 checksum
How to use checksums
cbd08f889b10ca84653431cf5c04fd03dae7cda741952c519b00f15a25bd61a8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.

Transparency log

Release files / lmux_aws_bedrock-0.13.6-py3-none-any.whl

Download URL lmux_aws_bedrock-0.13.6-py3-none-any.whl
Size 31.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
80d575b0eed6d4c810a23924132b1d3d8ec60b82f03a9280ab096c770f77deb1
BLAKE2b-256 checksum
How to use checksums
222edcb4895d94d319e6f1b372c494231c145b17af98cf73f9efbc12d47c005b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.13.6 This release

2 release files

0.13.5

2 release files

0.13.4

2 release files

0.13.1

2 release files

0.13.0

2 release files

0.12.1

2 release files

0.12.0

2 release files

0.11.0

2 release files

0.10.0

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page