Skip to main content

lmux-aws-bedrock

AWS Bedrock provider for lmux. Talks to the Bedrock Converse and InvokeModel REST endpoints directly over httpx, using boto3 only to resolve AWS credentials for request signing.

Supports chat completions, streaming, and embeddings.

Part of the lmux ecosystem: standardized interface, cost tracking on every response, and registry-based routing across providers.

Optional Extras

  • lmux-aws-bedrock[async]: aiobotocore for async AWS credential resolution in the auth providers. Not required for achat/aembed/achat_stream themselves.

Auth

Two authentication modes are supported, resolved on the first request:

  1. Bedrock API key (simplest): set AWS_BEARER_TOKEN_BEDROCK and each request is sent with an Authorization: Bearer <token> header — no signing required.
  2. SigV4 (fallback): otherwise AWS credentials are resolved through boto3's default credential chain (env vars, AWS config, instance metadata) and each request is SigV4-signed. No extra setup needed if your AWS credentials are already configured.
from lmux_aws_bedrock import BedrockProvider

provider = BedrockProvider()

# Or specify a region
provider = BedrockProvider(region="us-east-1")

For explicit session configuration:

from lmux_aws_bedrock import BedrockSessionAuthProvider

provider = BedrockProvider(auth=BedrockSessionAuthProvider(profile_name="my-profile"))

Usage

Chat

from lmux import UserMessage

response = provider.chat("anthropic.claude-sonnet-4-20250514-v1:0", [UserMessage(content="Hello")])
print(response.content)
print(response.cost)

Streaming

for chunk in provider.chat_stream("anthropic.claude-sonnet-4-20250514-v1:0", [UserMessage(content="Hello")]):
    if chunk.delta:
        print(chunk.delta, end="")

Tool continuations

Bedrock reasoning blocks carry signatures that must be returned unmodified when a tool result continues the assistant turn. lmux-aws-bedrock preserves the native ordered Converse blocks in response.continuation; use to_assistant_message() to keep them with the normalized response:

from lmux import ToolMessage

response = provider.chat(model, messages, tools=tools, reasoning_effort="high")
messages.append(response.to_assistant_message())
messages.append(ToolMessage(content=tool_result, tool_call_id=response.tool_calls[0].id))

A matching Converse continuation is replayed exactly. Other providers ignore it and use the normalized content and tool calls.

Embeddings

response = provider.embed("amazon.titan-embed-text-v2:0", "Hello")
print(response.embeddings)

Async

All methods have async variants: achat, achat_stream, aembed. These run over httpx's async client; credentials are resolved synchronously (via boto3) even on the async path.

Bedrock also supports lmux response_format, mapped to Converse outputConfig.textFormat.

Registry

Use with the lmux registry to route across multiple providers:

from lmux import Registry

registry = Registry()
registry.register("bedrock", provider)
response = registry.chat("bedrock/anthropic.claude-sonnet-4-20250514-v1:0", messages)

Provider Params

from lmux_aws_bedrock import BedrockParams, GuardrailConfig

response = provider.chat(
    "anthropic.claude-sonnet-4-20250514-v1:0",
    messages,
    provider_params=BedrockParams(
        guardrail_config=GuardrailConfig(
            guardrail_identifier="my-guardrail",
            guardrail_version="1",
        ),
    ),
)
Parameter Type Description
guardrail_config GuardrailConfig Bedrock guardrail to apply
additional_model_request_fields dict Extra fields passed to the model
additional_model_response_field_paths list[str] Extra response fields to return
pricing_as_of datetime.date Override the date used for dated pricing; defaults to the current date

For Claude 4.5 and older, reasoning_effort maps to a manual thinking budget capped below maxTokens. When max_tokens is omitted, lmux sends maxTokens: 4096, so medium and high effort both use a 4095-token thinking budget. Pass a larger max_tokens to use their full mapped budgets.

An integer manual-thinking budget in additional_model_request_fields raises the default maxTokens when needed. An explicit max_tokens is preserved instead, along with the provider-specific fields. Ensure those explicit values are compatible with the deployed model; manual thinking normally requires budget_tokens < maxTokens, except when interleaved thinking applies.

Prompt Caching

Place CachePointContent parts in UserMessage content to emit Converse cachePoint blocks marking the end of a stable prompt prefix. A cache point with no preceding block in its message is placed after whatever came before it (the prior message, or the system blocks). Markers with nothing cacheable before them are dropped, and adjacent duplicates are coalesced — the first marker wins.

from lmux import CachePointContent, TextContent, UserMessage

messages = [
    UserMessage(content=[TextContent(text=big_stable_context), CachePointContent()]),
    UserMessage(content="What changed since yesterday?"),
]

Cache points are emitted for whatever model the request targets; models without prompt-caching support reject them at request validation. Cache reads/writes are reported on response.usage (cache_read_tokens, cache_creation_tokens, and the per-TTL cache_creation_tokens_by_ttl breakdown from cacheDetails) and priced into response.cost, including per-TTL write rates where the pricing data carries them.

Constructor Options

BedrockProvider(
    auth=...,          # AuthProvider, default: BedrockEnvAuthProvider()
    region=...,        # AWS region
    endpoint_url=...,  # Custom endpoint URL (overrides region/FIPS host selection)
    use_fips=...,      # bool, default False: target the FIPS 140-3 endpoint (bedrock-runtime-fips.<region>.amazonaws.com)
    timeout=...,       # request timeout in seconds
    max_retries=...,   # retry count for transient failures
    default_headers=...,  # Optional headers included with every request
    transport=...,        # Optional httpx.BaseTransport for the sync client (proxies, testing)
    async_transport=...,  # Optional httpx.AsyncBaseTransport for the async client
)

default_headers is useful for gateway authentication, tracing, and routing. Bedrock-managed authentication and content-type headers take precedence over caller values, case-insensitively. With SigV4 authentication, custom headers are included in the request signature.

Metadata

Release files for lmux-aws-bedrock 0.13.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for lmux-aws-bedrock 0.13.5
File Size Uploaded
lmux_aws_bedrock-0.13.5.tar.gz 28.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for lmux-aws-bedrock 0.13.5
File Interpreter ABI Platform
lmux_aws_bedrock-0.13.5-py3-none-any.whl Python 3 none any Details

Total release size: 59.1 kB

Release files / lmux_aws_bedrock-0.13.5.tar.gz

Download URL lmux_aws_bedrock-0.13.5.tar.gz
Size 28.1 kB
Tags Source
SHA-256 checksum
How to use checksums
172746481728bfc63d7004836bad15c9cd37b52f80d4f251f90796d47c3b4746
BLAKE2b-256 checksum
How to use checksums
3c70d2e71f70bd3db05dfa33f220a035be94ac147a348b7cd5125e4ad822fab0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.

Transparency log

Release files / lmux_aws_bedrock-0.13.5-py3-none-any.whl

Download URL lmux_aws_bedrock-0.13.5-py3-none-any.whl
Size 31.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c21490dbbfe6e2350ee9761427434f7d3ad5a1ab7fbf83a5cdeb824b8edaf090
BLAKE2b-256 checksum
How to use checksums
4a2389f08f299350e5a812ef8479a62d351260fc6de0670cdd83fa5a0af2ba9b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.13.5 This release

2 release files

0.13.4

2 release files

0.13.1

2 release files

0.13.0

2 release files

0.12.1

2 release files

0.12.0

2 release files

0.11.0

2 release files

0.10.0

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page