Skip to main content

lmux-aws-bedrock

AWS Bedrock provider for lmux. Talks to the Bedrock Converse and InvokeModel REST endpoints directly over httpx, using boto3 only to resolve AWS credentials for request signing.

Supports chat completions, streaming, and embeddings.

Part of the lmux ecosystem: standardized interface, cost tracking on every response, and registry-based routing across providers.

Optional Extras

  • lmux-aws-bedrock[async]: aiobotocore for async AWS credential resolution in the auth providers. Not required for achat/aembed/achat_stream themselves.

Auth

Two authentication modes are supported, resolved on the first request:

  1. Bedrock API key (simplest): set AWS_BEARER_TOKEN_BEDROCK and each request is sent with an Authorization: Bearer <token> header — no signing required.
  2. SigV4 (fallback): otherwise AWS credentials are resolved through boto3's default credential chain (env vars, AWS config, instance metadata) and each request is SigV4-signed. No extra setup needed if your AWS credentials are already configured.
from lmux_aws_bedrock import BedrockProvider

provider = BedrockProvider()

# Or specify a region
provider = BedrockProvider(region="us-east-1")

For explicit session configuration:

from lmux_aws_bedrock import BedrockSessionAuthProvider

provider = BedrockProvider(auth=BedrockSessionAuthProvider(profile_name="my-profile"))

Usage

Chat

from lmux import UserMessage

response = provider.chat("anthropic.claude-sonnet-4-20250514-v1:0", [UserMessage(content="Hello")])
print(response.content)
print(response.cost)

Streaming

for chunk in provider.chat_stream("anthropic.claude-sonnet-4-20250514-v1:0", [UserMessage(content="Hello")]):
    if chunk.delta:
        print(chunk.delta, end="")

Tool continuations

Bedrock reasoning blocks carry signatures that must be returned unmodified when a tool result continues the assistant turn. lmux-aws-bedrock preserves the native ordered Converse blocks in response.continuation; use to_assistant_message() to keep them with the normalized response:

from lmux import ToolMessage

response = provider.chat(model, messages, tools=tools, reasoning_effort="high")
messages.append(response.to_assistant_message())
messages.append(ToolMessage(content=tool_result, tool_call_id=response.tool_calls[0].id))

A matching Converse continuation is replayed exactly. Other providers ignore it and use the normalized content and tool calls.

Embeddings

response = provider.embed("amazon.titan-embed-text-v2:0", "Hello")
print(response.embeddings)

Async

All methods have async variants: achat, achat_stream, aembed. These run over httpx's async client; credentials are resolved synchronously (via boto3) even on the async path.

Bedrock also supports lmux response_format, mapped to Converse outputConfig.textFormat.

Registry

Use with the lmux registry to route across multiple providers:

from lmux import Registry

registry = Registry()
registry.register("bedrock", provider)
response = registry.chat("bedrock/anthropic.claude-sonnet-4-20250514-v1:0", messages)

Provider Params

from lmux_aws_bedrock import BedrockParams, GuardrailConfig

response = provider.chat(
    "anthropic.claude-sonnet-4-20250514-v1:0",
    messages,
    provider_params=BedrockParams(
        guardrail_config=GuardrailConfig(
            guardrail_identifier="my-guardrail",
            guardrail_version="1",
        ),
    ),
)
Parameter Type Description
guardrail_config GuardrailConfig Bedrock guardrail to apply
additional_model_request_fields dict Extra fields passed to the model
additional_model_response_field_paths list[str] Extra response fields to return
pricing_as_of datetime.date Override the date used for dated pricing; defaults to the current date

For Claude 4.5 and older, reasoning_effort maps to a manual thinking budget capped below maxTokens. When max_tokens is omitted, lmux sends maxTokens: 4096, so medium and high effort both use a 4095-token thinking budget. Pass a larger max_tokens to use their full mapped budgets.

An integer manual-thinking budget in additional_model_request_fields raises the default maxTokens when needed. An explicit max_tokens is preserved instead, along with the provider-specific fields. Ensure those explicit values are compatible with the deployed model; manual thinking normally requires budget_tokens < maxTokens, except when interleaved thinking applies.

Prompt Caching

Place CachePointContent parts in UserMessage content to emit Converse cachePoint blocks marking the end of a stable prompt prefix. A cache point with no preceding block in its message is placed after whatever came before it (the prior message, or the system blocks). Markers with nothing cacheable before them are dropped, and adjacent duplicates are coalesced — the first marker wins.

from lmux import CachePointContent, TextContent, UserMessage

messages = [
    UserMessage(content=[TextContent(text=big_stable_context), CachePointContent()]),
    UserMessage(content="What changed since yesterday?"),
]

Cache points are emitted for whatever model the request targets; models without prompt-caching support reject them at request validation. Cache reads/writes are reported on response.usage (cache_read_tokens, cache_creation_tokens, and the per-TTL cache_creation_tokens_by_ttl breakdown from cacheDetails) and priced into response.cost, including per-TTL write rates where the pricing data carries them.

Constructor Options

BedrockProvider(
    auth=...,          # AuthProvider, default: BedrockEnvAuthProvider()
    region=...,        # AWS region
    endpoint_url=...,  # Custom endpoint URL (overrides region/FIPS host selection)
    use_fips=...,      # bool, default False: target the FIPS 140-3 endpoint (bedrock-runtime-fips.<region>.amazonaws.com)
    timeout=...,       # request timeout in seconds
    max_retries=...,   # retry count for transient failures
    default_headers=...,  # Optional headers included with every request
    transport=...,        # Optional httpx.BaseTransport for the sync client (proxies, testing)
    async_transport=...,  # Optional httpx.AsyncBaseTransport for the async client
)

default_headers is useful for gateway authentication, tracing, and routing. Bedrock-managed authentication and content-type headers take precedence over caller values, case-insensitively. With SigV4 authentication, custom headers are included in the request signature.

Metadata

Release files for lmux-aws-bedrock 0.13.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for lmux-aws-bedrock 0.13.3
File Size Uploaded
lmux_aws_bedrock-0.13.3.tar.gz 27.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for lmux-aws-bedrock 0.13.3
File Interpreter ABI Platform
lmux_aws_bedrock-0.13.3-py3-none-any.whl Python 3 none any Details

Total release size: 58.3 kB

Release files / lmux_aws_bedrock-0.13.3.tar.gz

Download URL lmux_aws_bedrock-0.13.3.tar.gz
Size 27.7 kB
Tags Source
SHA-256 checksum
How to use checksums
febadd1541d374eeb0b6f9c7273ca82415395b93fcbe5a54b4c22edb3b215339
BLAKE2b-256 checksum
How to use checksums
726fe3b19dbf0c7d9d5e50da54143a6b7e584fc444f9573b58786c273621ce9a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 4, 2026.

Transparency log

Release files / lmux_aws_bedrock-0.13.3-py3-none-any.whl

Download URL lmux_aws_bedrock-0.13.3-py3-none-any.whl
Size 30.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7241af07c99442b6d640ed6bea71c34fa14f8193b783ea956443d7c6c1d9183d
BLAKE2b-256 checksum
How to use checksums
3e278e50c70cba10086a6e22288704b950420b1084c09925b442f920f8393f07
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 4, 2026.

Transparency log

Release history Release notifications | RSS feed

0.13.5

2 release files

0.13.4

2 release files

This release

0.13.3 This release

2 release files

0.13.1

2 release files

0.13.0

2 release files

0.12.1

2 release files

0.12.0

2 release files

0.11.0

2 release files

0.10.0

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page