Skip to main content

llm-anthropic

PyPI Changelog Tests License

LLM access to models by Anthropic, including the Claude series

Installation

Install this plugin in the same environment as LLM.

llm install llm-anthropic

Usage

First, set an API key for Anthropic:

llm keys set anthropic
# Paste key here

You can also set the key in the environment variable ANTHROPIC_API_KEY

Run llm models to list the models, and llm models --options to include a list of their options.

To see the models currently available to your API key directly from the Anthropic models API, run:

llm anthropic models

Example output:

claude-sonnet-5-5: Claude Sonnet 5.5 (created 2026-09-28)
claude-opus-5-5: Claude Opus 5.5 (created 2026-09-21)
claude-fable-5-1: Claude Fable 5.1 (created 2026-08-28)

Add --json to see the full JSON returned by the API, including details of each model's capabilities. Use --key to pass a different API key.

Run prompts like this:

llm -m claude-opus-5.5 'Fun facts about walruses'
llm -m claude-sonnet-5.5 'Fun facts about pelicans'
llm -m claude-haiku-4.5 'Fun facts about cormorants'

Image attachments are supported too:

llm -m claude-sonnet-5.5 'describe this image' -a https://static.simonwillison.net/static/2024/pelicans.jpg
llm -m claude-haiku-4.5 'extract text' -a page.png

Claude 3.5 and later models can handle PDF files:

llm -m claude-sonnet-5.5 'extract text' -a page.pdf

Anthropic's models support schemas. Here's how to use Claude Sonnet 5.5 to invent a dog:

llm -m claude-sonnet-5.5 --schema 'name,age int,bio: one sentence' 'invent a surprising dog'

Example output:

{
  "name": "Whiskers the Mathematical Mastiff",
  "age": 7,
  "bio": "Whiskers is a mastiff who can solve complex calculus problems by barking in binary code and has won three international mathematics competitions against human competitors."
}

New models

This plugin includes built-in settings for each Claude model it supports. You can also use models that were released after the version of the plugin you have installed:

  • Model IDs listed by the anthropic Python library are registered automatically. Run llm install -U anthropic to upgrade that library and pick up newly released models.
  • Run llm anthropic refresh to fetch the models available to your API key from the Anthropic models API. Any that this plugin does not know about will be registered using the capabilities reported by the API: image and PDF input, thinking, effort, structured outputs and maximum output tokens.
llm anthropic refresh

Example output:

Saved 12 models to /Users/you/Library/Application Support/io.datasette.llm/anthropic_models.json
Added models: claude-opus-6

The list is cached in anthropic_models.json in your LLM user directory. Run the command again to update it.

Models without built-in settings are otherwise treated like the most recent Claude models, so they think by default using adaptive thinking. If the models API has not reported a model's output limit, max_tokens defaults to 64,000.

Newer models support Anthropic's web search tool for real-time information, using the -T WebSearch server-side tool:

llm -m claude-sonnet-5.5 -T WebSearch 'What is the current weather in San Francisco?'

The tool accepts optional configuration:

llm -m claude-sonnet-5.5 \
  -T 'WebSearch(max_uses=2, user_location={"city": "London", "country": "GB"})' \
  'Recent headlines'

Available arguments:

  • max_uses: maximum number of searches per request
  • allowed_domains / blocked_domains: lists of domains to allow or block (cannot be combined)
  • user_location: dictionary with optional city, region, country and timezone keys to localize results

Note that user_location affects the results of searches from the web search tool, but location information is not made directly available to the model.

On Claude 4.6 and later models this uses the web_search_20260318 tool version with dynamic filtering; older models use the basic web_search_20250305 version.

Web fetch

Models that support web search can also use Anthropic's web fetch tool to retrieve the full content of a URL, using the -T WebFetch server-side tool:

llm -m claude-sonnet-5.5 -T WebFetch 'Fetch https://www.example.com/ and quote its first heading'

For security reasons Claude can only fetch URLs that already appear in the conversation - provided by you or returned by a previous web search or fetch.

The tool accepts optional configuration:

llm -m claude-sonnet-5.5 \
  -T 'WebFetch(max_uses=2, max_content_tokens=20000)' \
  'Summarize https://www.example.com/'

Available arguments:

  • max_uses: maximum number of fetches per request
  • allowed_domains / blocked_domains: lists of domains to allow or block (cannot be combined)
  • citations: set to True to enable citations for fetched content
  • max_content_tokens: approximate cap on fetched content included in the context
  • use_cache: set to False to bypass Anthropic's fetch cache (Claude 4.6 and later models only)

On Claude 4.6 and later models this uses the web_fetch_20260318 tool version with dynamic filtering; older models use the basic web_fetch_20250910 version.

From Python, pass an instance of the WebFetch class in tools=:

import llm
from llm_anthropic import WebFetch

model = llm.get_model("claude-sonnet-5.5")
response = model.prompt(
    "Fetch https://www.example.com/ and quote its first heading",
    tools=[WebFetch(max_uses=1)],
)
print(response.text())

MCP connector

Models that support web search can also call tools on remote MCP servers using the AnthropicMCP server-side tool. Anthropic connects to the server from their own infrastructure - it must be reachable over HTTPS:

llm -m claude-sonnet-5.5 \
  -T 'AnthropicMCP(url="https://mcp.deepwiki.com/mcp", name="deepwiki")' \
  'Use the deepwiki tools to say what simonw/llm does, one sentence'

Available arguments:

  • url: the HTTPS URL of the remote MCP server (required)
  • name: an identifier for the server - defaults to the URL's hostname
  • authorization_token: OAuth bearer token, for servers that require authentication
  • allowed_tools: optional list of tool names - if provided, only those tools are enabled
llm -m claude-sonnet-5.5 \
  -T 'AnthropicMCP(url="https://mcp.deepwiki.com/mcp", name="deepwiki", allowed_tools=["ask_question"])' \
  'What does simonw/llm do?'

You can pass multiple MCP tools to connect to more than one server in the same request. Only MCP tool calls are supported - not MCP resources or prompts.

From Python, pass an instance of the AnthropicMCP class in tools=:

import llm
from llm_anthropic import AnthropicMCP

model = llm.get_model("claude-sonnet-5.5")
response = model.prompt(
    "Use the deepwiki tools to say what simonw/llm does, one sentence",
    tools=[AnthropicMCP(url="https://mcp.deepwiki.com/mcp", name="deepwiki")],
)
print(response.text())

This feature uses Anthropic's mcp-client-2025-11-20 beta.

Code execution

Claude 4.5 and later models support Anthropic's code execution tool, which runs Python and bash in a sandboxed server-side container. Use the -T CodeExecution server-side tool:

llm -m claude-sonnet-4.6 -T CodeExecution \
  'Compute the sha256 hex digest of the string "pelican"'

Each response that runs code reports a container ID in its response_json, visible with llm logs --json. Pass that ID back to reuse the container's files and state in a later prompt (containers expire after a period of inactivity):

llm -m claude-sonnet-4.6 -T 'CodeExecution(container="container_011CPd...")' \
  'Read /tmp/results.csv and summarize it'

Fast mode

Some models support fast mode for lower latency responses. Enable it with the -o fast 1 option:

llm -m claude-opus-5 -o fast 1 'Fun facts about walruses'

Counting tokens

The llm anthropic count command uses the Anthropic token counting API to count the input tokens for a prompt, without running it:

llm anthropic count 'Fun facts about walruses' -m claude-opus-5

It outputs the number of tokens:

15

The command accepts the same options as llm prompt, so the count reflects the exact request that would be sent. That includes system prompts, attachments, fragments, templates, tools, schemas, model options and previous messages in a conversation:

cat code.py | llm anthropic count -m claude-sonnet-5.5 -s 'Review this code'
llm anthropic count -m claude-opus-5 'Describe this' -a pelican.jpg
llm anthropic count -m claude-opus-5 --schema 'name, age int' 'Invent a dog'
llm anthropic count -m claude-opus-5 -T llm_time -o thinking_effort high 'What time is it?'
llm anthropic count -c 'Tell me more'

Nothing is logged to the database when counting tokens.

Usage from Python

Python code can access the models like this:

import llm

model = llm.get_model("claude-haiku-4.5")
print(model.prompt("Fun facts about chipmunks"))

Consult LLM's Python API documentation for more details.

To count the input tokens for a prompt without running it, call model.count_tokens(). It accepts the same arguments as model.prompt() and returns an integer:

model = llm.get_model("claude-opus-5")
count = model.count_tokens(
    "Describe this image",
    system="Be concise",
    attachments=[
        llm.Attachment(url="https://static.simonwillison.net/static/2024/pelicans.jpg")
    ],
    thinking_effort="high",
)

Pass conversation= to include the previous messages from a conversation:

conversation = model.conversation()
conversation.prompt("Fun facts about pelicans").text()
count = model.count_tokens("Tell me more", conversation=conversation)

For async models, use await model.count_tokens(...).

You can also import the model classes directly, which is useful if you want to point the base_url at a different Anthropic-compatible endpoint:

from llm_anthropic import ClaudeMessages

model = ClaudeMessages(
    "MiniMax-M2",
    base_url="https://api.minimax.io/anthropic"
)

print(model.prompt("Fun facts about pangolins", key="eyJh..."))

Mid-conversation system messages

Claude Opus 4.8 and the Claude 5 family models accept updated system instructions part-way through a conversation, which preserves prompt cache hits on earlier turns. Pass a system message in an explicit messages= chain:

import llm
from llm import system, user, assistant

model = llm.get_model("claude-opus-5")
response = model.prompt(messages=[
    system("You are a helpful assistant."),
    user("Say hi to me, briefly"),
    assistant("Hi there!"),
    system("New instruction: reply only in French from now on"),
    user("Say goodbye to me, briefly"),
])
print(response.text())  # Au revoir !

The first system message is sent as the top-level system prompt; later ones are sent inline in the messages array, positioned automatically to satisfy the API's placement rules. Models older than Opus 4.8 raise a ValueError if a system message appears anywhere other than the start of the chain.

Extended thinking

Anthropic models can spend thinking tokens reasoning through a prompt before producing their response. LLM streams that reasoning to standard error as it arrives - pass -R/--hide-reasoning to hide it. The reasoning is also logged, available as the reasoning field in llm logs --json.

Claude 5 models think by default. Tune how hard they think with the thinking_effort option - one of low, medium, high, xhigh or max:

llm -m claude-opus-5 -o thinking_effort max 'Design a fair algorithm for splitting rent between roommates with different sized rooms'

Sonnet 5 and Opus 5 can have thinking turned off entirely with -o thinking 0. Fable models and the 5.5 models always think - disabling it raises an error.

Claude 4.6 and older models do not think unless asked. Enable thinking with -o thinking 1:

llm -m claude-sonnet-4.6 -o thinking 1 'Write a convincing speech to congress about the need to protect the California Brown Pelican'

Claude 4.6 models (and Opus 4.5) also support thinking_effort, which implies thinking 1. Older models than that use a fixed 1,024 token thinking budget.

Claude 4.7 and later models leave the thinking trace out of the response by default, so this plugin asks the API for display: summarized whenever thinking is on. When -R/--hide-reasoning is set it passes display: omitted instead, which leaves the thinking trace out of the response entirely - it will not appear in your logs, though thinking tokens are still billed.

The thinking_budget, thinking_display and thinking_adaptive options were removed in llm-anthropic 0.26 - install llm-anthropic==0.25 if you need them for older models.

Refusals

Claude Opus 5 and the Fable models run safety classifiers that can decline a request. The API reports this as a successful response with stop_reason: "refusal" and empty content, so this plugin raises a llm_anthropic.ClaudeRefusal exception (a subclass of llm.ModelError) instead of returning an empty string. The exception message includes the category and explanation from the API, and the exception object exposes them as .category and .explanation:

import llm
from llm_anthropic import ClaudeRefusal

model = llm.get_model("claude-fable-5.1")
try:
    print(model.prompt("...").text())
except ClaudeRefusal as ex:
    print(ex.category, ex.explanation)

Model options

The following options can be passed using -o name value on the CLI or as keyword=value arguments to the Python model.prompt() method:

  • max_tokens: int

    The maximum number of tokens to generate before stopping

  • temperature: float

    Amount of randomness injected into the response. Defaults to 1.0. Ranges from 0.0 to 1.0. Use temperature closer to 0.0 for analytical / multiple choice, and closer to 1.0 for creative and generative tasks. Note that even with temperature of 0.0, the results will not be fully deterministic.

  • top_p: float

    Use nucleus sampling. In nucleus sampling, we compute the cumulative distribution over all the options for each subsequent token in decreasing probability order and cut it off once it reaches a particular probability specified by top_p. You should either alter temperature or top_p, but not both. Recommended for advanced use cases only. You usually only need to use temperature.

  • top_k: int

    Only sample from the top K options for each subsequent token. Used to remove 'long tail' low probability responses. Recommended for advanced use cases only. You usually only need to use temperature.

  • user_id: str

    An external identifier for the user who is associated with the request

  • prefill: str

    A prefill to use for the response

  • hide_prefill: boolean

    Do not repeat the prefill value at the start of the response

  • stop_sequences: array, str

    Custom text sequences that will cause the model to stop generating - pass either a list of strings or a single string

  • cache: boolean

    Use Anthropic prompt cache for any attachments or fragments

  • fast: boolean

    Use fast mode for lower latency responses: https://platform.claude.com/docs/en/build-with-claude/fast-mode

  • thinking: boolean

    Enable thinking mode. Claude 5 models think by default - set to false to disable thinking on models that allow it

Claude 5 models no longer accept sampling parameters - setting temperature, top_p or top_k on those models returns an error from the Anthropic API.

The prefill option can be used to set the first part of the response. To increase the chance of returning JSON, set that to {:

llm -m claude-haiku-4.5 'Fun data about pelicans as JSON' \
  -o prefill '{'

If you do not want the prefill token to be echoed in the response, set hide_prefill to true:

llm -m claude-haiku-4.5 'Short python function describing a pelican' \
  -o prefill '```python' \
  -o hide_prefill true \
  -o stop_sequences '```'

This example sets ``` as the stop sequence, so the response will be a Python function without the wrapping Markdown code block.

To pass a single stop sequence, send a string:

llm -m claude-sonnet-5.5 'Fun facts about pelicans' \
  -o stop_sequences "pouch"

For multiple stop sequences, pass a JSON array:

llm -m claude-sonnet-5.5 'Fun facts about pelicans' \
  -o stop_sequences '["beak", "feathers"]'

Development

To set up this plugin locally, first checkout the code. Then use uv to run the tests:

uv run pytest
uv run pytest -k test_tools

To execute the llm command via uv:

uv run llm 'your prompt here'

To enable debug logs while running (like this), set this environment variable:

export ANTHROPIC_LOG=debug

This project uses pytest-recording to record Anthropic API responses for the tests, and inline-snapshot for test assertions.

If you add a new test that calls the API you can capture the API response like this:

PYTEST_ANTHROPIC_API_KEY="$(llm keys get anthropic)" uv run pytest --record-mode once

You will need to have stored a valid Anthropic API key using this command first:

llm keys set anthropic
# Paste key here

To re-record all cassettes and update all inline snapshot assertions in one command:

rm tests/cassettes/test_anthropic/*.yaml
PYTEST_ANTHROPIC_API_KEY="$(llm keys get anthropic)" uv run pytest --record-mode all --inline-snapshot=fix

To re-record a single test:

rm tests/cassettes/test_anthropic/test_thinking_prompt.yaml
PYTEST_ANTHROPIC_API_KEY="$(llm keys get anthropic)" uv run pytest -k test_thinking_prompt --record-mode once --inline-snapshot=fix

Metadata

Release files for llm-anthropic 0.30

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llm-anthropic 0.30
File Size Uploaded
llm_anthropic-0.30.tar.gz 48.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llm-anthropic 0.30
File Interpreter ABI Platform
llm_anthropic-0.30-py3-none-any.whl Python 3 none any Details

Total release size: 78.2 kB

Release files / llm_anthropic-0.30.tar.gz

Download URL llm_anthropic-0.30.tar.gz
Size 48.6 kB
Tags Source
SHA-256 checksum
How to use checksums
40a3f94f76d393f0101d3aac649000944396b5c8ce4dfd4b36df54e4418c63d0
BLAKE2b-256 checksum
How to use checksums
d706056e2ada6816268180bf8b4fd5ea47b7db33b93c3b4c45c3f99dc1519ea6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / llm_anthropic-0.30-py3-none-any.whl

Download URL llm_anthropic-0.30-py3-none-any.whl
Size 29.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
da5e5484a2d7a09344701684a4ddd4ffcdd8cfda5e413285ae0fb7f553f84d7f
BLAKE2b-256 checksum
How to use checksums
d837225099883f7a054343310e12a8b4732cc0cd8bda03507f5dbcc5773373bd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page