llm-anthropic
LLM access to models by Anthropic, including the Claude series
Installation
Install this plugin in the same environment as LLM.
llm install llm-anthropic
Usage
First, set an API key for Anthropic:
llm keys set anthropic
# Paste key here
You can also set the key in the environment variable ANTHROPIC_API_KEY
Run llm models to list the models, and llm models --options to include a list of their options.
To see the models currently available to your API key directly from the Anthropic models API, run:
llm anthropic models
Example output:
claude-sonnet-5-5: Claude Sonnet 5.5 (created 2026-09-28)
claude-opus-5-5: Claude Opus 5.5 (created 2026-09-21)
claude-fable-5-1: Claude Fable 5.1 (created 2026-08-28)
Add --json to see the full JSON returned by the API, including details of each model's capabilities. Use --key to pass a different API key.
Run prompts like this:
llm -m claude-opus-5.5 'Fun facts about walruses'
llm -m claude-sonnet-5.5 'Fun facts about pelicans'
llm -m claude-haiku-4.5 'Fun facts about cormorants'
Image attachments are supported too:
llm -m claude-sonnet-5.5 'describe this image' -a https://static.simonwillison.net/static/2024/pelicans.jpg
llm -m claude-haiku-4.5 'extract text' -a page.png
Claude 3.5 and later models can handle PDF files:
llm -m claude-sonnet-5.5 'extract text' -a page.pdf
Anthropic's models support schemas. Here's how to use Claude Sonnet 5.5 to invent a dog:
llm -m claude-sonnet-5.5 --schema 'name,age int,bio: one sentence' 'invent a surprising dog'
Example output:
{
"name": "Whiskers the Mathematical Mastiff",
"age": 7,
"bio": "Whiskers is a mastiff who can solve complex calculus problems by barking in binary code and has won three international mathematics competitions against human competitors."
}
New models
This plugin includes built-in settings for each Claude model it supports. You can also use models that were released after the version of the plugin you have installed:
- Model IDs listed by the anthropic Python library are registered automatically. Run
llm install -U anthropicto upgrade that library and pick up newly released models. - Run
llm anthropic refreshto fetch the models available to your API key from the Anthropic models API. Any that this plugin does not know about will be registered using the capabilities reported by the API: image and PDF input, thinking, effort, structured outputs and maximum output tokens.
llm anthropic refresh
Example output:
Saved 12 models to /Users/you/Library/Application Support/io.datasette.llm/anthropic_models.json
Added models: claude-opus-6
The list is cached in anthropic_models.json in your LLM user directory. Run the command again to update it.
Models without built-in settings are otherwise treated like the most recent Claude models, so they think by default using adaptive thinking. If the models API has not reported a model's output limit, max_tokens defaults to 64,000.
Web search
Newer models support Anthropic's web search tool for real-time information, using the -T WebSearch server-side tool:
llm -m claude-sonnet-5.5 -T WebSearch 'What is the current weather in San Francisco?'
The tool accepts optional configuration:
llm -m claude-sonnet-5.5 \
-T 'WebSearch(max_uses=2, user_location={"city": "London", "country": "GB"})' \
'Recent headlines'
Available arguments:
max_uses: maximum number of searches per requestallowed_domains/blocked_domains: lists of domains to allow or block (cannot be combined)user_location: dictionary with optionalcity,region,countryandtimezonekeys to localize results
Note that user_location affects the results of searches from the web search tool, but location information is not made directly available to the model.
On Claude 4.6 and later models this uses the web_search_20260318 tool version with dynamic filtering; older models use the basic web_search_20250305 version.
Web fetch
Models that support web search can also use Anthropic's web fetch tool to retrieve the full content of a URL, using the -T WebFetch server-side tool:
llm -m claude-sonnet-5.5 -T WebFetch 'Fetch https://www.example.com/ and quote its first heading'
For security reasons Claude can only fetch URLs that already appear in the conversation - provided by you or returned by a previous web search or fetch.
The tool accepts optional configuration:
llm -m claude-sonnet-5.5 \
-T 'WebFetch(max_uses=2, max_content_tokens=20000)' \
'Summarize https://www.example.com/'
Available arguments:
max_uses: maximum number of fetches per requestallowed_domains/blocked_domains: lists of domains to allow or block (cannot be combined)citations: set toTrueto enable citations for fetched contentmax_content_tokens: approximate cap on fetched content included in the contextuse_cache: set toFalseto bypass Anthropic's fetch cache (Claude 4.6 and later models only)
On Claude 4.6 and later models this uses the web_fetch_20260318 tool version with dynamic filtering; older models use the basic web_fetch_20250910 version.
From Python, pass an instance of the WebFetch class in tools=:
import llm
from llm_anthropic import WebFetch
model = llm.get_model("claude-sonnet-5.5")
response = model.prompt(
"Fetch https://www.example.com/ and quote its first heading",
tools=[WebFetch(max_uses=1)],
)
print(response.text())
MCP connector
Models that support web search can also call tools on remote MCP servers using the AnthropicMCP server-side tool. Anthropic connects to the server from their own infrastructure - it must be reachable over HTTPS:
llm -m claude-sonnet-5.5 \
-T 'AnthropicMCP(url="https://mcp.deepwiki.com/mcp", name="deepwiki")' \
'Use the deepwiki tools to say what simonw/llm does, one sentence'
Available arguments:
url: the HTTPS URL of the remote MCP server (required)name: an identifier for the server - defaults to the URL's hostnameauthorization_token: OAuth bearer token, for servers that require authenticationallowed_tools: optional list of tool names - if provided, only those tools are enabled
llm -m claude-sonnet-5.5 \
-T 'AnthropicMCP(url="https://mcp.deepwiki.com/mcp", name="deepwiki", allowed_tools=["ask_question"])' \
'What does simonw/llm do?'
You can pass multiple MCP tools to connect to more than one server in the same request. Only MCP tool calls are supported - not MCP resources or prompts.
From Python, pass an instance of the AnthropicMCP class in tools=:
import llm
from llm_anthropic import AnthropicMCP
model = llm.get_model("claude-sonnet-5.5")
response = model.prompt(
"Use the deepwiki tools to say what simonw/llm does, one sentence",
tools=[AnthropicMCP(url="https://mcp.deepwiki.com/mcp", name="deepwiki")],
)
print(response.text())
This feature uses Anthropic's mcp-client-2025-11-20 beta.
Code execution
Claude 4.5 and later models support Anthropic's code execution tool, which runs Python and bash in a sandboxed server-side container. Use the -T CodeExecution server-side tool:
llm -m claude-sonnet-4.6 -T CodeExecution \
'Compute the sha256 hex digest of the string "pelican"'
Each response that runs code reports a container ID in its response_json, visible with llm logs --json. Pass that ID back to reuse the container's files and state in a later prompt (containers expire after a period of inactivity):
llm -m claude-sonnet-4.6 -T 'CodeExecution(container="container_011CPd...")' \
'Read /tmp/results.csv and summarize it'
Fast mode
Some models support fast mode for lower latency responses. Enable it with the -o fast 1 option:
llm -m claude-opus-5 -o fast 1 'Fun facts about walruses'
Counting tokens
The llm anthropic count command uses the Anthropic token counting API to count the input tokens for a prompt, without running it:
llm anthropic count 'Fun facts about walruses' -m claude-opus-5
It outputs the number of tokens:
15
The command accepts the same options as llm prompt, so the count reflects the exact request that would be sent. That includes system prompts, attachments, fragments, templates, tools, schemas, model options and previous messages in a conversation:
cat code.py | llm anthropic count -m claude-sonnet-5.5 -s 'Review this code'
llm anthropic count -m claude-opus-5 'Describe this' -a pelican.jpg
llm anthropic count -m claude-opus-5 --schema 'name, age int' 'Invent a dog'
llm anthropic count -m claude-opus-5 -T llm_time -o thinking_effort high 'What time is it?'
llm anthropic count -c 'Tell me more'
Nothing is logged to the database when counting tokens.
Usage from Python
Python code can access the models like this:
import llm
model = llm.get_model("claude-haiku-4.5")
print(model.prompt("Fun facts about chipmunks"))
Consult LLM's Python API documentation for more details.
To count the input tokens for a prompt without running it, call model.count_tokens(). It accepts the same arguments as model.prompt() and returns an integer:
model = llm.get_model("claude-opus-5")
count = model.count_tokens(
"Describe this image",
system="Be concise",
attachments=[
llm.Attachment(url="https://static.simonwillison.net/static/2024/pelicans.jpg")
],
thinking_effort="high",
)
Pass conversation= to include the previous messages from a conversation:
conversation = model.conversation()
conversation.prompt("Fun facts about pelicans").text()
count = model.count_tokens("Tell me more", conversation=conversation)
For async models, use await model.count_tokens(...).
You can also import the model classes directly, which is useful if you want to point the base_url at a different Anthropic-compatible endpoint:
from llm_anthropic import ClaudeMessages
model = ClaudeMessages(
"MiniMax-M2",
base_url="https://api.minimax.io/anthropic"
)
print(model.prompt("Fun facts about pangolins", key="eyJh..."))
Mid-conversation system messages
Claude Opus 4.8 and the Claude 5 family models accept updated system instructions part-way through a conversation, which preserves prompt cache hits on earlier turns. Pass a system message in an explicit messages= chain:
import llm
from llm import system, user, assistant
model = llm.get_model("claude-opus-5")
response = model.prompt(messages=[
system("You are a helpful assistant."),
user("Say hi to me, briefly"),
assistant("Hi there!"),
system("New instruction: reply only in French from now on"),
user("Say goodbye to me, briefly"),
])
print(response.text()) # Au revoir !
The first system message is sent as the top-level system prompt; later ones are sent inline in the messages array, positioned automatically to satisfy the API's placement rules. Models older than Opus 4.8 raise a ValueError if a system message appears anywhere other than the start of the chain.
Extended thinking
Anthropic models can spend thinking tokens reasoning through a prompt before producing their response. LLM streams that reasoning to standard error as it arrives - pass -R/--hide-reasoning to hide it. The reasoning is also logged, available as the reasoning field in llm logs --json.
Claude 5 models think by default. Tune how hard they think with the thinking_effort option - one of low, medium, high, xhigh or max:
llm -m claude-opus-5 -o thinking_effort max 'Design a fair algorithm for splitting rent between roommates with different sized rooms'
Sonnet 5 and Opus 5 can have thinking turned off entirely with -o thinking 0. Fable models and the 5.5 models always think - disabling it raises an error.
Claude 4.6 and older models do not think unless asked. Enable thinking with -o thinking 1:
llm -m claude-sonnet-4.6 -o thinking 1 'Write a convincing speech to congress about the need to protect the California Brown Pelican'
Claude 4.6 models (and Opus 4.5) also support thinking_effort, which implies thinking 1. Older models than that use a fixed 1,024 token thinking budget.
Claude 4.7 and later models leave the thinking trace out of the response by default, so this plugin asks the API for display: summarized whenever thinking is on. When -R/--hide-reasoning is set it passes display: omitted instead, which leaves the thinking trace out of the response entirely - it will not appear in your logs, though thinking tokens are still billed.
The thinking_budget, thinking_display and thinking_adaptive options were removed in llm-anthropic 0.26 - install llm-anthropic==0.25 if you need them for older models.
Refusals
Claude Opus 5 and the Fable models run safety classifiers that can decline a request. The API reports this as a successful response with stop_reason: "refusal" and empty content, so this plugin raises a llm_anthropic.ClaudeRefusal exception (a subclass of llm.ModelError) instead of returning an empty string. The exception message includes the category and explanation from the API, and the exception object exposes them as .category and .explanation:
import llm
from llm_anthropic import ClaudeRefusal
model = llm.get_model("claude-fable-5.1")
try:
print(model.prompt("...").text())
except ClaudeRefusal as ex:
print(ex.category, ex.explanation)
Model options
The following options can be passed using -o name value on the CLI or as keyword=value arguments to the Python model.prompt() method:
-
max_tokens:
intThe maximum number of tokens to generate before stopping
-
temperature:
floatAmount of randomness injected into the response. Defaults to 1.0. Ranges from 0.0 to 1.0. Use temperature closer to 0.0 for analytical / multiple choice, and closer to 1.0 for creative and generative tasks. Note that even with temperature of 0.0, the results will not be fully deterministic.
-
top_p:
floatUse nucleus sampling. In nucleus sampling, we compute the cumulative distribution over all the options for each subsequent token in decreasing probability order and cut it off once it reaches a particular probability specified by top_p. You should either alter temperature or top_p, but not both. Recommended for advanced use cases only. You usually only need to use temperature.
-
top_k:
intOnly sample from the top K options for each subsequent token. Used to remove 'long tail' low probability responses. Recommended for advanced use cases only. You usually only need to use temperature.
-
user_id:
strAn external identifier for the user who is associated with the request
-
prefill:
strA prefill to use for the response
-
hide_prefill:
booleanDo not repeat the prefill value at the start of the response
-
stop_sequences:
array, strCustom text sequences that will cause the model to stop generating - pass either a list of strings or a single string
-
cache:
booleanUse Anthropic prompt cache for any attachments or fragments
-
fast:
booleanUse fast mode for lower latency responses: https://platform.claude.com/docs/en/build-with-claude/fast-mode
-
thinking:
booleanEnable thinking mode. Claude 5 models think by default - set to false to disable thinking on models that allow it
Claude 5 models no longer accept sampling parameters - setting temperature, top_p or top_k on those models returns an error from the Anthropic API.
The prefill option can be used to set the first part of the response. To increase the chance of returning JSON, set that to {:
llm -m claude-haiku-4.5 'Fun data about pelicans as JSON' \
-o prefill '{'
If you do not want the prefill token to be echoed in the response, set hide_prefill to true:
llm -m claude-haiku-4.5 'Short python function describing a pelican' \
-o prefill '```python' \
-o hide_prefill true \
-o stop_sequences '```'
This example sets ``` as the stop sequence, so the response will be a Python function without the wrapping Markdown code block.
To pass a single stop sequence, send a string:
llm -m claude-sonnet-5.5 'Fun facts about pelicans' \
-o stop_sequences "pouch"
For multiple stop sequences, pass a JSON array:
llm -m claude-sonnet-5.5 'Fun facts about pelicans' \
-o stop_sequences '["beak", "feathers"]'
Development
To set up this plugin locally, first checkout the code. Then use uv to run the tests:
uv run pytest
uv run pytest -k test_tools
To execute the llm command via uv:
uv run llm 'your prompt here'
To enable debug logs while running (like this), set this environment variable:
export ANTHROPIC_LOG=debug
This project uses pytest-recording to record Anthropic API responses for the tests, and inline-snapshot for test assertions.
If you add a new test that calls the API you can capture the API response like this:
PYTEST_ANTHROPIC_API_KEY="$(llm keys get anthropic)" uv run pytest --record-mode once
You will need to have stored a valid Anthropic API key using this command first:
llm keys set anthropic
# Paste key here
To re-record all cassettes and update all inline snapshot assertions in one command:
rm tests/cassettes/test_anthropic/*.yaml
PYTEST_ANTHROPIC_API_KEY="$(llm keys get anthropic)" uv run pytest --record-mode all --inline-snapshot=fix
To re-record a single test:
rm tests/cassettes/test_anthropic/test_thinking_prompt.yaml
PYTEST_ANTHROPIC_API_KEY="$(llm keys get anthropic)" uv run pytest -k test_thinking_prompt --record-mode once --inline-snapshot=fix
Metadata
Release files for llm-anthropic 0.30
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llm_anthropic-0.30.tar.gz | 48.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llm_anthropic-0.30-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 78.2 kB
Release files / llm_anthropic-0.30.tar.gz
| Download URL | llm_anthropic-0.30.tar.gz |
|---|---|
| Size | 48.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
40a3f94f76d393f0101d3aac649000944396b5c8ce4dfd4b36df54e4418c63d0
|
|
BLAKE2b-256 checksum How to use checksums |
d706056e2ada6816268180bf8b4fd5ea47b7db33b93c3b4c45c3f99dc1519ea6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency logRelease files / llm_anthropic-0.30-py3-none-any.whl
| Download URL | llm_anthropic-0.30-py3-none-any.whl |
|---|---|
| Size | 29.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
da5e5484a2d7a09344701684a4ddd4ffcdd8cfda5e413285ae0fb7f553f84d7f
|
|
BLAKE2b-256 checksum How to use checksums |
d837225099883f7a054343310e12a8b4732cc0cd8bda03507f5dbcc5773373bd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.
Transparency log