Skip to main content
LLMeter (Logo)

Measuring large language models latency and throughput

Latest Version Documentation: Online Supported Python Versions Code Style: Ruff

LLMeter is a pure-python library for simple latency and throughput testing of large language models (LLMs). It's designed to be lightweight to install; straightforward to run standard tests; and versatile to integrate - whether in notebooks, CI/CD, or other workflows.

📖 For full details, check out our documentation at: https://awslabs.github.io/llmeter

🛠️ Installation

LLMeter requires python>=3.10, please make sure your current version of python is compatible.

To install the basic metering functionalities, you can install the minimum package using pip or uv:

pip install llmeter

Or with uv (recommended for faster installation):

uv pip install llmeter

LLMeter also offers extra features that require additional dependencies. Currently these extras include:

  • plotting: Add methods to generate charts to summarize the results
  • openai: Enable testing endpoints offered by OpenAI
  • litellm: Enable testing a range of different models through LiteLLM
  • mlflow: Enable logging LLMeter experiments to MLFlow

You can install one or more of these extra options using pip:

pip install 'llmeter[plotting,openai,litellm,mlflow]'

Or with uv:

uv pip install 'llmeter[plotting,openai,litellm,mlflow]'

🚀 Quick-start

At a high level, you'll start by configuring an LLMeter "Endpoint" for whatever type of LLM you're connecting to:

# For example with Amazon Bedrock...
from llmeter.endpoints import BedrockConverse
endpoint = BedrockConverse(model_id="...")

# ...or OpenAI...
from llmeter.endpoints import OpenAIEndpoint
endpoint = OpenAIEndpoint(model_id="...", api_key="...")

# ...or via LiteLLM...
from llmeter.endpoints import LiteLLM
endpoint = LiteLLM("{provider}/{model_id}")

# ...and so on

You can then run the high-level "experiments" offered by LLMeter:

# Testing how throughput varies with concurrent request count:
from llmeter.experiments import LoadTest
load_test = LoadTest(
    endpoint=endpoint,
    payload={...},
    sequence_of_clients=[1, 5, 20, 50, 100, 500],
    output_path="local or S3 path"
)
load_test_results = await load_test.run()
load_test_results.plot_results()

Where payload can be a single dictionary, a list of dictionary, or a path to a JSON Line file that contains a payload for every line.

Each LLMeter Endpoint type offers a create_payload() function you can use to help build your inputs, in case you're not sure of the request JSON format for your target API. For example with Amazon Bedrock Converse:

from llmeter.prompt_utils import ImageContent
payload = BedrockConverse.create_payload(
    user_messages=[
        "Describe the following image:",
        ImageContent.from_path("photo.jpg"),
    ],
    max_tokens=1024,
)

As well as the high-level Experiments, you can use the low-level llmeter.runner.Runner class to run and analyze request batches - and build your own custom experiments.

from llmeter.runner import Runner

endpoint_test = Runner(
    endpoint,
    tokenizer=tokenizer,
    output_path="local or S3 path",
)
result = await endpoint_test.run(
    payload={...},
    n_requests=3,
    clients=3,
)

print(result.stats)

Additional functionality like cost modelling and MLFlow experiment tracking is enabled through llmeter.callbacks, and you can write your own callbacks to hook other custom logic into LLMeter test runs.

For more details, check out the LLMeter user guide and our selection of end-to-end code examples in the examples folder!

Analyze and compare results

You can analyze the results of a single run or a load test by generating interactive charts. You can find examples in in the examples folder.

Load testing

You can generate a collection of standard charts to visualize the result of a load test:

# Load test results
from llmeter.experiments import LoadTestResult
load_test_result = LoadTestResult.load("local or S3 path", test_name="Test result")

figures = load_test_result.plot_results()
Average input tokens Average output tokens
Error rate Request per minute
--- ---
Time to first token Time to last token

You can see how to compare two load test in Compare load test.

Single Run visualizations

Metrics like time to first token (TTFT) and time per output token (TPOT) are described as distributions. While statistical descriptions of these distributions (median, 90th percentile, average, etc.) are a convenient way to compare them, visualizations provide insights on the endpoint behavior.

Boxplot

import plotly.graph_objects as go
from llmeter.plotting import boxplot_by_dimension

result = Result.load("local or S3 path")

fig = go.Figure()
trace = boxplot_by_dimension(result=result, dimension="time_to_first_token")
fig.add_trace(trace)

Multiple traces can easily be combined into the same figure.

alt text

Histograms

import plotly.graph_objects as go
from llmeter.plotting import histogram_by_dimension

result = Result.load("local or S3 path")

fig = go.Figure()
trace = histogram_by_dimension(result=result, dimension="time_to_first_token", xbins={"size":0.02})
fig.add_trace(trace)

Multiple traces can easily be combined into the same figure.

alt text

Security

See CONTRIBUTING for more information.

License

This project is licensed under the Apache-2.0 License.

Metadata

Release files for llmeter 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llmeter 0.2.0
File Size Uploaded
llmeter-0.2.0.tar.gz 604.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llmeter 0.2.0
File Interpreter ABI Platform
llmeter-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 721.3 kB

Release files / llmeter-0.2.0.tar.gz

Download URL llmeter-0.2.0.tar.gz
Size 604.0 kB
Tags Source
SHA-256 checksum
How to use checksums
b00739cab732bb98b70dc4a4003ad854ab7ea1c68305273f7e9b80a34fcc8fab
BLAKE2b-256 checksum
How to use checksums
33d7fe679a44422fb2417fc4ef35081f8235c0777f88d4d8228a5d1ca6802b05
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 30, 2026.

Transparency log

Release files / llmeter-0.2.0-py3-none-any.whl

Download URL llmeter-0.2.0-py3-none-any.whl
Size 117.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f96a149625b9fbcd79403a9a073a8a2a30ecf84ff45abde78842220faa62c644
BLAKE2b-256 checksum
How to use checksums
8187764f299e0a0f75dbc3b54038767e5281376b264633feee3a73424b7106f5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 30, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.12

2 release files

0.1.11

2 release files

0.1.10

2 release files

0.1.9

2 release files

0.1.8

2 release files

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page