llmfaker
A mock server and in-process faker for OpenAI and Anthropic APIs, built for Python testing. Monkey-patches official LLM client libraries to intercept calls without network overhead.
Features
- In-process patching of
openai,anthropic,litellm, andlangchainclients - Fluent builder API for configuring responses
- Pattern matching: exact, regex, predicate-based, and template rendering
- Streaming support with realistic SSE emission (OpenAI and Anthropic formats)
- Failure injection: rate limits, timeouts, mid-stream disconnects, malformed JSON
- Latency simulation with configurable TTFT and inter-token delays
- Record/replay cassettes for integration testing
- Multi-turn conversation scripting and tool-call sequences
- Token counting and cost estimation via pricing tables
- Pytest plugin with
llm_fakerandllm_recordingfixtures - Standalone mock server mode via CLI
Installation
pip install llmfaker
Quick Start
In-process (for unit tests)
from llmfaker import LLMFaker
import openai
client = openai.OpenAI(api_key="fake")
with LLMFaker() as faker:
faker.when(prompt_contains="weather").respond("It's sunny!")
response = client.chat.completions.create(
model="gpt-4",
messages=[{"role": "user", "content": "What's the weather?"}],
)
print(response.choices[0].message.content) # "It's sunny!"
Pytest plugin
def test_my_feature(llm_faker):
llm_faker.when(prompt_contains="hello").respond("Hi!")
response = client.chat.completions.create(
model="gpt-4",
messages=[{"role": "user", "content": "hello"}],
)
assert "Hi!" in response.choices[0].message.content
assert len(llm_faker.calls) == 1
Cassette record/replay
def test_real_api_behavior(llm_recording):
# First run: calls real API and records to cassette
# Subsequent runs: replays from cassette file
response = client.chat.completions.create(
model="gpt-4",
messages=[{"role": "user", "content": "hello"}],
)
assert response.choices[0].message.content
Failure injection
with LLMFaker() as faker:
with faker.fail(rate=1.0, status=429, retry_after=30):
# All calls will get a 429 rate limit error
...
Standalone mock server
mockllm start --responses responses.yml --port 8000
YAML Configuration
responses:
"what colour is the sky?": "The sky is blue due to Rayleigh scattering."
"tell me a joke": "Why don't programmers like nature? Too many bugs!"
defaults:
unknown_response: "I don't know the answer to that."
settings:
lag_enabled: true
lag_factor: 10
Development
pip install -r requirements.txt
pip install -e .
# Run tests
python -m pytest tests/ -v
License
Inspired by mockllm.
Metadata
Release files for llmfaker 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llmfaker-0.1.0.tar.gz | 114.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llmfaker-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 203.4 kB
Release files / llmfaker-0.1.0.tar.gz
| Download URL | llmfaker-0.1.0.tar.gz |
|---|---|
| Size | 114.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
eead8e3481eafebccf96382f0f83fc24d781f834d2cc9958c240f358605d8e5c
|
|
BLAKE2b-256 checksum How to use checksums |
b97cd47a63be466d03bf87f14203fdb0e2bd41c508c664e63b4d610617ab7bd8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.12
|
Release files / llmfaker-0.1.0-py3-none-any.whl
| Download URL | llmfaker-0.1.0-py3-none-any.whl |
|---|---|
| Size | 88.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
67562fbf8d0292d64891ce3785d3ab3fed1d5ba6f3307f45891e3353e578d1fd
|
|
BLAKE2b-256 checksum How to use checksums |
12d57a846ecd570a9dda5ff1842d0ab374c48f0466bd2da57c1ba35ae7b58a54
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.12.12
|