Skip to main content

Langchain OpenAI limiter

Goal

By default langchain only do retries if OpenAI queries hit limits. Which could lead to spending many resources in some cases.

Moreover, OpenAI have very different tiers for different users.

Like by default for GPT-4 it's something like 10 000 TPM (token per minute) and 1000 RPM (request per minute) - and up to something like 10 000 RPM and 150 000 TPM (in my personal case).

Fortunately, they provide response headers with all the required info, so we don't have to monitor it ourselves.

Unfortunately, neither OpenAI python library nor LangChain built on top of that do not provide easy built-in access to them.

So I made this package

Installation

You should be able to install it via pip, like

pip install langchain_openai_limiter

Examples

You could see example.ipynb notebook for examples. However:

Chat completion

# LangChain built-in model
chat_model = ChatOpenAI(
    model_name="gpt-4-0613",
    streaming=True,
)
# Thing which will await for rate/token limits
chat_model_limit_await = LimitAwaitChatOpenAI(
    chat_openai=chat_model,
    limit_await_timeout=60.0,
    limit_await_sleep=0.1,
)
# Thing which will do key rotation
chat_model_key_choose = ChooseKeyChatOpenAI(
    chat_openai=chat_model_limit_await,
    openai_api_keys=[
        os.environ["OPENAI_API_KEY0"],
        os.environ["OPENAI_API_KEY1"],
    ]
)

all three things is compatible with LangChain's ChatModel, so:

history = [
    SystemMessage(
        content="You are a helpful assistant that translates English to French."
    ),
    HumanMessage(
        content="Translate this sentence from English to French. I love programming."
    ),
]
print(chat_model_key_choose.invoke(history).content)

J'aime la programmation.

Async and streaming methods implemented as well.

Embeddings

Pretty often we do not only need chat models - we need embeddings (for RAG, for instance) too:

# LangChain built-in model
embedder_model = OpenAIEmbeddings(
    model="text-embedding-ada-002",
)
# Thing which will await for rate/token limits
embedder_model_limit_await = LimitAwaitOpenAIEmbeddings(
    openai_embeddings=embedder_model,
    limit_await_timeout=60.0,
    limit_await_sleep=0.1,
)
# Thing which will do key rotation
embedder_model_key_choose = ChooseKeyOpenAIEmbeddings(
    openai_embeddings=embedder_model_limit_await,
    openai_api_keys=[
        os.environ["OPENAI_API_KEY0"],
        os.environ["OPENAI_API_KEY1"],
    ]
)
docs = embedder_model_key_choose.embed_documents([
    "Markdown is a lightweight markup language",
    "Brainfuck is an esoteric programming language",
])
query = embedder_model_key_choose.embed_query("What is Markdown?")

-0.01 0.03 -0.00 -0.00 0.00 ... -0.02 0.00 -0.01 -0.00 -0.00 ... -0.01 0.01 0.00 -0.01 0.00 ...

Testing

To run tests - you can do the following stuff

pip install langchain_openai_limiter[dev]
pytest

Metadata

Release files for langchain-openai-limiter 0.0.2.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for langchain-openai-limiter 0.0.2.5
File Size Uploaded
langchain_openai_limiter-0.0.2.5.tar.gz 12.0 kB Details

Release files / langchain_openai_limiter-0.0.2.5.tar.gz

Download URL langchain_openai_limiter-0.0.2.5.tar.gz
Size 12.0 kB
Tags Source
SHA-256 checksum
How to use checksums
51b618e119e8ea701f53fb10a86cd5210f6b77179830f2a1183a98c1cdec7d6f
BLAKE2b-256 checksum
How to use checksums
ca6364f673530049bda0148d759f0284d50910537dc181826a6cba8b09d129f2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.11.5

Release history Release notifications | RSS feed

This release

0.0.2.5 This release

1 release file

0.0.2.4

1 release file

0.0.2

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page