This release is a pre-release and may not be stable for production use.
llm-chat-completions-server
LLM plugin to serve an OpenAI Chat Completions API endpoint
Installation
Install this plugin in the same environment as LLM.
llm install llm-chat-completions-server
Usage
Start the server:
llm chat-completions-server
It listens on http://127.0.0.1:8002 by default. Use -p or --port to
select a different port:
llm chat-completions-server --port 9000
For development, --reload restarts the server when Python files change:
llm chat-completions-server --reload
The server does not require an API token. Models may still need credentials
configured on the server using the usual llm keys set commands.
List models
GET /v1/models lists the models registered with LLM that provide an async
implementation. Sync-only models are not served.
curl http://127.0.0.1:8002/v1/models
Create a chat completion
POST /v1/chat/completions implements the classic OpenAI Chat Completions
request and response format using LLM's async model API.
curl -s http://127.0.0.1:8002/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "gemini/gemini-2.5-flash",
"messages": [
{"role": "user", "content": "Tell me a short joke about pelicans"}
],
"stream": false
}'
Streaming responses use server-sent events and finish with data: [DONE]:
curl -N http://127.0.0.1:8002/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "gemini/gemini-2.5-flash",
"messages": [
{"role": "user", "content": "Count to three"}
],
"stream": true,
"stream_options": {"include_usage": true}
}'
The endpoint accepts system, developer, user, assistant and tool messages,
including image URL content and function tool calls. Sampling options are
passed through when they are supported by the selected LLM model. Only
n=1 is currently supported.
Completed responses—including streamed responses—are written to LLM's normal
logs.db. This populates both the legacy response tables and the current
content-addressed message and turn tables. The global llm logs off setting
is respected.
Development
This checkout is configured to use the editable in-development LLM checkout
at ~/dev/llm. Run the tests with:
cd llm-chat-completions-server
uv run pytest
To run the server with additional model plugins under development:
uv run \
--with-editable . \
--with-editable ~/dev/llm \
--with-editable ~/dev/ecosystem/llm-gemini \
--with-editable ~/dev/ecosystem/llm-anthropic \
--with-editable ~/dev/ecosystem/llm-apple-foundation \
llm chat-completions-server --reload
Release files for llm-chat-completions-server 0.1a0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llm_chat_completions_server-0.1a0.tar.gz | 15.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llm_chat_completions_server-0.1a0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 27.7 kB
Release files / llm_chat_completions_server-0.1a0.tar.gz
| Download URL | llm_chat_completions_server-0.1a0.tar.gz |
|---|---|
| Size | 15.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3c70ef5fab6ef9cf3330b4c335f0650f2daefc7bf12e2447dbb657ff9e6efd23
|
|
BLAKE2b-256 checksum How to use checksums |
368a10251f5d45e64c445265e742f880772df483d19b5e7cad717aa7a70a5de0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 30, 2026.
Transparency logRelease files / llm_chat_completions_server-0.1a0-py3-none-any.whl
| Download URL | llm_chat_completions_server-0.1a0-py3-none-any.whl |
|---|---|
| Size | 12.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
46a864d8a152019cf105765c5b18429fadfdc9061cd1cab38f537512a4b8ad92
|
|
BLAKE2b-256 checksum How to use checksums |
9a7823c0334e9b006c7057793cf7e0111745e1e9abb3702448afaad976ad46f6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 30, 2026.
Transparency log