MESA local
Serve MESA models locally.
-
⬇️ Downloads weights from S3
-
📦 Unpacks
-
🚀 Serves via a local OpenAI-compatible server
Prerequisites
Software
- Python 3.12
Hardware
- A GPU with >=24GB VRAM (tested on NVIDIA A30).
- An NVIDIA driver supporting CUDA >=13; hosts limited to CUDA 12.x need the alternative install in CUDA 12.x hosts below.
Configuration
Create a file called .env in the directory where you intend to run this package.
Populate it with the details you have been provided with in the following format:
MODEL_NAME=
WEIGHTS_ID=
WEIGHTS_KEY=
(Alternative) S3 URI
Download weights directly from an S3 bucket:
MODEL_NAME=
WEIGHTS_URI=
WEIGHTS_REGION= # optional, defaults to eu-west-2
(Optional) Caching
Download weights and cache to S3 for faster subsequent downloads:
MODEL_NAME=
WEIGHTS_ID=
WEIGHTS_KEY=
WEIGHTS_URI=
WEIGHTS_REGION= # optional, defaults to eu-west-2
With this configuration:
-
First run: Downloads weights and uploads to S3 cache
-
Subsequent runs: Downloads directly from S3 cache (faster)
vLLM configuration
The package provides a set of vLLM configuration files for running a specific model on a specific GPU.
In addition to MODEL_NAME, this can be specified by adding GPU to the .env.
Individual vLLM settings can also be overridden by adding them to the .env file:
| Setting | Alias | Type | Default |
|---|---|---|---|
MODEL |
MODEL_NAME |
str |
mesalocal |
GPU |
str |
None |
|
MAX_MODEL_LEN |
MODEL_LENGTH |
int |
41152 |
ENFORCE_EAGER |
bool |
False |
|
ENABLE_CHUNKED_PREFILL |
bool |
True |
|
ENABLE_PREFIX_CACHING |
bool |
True |
|
GPU_MEMORY_UTILIZATION |
float |
0.9 |
|
MAX_NUM_SEQS |
int |
256 |
|
MAX_NUM_BATCHED_TOKENS |
int |
None |
|
ENABLE_LOG_REQUESTS |
bool |
False |
|
UVICORN_LOG_LEVEL |
str |
warning |
|
HTTP_TIMEOUT_KEEP_ALIVE |
int |
30 |
Installation
-
(Recommended) Create a virtual environment and activate it:
python -m venv .venv source .venv/bin/activate
-
Install this package:
pip install londonaicentre-mesa-local.
CUDA 12.x hosts
-
Install
uv. -
Create a uv-managed environment and install the CUDA 12.9 build (last supported by the current vLLM version) in place of the default above:
uv venv --python 3.12 .venv && source .venv/bin/activate uv pip install --torch-backend=cu129 \ "vllm @ https://github.com/vllm-project/vllm/releases/download/v0.30.0/vllm-0.30.0+cu129-cp38-abi3-manylinux_2_28_x86_64.whl" \ londonaicentre-mesa-local
-
(Optional) Forward compatibility is required if the current CUDA version is below 12.9; a driver already at 12.9 runs these wheels directly and can skip this step. Where needed, unpack
cuda-compat-12-9(no root needed; download the.debfrom NVIDIA's CUDA repo, copying it onto the host first if egress is restricted), prepend it to the library path in the shell that will runmesalocal, and verify (expectTrue 12.9):dpkg-deb -x cuda-compat-12-9_575.57.08-0ubuntu1_amd64.deb ~/cuda-compat export LD_LIBRARY_PATH=$HOME/cuda-compat/usr/local/cuda-12.9/compat:$LD_LIBRARY_PATH python -c "import torch; print(torch.cuda.is_available(), torch.version.cuda)"
Hosts without a CUDA compiler (nvcc)
vLLM JIT-compiles some FlashInfer and DeepGEMM kernels at runtime, which requires nvcc on the host.
Where nvcc cannot be installed, disable those kernels so vLLM uses its built-in, non-JIT paths instead:
export VLLM_USE_FLASHINFER_SAMPLER=0
export VLLM_USE_DEEP_GEMM=0
Usage
CLI (primary)
-
Note command line arguments:
Argument Description -v, --verbose Enable debug output (optional) -
Start the server as follows:
mesalocal [args].
Library (secondary)
- Import and use the logic of this package as a library:
import asyncio
from mesalocal.weights import Weights
from mesalocal.inferrer import VLLM
vllm_config: VLLMConfig = VLLMConfig() # VLLMConfig(model_name="foo", gpu="bar") to use a vLLM config without a .env file
weights: Weights = Weights(vllm_config.model)
if weights.unpack():
vllm: VLLM = VLLM(weights.get_model_folder(), vllm_config)
async def run():
async for output in vllm.generate(prompt):
print(output.outputs[0].text)
asyncio.run(run())
Clients
OpenAI (example with Oncollama)
-
Interact with the server using the OpenAI client in python:
from openai import OpenAI from oncoschema.prompt_builder import PromptBuilder # pip install londonaicentre-oncoschema client = OpenAI( base_url="http://localhost:5000/v1", api_key="blank" ) response = client.chat.completions.create( model="oncollama3betav01", messages=[ {"role": "system", "content": PromptBuilder().build_main_prompt()}, {"role": "user", "content": "Diagnosis 01/01/26..."} ] ) print(response.choices[0].message.content)
License
This project uses a proprietary license (see LICENSE).
Metadata
Release files for londonaicentre-mesa-local 2.11.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| londonaicentre_mesa_local-2.11.0.tar.gz | 29.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| londonaicentre_mesa_local-2.11.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 54.9 kB
Release files / londonaicentre_mesa_local-2.11.0.tar.gz
| Download URL | londonaicentre_mesa_local-2.11.0.tar.gz |
|---|---|
| Size | 29.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
6983256e15d6e9b50b63dec193049b88bba6f5bea71275965159477f2e99293f
|
|
BLAKE2b-256 checksum How to use checksums |
7da10919259444c1d4779c9fee25b3d644b0b0e734557fbe0d40f655b0823d55
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Amazon Linux","version":"2023","id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / londonaicentre_mesa_local-2.11.0-py3-none-any.whl
| Download URL | londonaicentre_mesa_local-2.11.0-py3-none-any.whl |
|---|---|
| Size | 26.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
555111e5c4c9a009934ea05335bac0cbae8a94f57c9a254c39295b8bf5e6bd32
|
|
BLAKE2b-256 checksum How to use checksums |
95e84e6315bc78e155b0c801d2331a9001cca6ba53d7235484dd2f7e8d193063
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.12.23 {"installer":{"name":"uv","version":"0.12.23","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Amazon Linux","version":"2023","id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|