Llama serve
[!WARNING] This package is now deprecated. Use
londonaicentre-mesa-localrather thanlondonaicentre-llama-servefor pip commands.
Serve llama models locally.
-
⬇️ Downloads weights from S3
-
📦 Unpacks
-
🚀 Serves via a local OpenAI-compatible server
Prerequisites
Software
- Python 3.12
Hardware
- A GPU with >=24GB VRAM (tested on NVIDIA A30)
Configuration
- Create a file called
.envin the directory where you intend to run this package. Populate it with the details you have been provided with in the following format:
MODEL_NAME=
WEIGHTS_ID=
WEIGHTS_KEY=
Installation
-
(Recommended) Create a virtual environment and activate it:
python -m venv .venv source .venv/bin/activate
-
Install this package:
pip install londonaicentre-llama-serve.
Usage
CLI
-
Note command line arguments:
Argument Description -v, --verbose Enable debug output (optional) -
Start the server as follows:
llamaserve [args].
Clients
OpenAI (example)
-
Interact with the server using the OpenAI client in python:
from openai import OpenAI client = OpenAI( base_url="http://localhost:5000/v1", api_key="blank" ) response = client.chat.completions.create( model="<MODEL_NAME>", messages=[ {"role": "system", "content": "You are an LLM named gpt-4o"}, {"role": "user", "content": "Hello"} ] ) print(response.choices[0].message.content)
License
This project uses a proprietary license (see LICENSE).
Metadata
Release files for londonaicentre-llama-serve 1.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| londonaicentre_llama_serve-1.2.0.tar.gz | 20.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| londonaicentre_llama_serve-1.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 39.1 kB
Release files / londonaicentre_llama_serve-1.2.0.tar.gz
| Download URL | londonaicentre_llama_serve-1.2.0.tar.gz |
|---|---|
| Size | 20.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4511e2b646f1ddccf55bdd66005ad55b68dcd6ffbbafdf8878eddfa571c08665
|
|
BLAKE2b-256 checksum How to use checksums |
9029b732b6c820b8d873dabe5d9d54fa96e17601c4cb0d560f6e45b9468674d8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.9.26 {"installer":{"name":"uv","version":"0.9.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Amazon Linux","version":"2023","id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / londonaicentre_llama_serve-1.2.0-py3-none-any.whl
| Download URL | londonaicentre_llama_serve-1.2.0-py3-none-any.whl |
|---|---|
| Size | 19.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
546187f60b35f016e9e6232c19451aafac74f59ae1b34be0dd5b6807a6d6e494
|
|
BLAKE2b-256 checksum How to use checksums |
663a85553467575e2ae37e55201c01df951cbcea8e4307fc6366a91f438c7078
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.9.26 {"installer":{"name":"uv","version":"0.9.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Amazon Linux","version":"2023","id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|