Remote inference for language models
Project description
Remoteinference
Simple package to perform remote inference on language models of different providers.
Getting Started
Install the package
pip install remoteinference
To access an OpenAI model simply import the OpenAILLM and use the chat_completion endpoint to send your contents to the server endpoint. As a response you will receive a valid JSON containing the typicall OpenAI API conform response in a dictionary:
import os
from remoteinference.models import OpenAILLM
from remoteinference.util import user_prompt
model_type = 'gpt-4o-mini'
model = OpenAILLM(
api_key=os.environ.get('OPEANI_API_KEY'),
model=model_type
)
response = model.chat_completion(
prompt=[user_prompt('Who are you?')],
temperature=0.5,
max_tokens=50
)
print(response['choices'][0]['message']['content'])
If you have a LLM running on a remote server using llama.cpp you can initalize the model by running:
from remoteinference.models import LlamaCPPLLM
from remoteinference.util import user_prompt
# initalize the model
model = LlamaCPPLLM(
server_address='localhost',
server_port=8080
)
# run simple completion
response = model.chat_completion(
prompt=[user_prompt('Who are you?')],
temperature=0.5,
max_tokens=50
)
print(response['choices'][0]['message']['content'])
Supported Models
OpenAI
Initialize an OpenAI model by calling:
from remoteinference.models import OpenAILLM
model = OpenAILLM(
api_key='your_key',
model='gpt-4o-mini'
)
To view a full list of available models for the OpenAI endpoint see OpenAI docs
TogetherAI
Initialize an OpenAI model by calling:
from remoteinference.models import TogetherAILLM
model = TogetherAILLM(
api_key='your_key',
model='meta-llama/Llama-3-8b-hf'
)
To view a full list of available models for the OpenAI endpoint see TogetherAI docs
LlamaCPP
This package also provides functionality to query a self-hosted language model via llama.cpp
To initalize a model which is hosted locally just do:
from remoteinference.models import LlamaCPPLLM
model = LlamaCPPLLM(
server_address='localhost',
server_port=8080
)
To see the full specifications of the llama.cpp webserver see server docs.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
File details
Details for the file remoteinference-0.0.8.tar.gz.
File metadata
- Download URL: remoteinference-0.0.8.tar.gz
- Upload date:
- Size: 6.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/5.1.1 CPython/3.10.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
29bca8adf760377f69b79193669d35f1d9ac2883d22b4c5778a6bd1dc4f611d6
|
|
| MD5 |
b1d92485a6c26cb633092b1e34825439
|
|
| BLAKE2b-256 |
7c77797c84308e77673787f04e0912b16711c495fa1541f78186f2f913b4415d
|