Skip to main content

A minimal API server for local HuggingFace LLMs or VLLM LLMs

Project description

Minimal LLM Server, for API calls PyPl Total Downloads

The simplest possible Python code for running local LLM inference as a REST API server and a simple client.

This package lets you start an inference server for Hugging Face–compatible models (like LLaMA, Qwen, GPT-OSS, etc.) on your own computer or server, and make it accessible to applications via HTTP. It supports both standard HuggingFace Transformers and high-performance vLLM backends.

See the Tutorial page for extented info.

Backend Options

This package now supports two inference backends:

1. HuggingFace Transformers (Standard)

  • ✓ Widely compatible
  • ✓ CPU support available
  • ✓ Smaller installation size
  • ✓ Good for development and testing

2. vLLM Optimized (High-Performance)

  • ✓ Up to 24x faster throughput than standard transformers
  • ✓ Lower latency for single requests
  • ✓ Better GPU memory utilization with PagedAttention
  • ✓ Automatic multi-GPU support with tensor parallelism
  • ✓ Continuous batching for higher throughput
  • 🚀 Automatic optimization for quantized models - Detects GPTQ/AWQ/Int4 models and applies optimal vLLM parameters
  • ⚠ Requires CUDA GPUs (no CPU support)
  • ⚠ Best for production deployments

In comparison to the original vLLM min_llm_server_client:

  • ✓ Automatic GPU selection based on free VRAM
  • ✓ Auto-configured multi-GPU tensor parallelism
  • Smart quantized model detection - Automatically applies enforce_eager=False and kv_cache_dtype=fp8 for quantized multi-GPU setups
  • ✓ Ultra-lightweight API with minimal setup and dependencies, allows just setup and run with minimal or no configuration
  • ✓ Easier to customize and integrate into research or internal AI pipelines in research clusters.

Installation by pip

Prerequisite

uv venv --python 3.12
source .venv/bin/activate

Standard light weight Installation (HuggingFace):

uv pip install min-llm-server-client

With vLLM Support:

uv pip install "min-llm-server-client[vllm]"

Installation From Source:

git clone https://github.com/afshinsadeghi/min_llm_server_client.git
cd min_llm_server_client

# Standard installation
uv pip install .

# Or with vLLM support
uv pip install ".[vllm]"

Usage

Starting the Server

Standard HuggingFace Transformers Server

uv run min-llm-server --model_name meta-llama/Llama-3.3-70B-Instruct --max_new_tokens 100 --device cuda:0

vLLM Optimized infernce Server

uv run min-llm-server-vllm --model_name openai/gpt-oss-20b --max_new_tokens 100 --device cuda:2

Command Options:

  • --model_name : Hugging Face model name or local path suggested models: openai/gpt-oss-20b openai/gpt-oss-120b meta-llama/Llama-3.3-70B-Instruct
    casperhansen/llama-3.3-70b-instruct-awq The VLLM version runs this only on ONE A100 core" meta-llama/Llama-3.1-8B Qwen/Qwen3-0.6B Qwen/Qwen2-VL-72B-Instruct-AWQ deepseek-ai/DeepSeek-R1-Distill-Qwen-32B Qwen/Qwen3-235B-A22B-FP8 Runs it with 4 A100 cores Qwen/Qwen3-30B-A3B-Instruct-250 Qwen/Qwen3.5-397B-A17B-GPTQ-Int4 Runs it with 4 A100 cores and optimzed setting for faster infernce or it can use a local model on your device with /path/to/model.

  • --max_new_tokens : maximum number of tokens to generate in response.

  • --device : Device selection

    • auto - Auto-detect available GPUs (default)
    • cpu, - Force CPU (HuggingFace only, vLLM requires GPU)
    • cuda:0, cuda:1 , or a list of GPU cores: cuda:2,3,4,5,6,7.
  • Specific to vLLM :

    • --max_model_len : Maximum model context length. If not specified, will auto-detect from model config. Example: 8192
    • --gpu_memory_utilization : Fraction of GPU memory to use (0.0 to 1.0). Default: 0.90 (90%). Lower this value if sharing GPU with other processes. Examples: 0.85, 0.80, 0.75
    • --max_num_seqs : Maximum number of sequences to process in parallel. Lower this if you get 'max_num_seqs exceeds available cache blocks' errors. Examples: 256, 396, 512, 1024

If the device parameter is not given or is auto, it finds the available GPU cores and uses them and if no gpu is available, it uses CPU instead.

Example run:

Standard server with default settings (auto GPU detection):

min-llm-server 

Standard server on a specific GPU (e.g., GPU 0):

min-llm-server --model_name openai/gpt-oss-20b --device cuda:0

Standard server on a specific GPU (e.g., GPU 1):

min-llm-server --model_name openai/gpt-oss-120b --device cuda:1

Standard server forced on CPU:

min-llm-server --model_name openai/gpt-oss-20b --max_new_tokens 50 --device cpu

vLLM server with auto GPU detection (uses all available GPUs):

min-llm-server-vllm --model_name meta-llama/Llama-3.3-70B-Instruct

vLLM server on a specific GPU (e.g., GPU 2):

min-llm-server-vllm --model_name meta-llama/Llama-3.3-70B-Instruct --device cuda:2

vLLM server with reduced GPU memory usage (for shared GPU scenarios):

min-llm-server-vllm --model_name meta-llama/Llama-3.3-70B-Instruct --device cuda:0 --gpu_memory_utilization 0.85

Standard server on a several GPUs:

min-llm-server --model_name meta-llama/Llama-3.3-70B-Instruct --device cuda:2,3,4,5,6,7

Sending Queries

Once the server is running (default: http://127.0.0.1:5000/llm/q), you can query it with curl or Python.

Basic Curl Example:

curl -X POST http://127.0.0.1:5000/llm/q \
  -H "Content-Type: application/json" \
  -d '{"query": "What is Earth?", "key": "key1"}'

Advanced Curl Example with Generation Parameters:

curl -X POST http://127.0.0.1:5000/llm/q \
  -H "Content-Type: application/json" \
  -d '{
    "query": "Explain quantum computing in simple terms",
    "key": "key1",
    "temperature": 0.7,
    "top_p": 0.95,
    "top_k": 50,
    "repetition_penalty": 1.1,
    "presence_penalty": 0.5
  }'

Python Client - Using LLMClient Class (Recommended):

from min_llm_server_client import LLMClient

# Initialize the client
client = LLMClient(base_url="http://127.0.0.1:5000", user_key="key1")

# Ask a question
answer = client.ask_question("What is the capital of France?")
print(answer)

Python Client - Advanced with Custom Parameters:

import requests
import json

url = "http://127.0.0.1:5000/llm/q"
payload = {
    "query": "Write a short poem about AI",
    "key": "key1",
    "temperature": 0.8,
    "top_p": 0.9,
    "top_k": 40,
    "repetition_penalty": 1.2,
    "presence_penalty": 0.6
}

response = requests.post(url, json=payload)
result = response.json()
print(result['answer'])

Available Generation Parameters:

  • query (required): The text prompt/question
  • key (required): API authentication key (default: "key1")
  • temperature (optional, default: 0.1): Controls randomness (0.0 = deterministic, 1.0 = very random)
  • top_p (optional, default: 0.9): Nucleus sampling threshold (0.0-1.0)
  • top_k (optional, default: None): Top-k sampling - limits to k most likely tokens
  • repetition_penalty (optional, default: 1.2): Penalty for repeating tokens (1.0 = no penalty)
  • presence_penalty (optional, default: None): Penalty for using tokens that have appeared (vLLM only)
  • extra_body (optional, default: None): Dictionary of additional custom parameters

Note: The server now handles requests asynchronously using a thread pool, preventing blocking during inference.


Performance Comparison

LLaMA 3.1 8B - Standard HuggingFace Backend:

  • Intel CPU → ~30 seconds per request, ~2.4 GB RAM
  • A100 GPU → <1 second per request, ~34 GB GPU memory, ~4.8 GB CPU RAM

LLaMA 3.1 8B - vLLM Optimized Backend:

  • A100 GPU → ~0.1-0.3 seconds per request (3-10x faster)
  • Better memory efficiency with PagedAttention
  • Supports higher concurrent request throughput

Performance Tips:

  • Use vLLM for production deployments with high request volumes
  • Use standard backend for development, testing, or CPU-only environments
  • Both the deployement method based on Hugging face and vLLM automatically utilize multiple GPUs, vLLM with tensor parallelism
  • Both backends support the same API, making it easy to switch
  • 🚀 Quantized models (GPTQ/AWQ/Int4) on multi-GPU setups are automatically optimized - No manual configuration needed!

Automatic Quantized Model Optimization

The vLLM backend now includes intelligent model detection that automatically optimizes inference for quantized models running on multiple GPUs:

How it works:

  1. Automatic Detection: When you load a model, the server inspects the model's config.json to detect quantization (GPTQ, AWQ, Int4, etc.)
  2. Multi-GPU Check: Determines if running on multiple GPUs via tensor parallelism
  3. Smart Optimization: When both conditions are met, automatically applies:
    • enforce_eager=False - Enables CUDA graph optimization for better performance
    • kv_cache_dtype=fp8 - Uses FP8 for KV cache to maximize memory efficiency

Example - Running Qwen3.5-397B-A17B-GPTQ-Int4:

# Just run it - optimizations are applied automatically!
min-llm-server-vllm --model_name Qwen/Qwen3.5-397B-A17B-GPTQ-Int4 --device cuda:0,1,2,3,4,5,6,7

What you'll see:

🚀 AUTOMATIC OPTIMIZATION ENABLED
Detected: Quantized model + Multi-GPU setup
Applying optimizations:
  • enforce_eager = False  (Enable CUDA graph optimization)
  • kv_cache_dtype = fp8   (Use FP8 for KV cache)

Benefits:

  • Zero configuration - Works out of the box
  • Faster inference - CUDA graphs reduce overhead
  • Better memory efficiency - FP8 KV cache saves VRAM
  • Optimal for large quantized models - Perfect for 70B+, 100B+, 400B+ models
  • Transparent operation - Clear logging shows when optimizations are applied

Supported quantization formats:

  • GPTQ (e.g., Qwen3.5-397B-A17B-GPTQ-Int4)
  • AWQ (e.g., llama-3.3-70b-instruct-awq)
  • Int4/Int8 quantization
  • Any model with quantization_config or compression_config in its config

Project Structure

min_llm_server_client/
├── src/
│   ├── local_llm_inference_api_client.py
│   ├── local_llm_inference_server_api.py
│   └── ...
└── README.md

License

This project is open source under the Apache 2.0 License.


Author

Afshin Sadeghi
🔗 GitHub
🔗 Google Scholar
🔗 LinkedIn

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

min_llm_server_client-0.4.8.tar.gz (21.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

min_llm_server_client-0.4.8-py3-none-any.whl (21.1 kB view details)

Uploaded Python 3

File details

Details for the file min_llm_server_client-0.4.8.tar.gz.

File metadata

  • Download URL: min_llm_server_client-0.4.8.tar.gz
  • Upload date:
  • Size: 21.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.12

File hashes

Hashes for min_llm_server_client-0.4.8.tar.gz
Algorithm Hash digest
SHA256 bd74957dc391f1e566ccf9d235a13aa2d20bd8c6b49b46431f5b9232e0eb9146
MD5 429bbf00951da9cef1d006b31d38ff42
BLAKE2b-256 f292f2d4087d0608a0ec153df4b2e0eef4420191a01ce3402a9451670c898c13

See more details on using hashes here.

File details

Details for the file min_llm_server_client-0.4.8-py3-none-any.whl.

File metadata

File hashes

Hashes for min_llm_server_client-0.4.8-py3-none-any.whl
Algorithm Hash digest
SHA256 62ac2b6e4aa5b9500184bb2466fd31c373fc6c403420e1f87b2d490784697d90
MD5 203e5c46d7ac6621772948e1922f3d6d
BLAKE2b-256 f1d49e729a1ec35912e964d72198f91e2f96f96a78279104bfa7ee863b3908e8

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page