Skip to main content

LLMmodelHub

Turn any local machine or Google Colab notebook into a production-ready AI model server with a single line of code.

pip install LLMmodelHub
from LLMmodelHub import load_model

hub = load_model("TheBloke/Llama-2-7B-Chat-GGUF")

print(hub.local_url)    # http://0.0.0.0:5000
print(hub.public_url)   # https://xxxx.trycloudflare.com -- free, zero setup

hub.chat()               # interactive terminal chat (or a widget UI in Colab)

Features

  • One-line model loading — pulls models straight from Hugging Face, including quantized GGUF files for consumer hardware.
  • LLM + VLM support — text models and vision-language models behind one API.
  • Automatic REST API — an OpenAI-compatible /v1/chat/completions endpoint is spun up the moment the model loads.
  • Public tunneling, zero setup — by default, a free cloudflared quick tunnel is used (the binary is auto-downloaded on first run; no account needed). If you've already configured a free ngrok authtoken, ngrok is used instead. (Note: ngrok ended anonymous/no-signup tunnels, so it now requires from pyngrok import ngrok; ngrok.set_auth_token("...") — set that up once at https://dashboard.ngrok.com/signup if you prefer ngrok's stable domains.)
  • Terminal chat UI — color-coded, with /image, /clear, /url commands.
  • Colab widget UI — chat box, image upload, and URL display inside the notebook.
  • Caching — downloaded models are cached locally and reused.
  • Optional security — API key auth, CORS, and rate limiting.

Installation

pip install LLMmodelHub                 # core: API + downloader + GGUF support (llama-cpp-python)
pip install "LLMmodelHub[llm]"          # + transformers/torch for full (non-GGUF) HF models
pip install "LLMmodelHub[colab]"        # + ipywidgets for the notebook UI
pip install "LLMmodelHub[vision]"       # + pillow for VLM image input
pip install "LLMmodelHub[full]"         # everything

Note: llama-cpp-python is a core dependency (needed to run GGUF models) and compiles native code on install. On Linux/Colab this usually just works via a prebuilt wheel. On some platforms (older Windows Python versions, some Macs) pip may need to compile it from source, which requires a C++ compiler (e.g. xcode-select --install on macOS, or Visual Studio Build Tools on Windows) and can take a few minutes the first time. For GPU acceleration, install with: CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python --force-reinstall --no-cache-dir (after pip install LLMmodelHub) to rebuild it with CUDA support.

Usage

Load a full Hugging Face model

from LLMmodelHub import load_model

hub = load_model("meta-llama/Llama-2-7b-chat-hf", model_type="llm")
reply = hub.generate("Explain quantum computing in one sentence.")
print(reply)

Load a quantized GGUF model

hub = load_model("TheBloke/Llama-2-7B-Chat-GGUF", filename="llama-2-7b-chat.Q4_K_M.gguf")

Load a vision-language model

hub = load_model("Salesforce/blip2-opt-2.7b", model_type="vlm")
reply = hub.generate("What is in this image?", image_path="cat.jpg")

Call the REST API from anywhere

curl -X POST "$PUBLIC_URL/v1/chat/completions" \
  -H "Content-Type: application/json" \
  -d '{"messages": [{"role": "user", "content": "Hello!"}]}'

Stop the server

hub.stop()

License

MIT

Metadata

Release files for LLMmodelHub 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for LLMmodelHub 0.2.0
File Size Uploaded
llmmodelhub-0.2.0.tar.gz 14.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for LLMmodelHub 0.2.0
File Interpreter ABI Platform
llmmodelhub-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 30.3 kB

Release files / llmmodelhub-0.2.0.tar.gz

Download URL llmmodelhub-0.2.0.tar.gz
Size 14.2 kB
Tags Source
SHA-256 checksum
How to use checksums
35cb093b31e7e18095d8353f6040129e4d7a159cf4f2bb9b96a25e332e59c2b4
BLAKE2b-256 checksum
How to use checksums
29273cdf32b8e2bb85bbfce164da2255284ec363c9336fa655d5f2351ad9be08
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.0

Release files / llmmodelhub-0.2.0-py3-none-any.whl

Download URL llmmodelhub-0.2.0-py3-none-any.whl
Size 16.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
070f55e88eea1b5e8fc26183a903873926fc01d9a909b8d4751873ebbfc0da46
BLAKE2b-256 checksum
How to use checksums
78e001ceeab540208e87a0138ee011a0bc97f136e3febeefcb559f779f087351
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.0

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page