olmon
A lightweight CLI to monitor your Ollama models and usage in real time.
$ olmon top
╭─────────────────────────────────────── olmon top 01:09:32 ───────────────────────────────────────╮
│ Model VRAM VRAM % Expires In Status │
│ ornith:9b 5.0 GB ██░░░ 41% 0m 22s ● expiring │
│ phi4-mini:latest 2.9 GB █░░░░ 23% 0m 17s ● expiring │
╰────────────────────────────── VRAM: 7.8 GB / 12.0 GB | 2 running ───────────────────────────────╯
Features
- 🔄 Real-time dashboard — live auto-refreshing view of running models
- 📋 Model listing — browse all installed models with size, family, and quantization
- 🔍 Model inspection — full details on any installed model
- 🟢 Status indicators — green / blue / red at a glance
- ⚙️ Configurable — set your API host and refresh interval
- 🪶 Lightweight — minimal dependencies, works over SSH on headless servers
- 🖥️ htop-style monitoring —
olmon topwith VRAM usage and expiry countdown - 🛑 Model control — force unload models from VRAM
- ⚖️ Model comparison — side by side spec comparison of multiple models
- 🔧 Scripting friendly —
--jsonflag and exit codes on every command
Requirements
- Python 3.13+
- Ollama installed and running
Installation
pip
pip install olmon
curl (Linux / macOS)
curl -fsSL https://raw.githubusercontent.com/glemiu6/olmon/master/scripts/install.sh | sh
From source
git clone https://github.com/glemiu6/olmon.git
cd olmon
uv pip install -e .
Windows
Native Windows binary is not currently supported.
Use WSL and follow the Linux installation instructions.
Usage
olmon status # quick health check
olmon models # list all installed models
olmon models --sort size # sort by size
olmon models --filter llama # filter by name or family
olmon inspect llama3:latest # full details on a model
olmon ps # show currently running models
olmon watch # live auto-refreshing dashboard
olmon watch --interval 5 # refresh every 5 seconds
olmon top # htop-style live monitoring
olmon stop qwen2.5:7b # unload a model from VRAM
olmon compare qwen2.5:7b llama3.2:latest # compare models
olmon --no-color models # pipe-friendly output
olmon models --json # output as JSON
Global flags
olmon --host http://192.168.1.10:11434 status # connect to remote Ollama
olmon --version # print version
Status Indicators
| Indicator | Meaning |
|---|---|
| 🟢 Green | Models are loaded and running |
| 🔵 Blue | Ollama is idle, no models loaded |
| 🔴 Red | Ollama is offline or unreachable |
Configuration
olmon init # create default config file
Config is stored at ~/.config/olmon/config.json:
{
"host": "http://localhost:11434",
"interval": 2,
"no_color": false,
"default_sort": "name"
}
Update & Uninstall
olmon update # update to latest version
olmon uninstall # remove olmon and config
Why olmon?
Most Ollama monitoring tools are GUI or system tray apps. olmon is built for:
- Headless Linux servers — no GUI required
- Remote monitoring — works over SSH
- Shell scripting — pipe-friendly with
--jsonflag and exit codes - DevOps workflows — integrate into scripts and cron jobs
GPU Support
- NVIDIA — full VRAM monitoring via
nvidia-smi - AMD — coming soon
- CPU only — VRAM stats not available
Roadmap
See ROADMAP.md for the full plan.
Contributing
Contributions are welcome. Feel free to open an issue or submit a pull request.
License
MIT — see LICENSE for details.
Made with ❤️ by Vlad Digori
Metadata
Release files for olmon 0.2.6
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| olmon-0.2.6.tar.gz | 21.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| olmon-0.2.6-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 47.6 kB
Release files / olmon-0.2.6.tar.gz
| Download URL | olmon-0.2.6.tar.gz |
|---|---|
| Size | 21.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
874e72149bfc110ff4d98faddcab55b92406b10d929f9d5a03051652152731ed
|
|
BLAKE2b-256 checksum How to use checksums |
a9c01817c410293c539693c2f3ad203f7e3a5e29f1905d3a68e4b02a38786ab5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 28, 2026.
Transparency logRelease files / olmon-0.2.6-py3-none-any.whl
| Download URL | olmon-0.2.6-py3-none-any.whl |
|---|---|
| Size | 25.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6a2f2f48df7e5efd2d4126b9e8d79fff34e552dd9027ffdf90898a52c08aede7
|
|
BLAKE2b-256 checksum How to use checksums |
91027aa28af42c6caaf1ac5013b422fb72023db0f148ca78cebcfe2ffb59b777
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 28, 2026.
Transparency log