ollama-scout
One command to find the right LLMs for your hardware.
ollama-scout scans your GPU VRAM, CPU, and RAM, then recommends compatible Ollama models grouped by use case — with an interactive guided mode for new users and full CLI flags for power users.
Demo
$ ollama-scout
ollama-scout v0.2.0 | LLM Hardware Advisor (Interactive Mode)
Welcome! Let's find the best LLMs for your hardware.
Press Enter to scan your system, or Ctrl+C to exit.
System Hardware
┌─────────────────┬────────────────────────────────────────┐
│ OS │ Linux │
│ CPU │ AMD Ryzen 9 5900X │
│ Cores / Threads │ 12 cores / 24 threads │
│ RAM │ 32.0 GB │
│ GPU │ NVIDIA RTX 3080 (10.0 GB VRAM) │
└─────────────────┴────────────────────────────────────────┘
What are you mainly using this for?
1. All categories 2. Coding 3. Reasoning 4. Chat
Coding Models
Model Tag Quant Size Fit Mode Status
deepseek-coder 6.7b Q4_K_M 3.8GB Excellent GPU Available
codellama 7b Q4_K_M 3.8GB Excellent GPU Pulled
Would you like to compare two models? [y/N]:
Save results as a Markdown report? [y/N]:
Quick Start
# Recommended
pipx install ollama-scout
# Or with pip
pip install ollama-scout
# Or from source
git clone https://github.com/sandy-sp/ollama-scout.git
cd ollama-scout && pip install -e .
Then run:
ollama-scout # interactive guided mode — no flags needed
Requires Python 3.10+ and Ollama installed.
Features
| Feature | Description |
|---|---|
| Interactive guided mode | Step-by-step session with no flags needed (ollama-scout or -i) |
| Hardware detection | NVIDIA, AMD ROCm, Apple Silicon unified memory, multi-GPU |
| Live + offline models | Fetches from Ollama API with 24hr cache; --offline uses built-in list |
| Smart scoring | GPU > Multi-GPU > CPU+GPU offload > CPU-only, with time estimates |
| Use-case grouping | Coding, Reasoning, Chat with per-category tables |
| Real benchmark timing | Measures actual tokens/sec on pulled models via ollama run |
| Model comparison | --compare model1 model2 for side-by-side analysis with verdict |
| Model detail view | --model NAME shows all variants scored against your hardware |
| Markdown export | --export saves a formatted report |
| Persistent config | XDG-compliant paths with --config and --config-set |
| Auto-pull | Pull recommended models interactively or via --pull |
CLI Reference
| Flag | Description | Example |
|---|---|---|
-i, --interactive |
Launch guided mode (default with no args) | ollama-scout -i |
--use-case |
Filter by category | --use-case coding |
--flat |
Flat list instead of grouped tables | --flat |
--top N |
Limit number of results | --top 20 |
--offline |
Use built-in fallback model list | --offline |
--benchmark |
Show inference speed estimates | --benchmark |
--model NAME |
Detail view for a specific model | --model deepseek-coder |
--compare M1 M2 |
Side-by-side model comparison | --compare llama3.2 mistral |
--export |
Auto-export Markdown report | --export |
--output PATH |
Export to a specific file | --output ~/report.md |
--pull MODEL |
Pull a model via ollama | --pull llama3.2:3b |
--update-models |
Force-refresh model list cache | --update-models |
--config |
Show current configuration | --config |
--config-set K=V |
Set a config value | --config-set offline_mode=true |
--no-pull-prompt |
Skip interactive pull prompt | --no-pull-prompt |
--version |
Show version | --version |
Run ollama-scout --help for the full list. See docs/USAGE.md for detailed examples.
How Scoring Works
| Fit | Mode | Meaning |
|---|---|---|
| Excellent | GPU | Model fits fully in single GPU VRAM |
| Excellent | Multi-GPU | Model distributed across multiple GPUs |
| Good | CPU+GPU | Partially offloaded to RAM — usable but slower |
| Possible | CPU | CPU-only inference with time estimate |
| (excluded) | — | Model too large for available memory |
Apple Silicon: unified memory is treated as VRAM, so M1/M2/M3/M4 Macs get GPU-tier scoring with a 4GB system reserve.
Roadmap
- Interactive guided mode
- Real benchmark timing (
ollama run) - Model comparison mode (
--compare) - Multi-GPU support
- XDG-compliant config paths
- Model list caching (24hr TTL)
- pip installable on PyPI
- Live streaming benchmarks with progress bar
- Model search and filter by keyword
- Web UI version
- Config profiles (work / gaming / minimal)
Contributing
Contributions welcome! See CONTRIBUTING.md for setup instructions and guidelines.
License
MIT — see LICENSE.
Generated by ollama-scout
Metadata
Release files for ollama-scout 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ollama_scout-0.3.0.tar.gz | 46.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ollama_scout-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 81.7 kB
Release files / ollama_scout-0.3.0.tar.gz
| Download URL | ollama_scout-0.3.0.tar.gz |
|---|---|
| Size | 46.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1ceb626853022021d7341cffefe509f18fa70ca6d11c8cb991300ccf626df9da
|
|
BLAKE2b-256 checksum How to use checksums |
5c5f43bb6f9afac77b8c5dcbddb307a5e575e5d261d3a97442e4e22be941d948
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Feb 23, 2026.
Transparency logRelease files / ollama_scout-0.3.0-py3-none-any.whl
| Download URL | ollama_scout-0.3.0-py3-none-any.whl |
|---|---|
| Size | 35.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a1e4f10a9ad160d077b4395dbeea59d258a2ade8fa63eed69c5c7cd19d875373
|
|
BLAKE2b-256 checksum How to use checksums |
913e8341763a980bddcac51fab2b2d6cfdec93686bbf363bf6fcc7c4ea6a7c4b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Feb 23, 2026.
Transparency log