Skip to main content

ollama-scout

Python 3.10+ PyPI License: MIT Platform CI

One command to find the right LLMs for your hardware.

ollama-scout scans your GPU VRAM, CPU, and RAM, then recommends compatible Ollama models grouped by use case — with an interactive guided mode for new users and full CLI flags for power users.


Demo

$ ollama-scout

  ollama-scout v0.2.0  |  LLM Hardware Advisor  (Interactive Mode)

  Welcome! Let's find the best LLMs for your hardware.
  Press Enter to scan your system, or Ctrl+C to exit.

                   System Hardware
  ┌─────────────────┬────────────────────────────────────────┐
  │ OS              │ Linux                                  │
  │ CPU             │ AMD Ryzen 9 5900X                      │
  │ Cores / Threads │ 12 cores / 24 threads                  │
  │ RAM             │ 32.0 GB                                │
  │ GPU             │ NVIDIA RTX 3080 (10.0 GB VRAM)         │
  └─────────────────┴────────────────────────────────────────┘

  What are you mainly using this for?
    1. All categories    2. Coding    3. Reasoning    4. Chat

  Coding Models
  Model             Tag    Quant    Size    Fit        Mode     Status
  deepseek-coder    6.7b   Q4_K_M   3.8GB   Excellent  GPU      Available
  codellama         7b     Q4_K_M   3.8GB   Excellent  GPU      Pulled

  Would you like to compare two models? [y/N]:
  Save results as a Markdown report? [y/N]:

Quick Start

# Recommended
pipx install ollama-scout

# Or with pip
pip install ollama-scout

# Or from source
git clone https://github.com/sandy-sp/ollama-scout.git
cd ollama-scout && pip install -e .

Then run:

ollama-scout          # interactive guided mode — no flags needed

Requires Python 3.10+ and Ollama installed.


Features

Feature Description
Interactive guided mode Step-by-step session with no flags needed (ollama-scout or -i)
Hardware detection NVIDIA, AMD ROCm, Apple Silicon unified memory, multi-GPU
Live + offline models Fetches from Ollama API with 24hr cache; --offline uses built-in list
Smart scoring GPU > Multi-GPU > CPU+GPU offload > CPU-only, with time estimates
Use-case grouping Coding, Reasoning, Chat with per-category tables
Real benchmark timing Measures actual tokens/sec on pulled models via ollama run
Model comparison --compare model1 model2 for side-by-side analysis with verdict
Model detail view --model NAME shows all variants scored against your hardware
Markdown export --export saves a formatted report
Persistent config XDG-compliant paths with --config and --config-set
Auto-pull Pull recommended models interactively or via --pull

CLI Reference

Flag Description Example
-i, --interactive Launch guided mode (default with no args) ollama-scout -i
--use-case Filter by category --use-case coding
--flat Flat list instead of grouped tables --flat
--top N Limit number of results --top 20
--offline Use built-in fallback model list --offline
--benchmark Show inference speed estimates --benchmark
--model NAME Detail view for a specific model --model deepseek-coder
--compare M1 M2 Side-by-side model comparison --compare llama3.2 mistral
--export Auto-export Markdown report --export
--output PATH Export to a specific file --output ~/report.md
--pull MODEL Pull a model via ollama --pull llama3.2:3b
--update-models Force-refresh model list cache --update-models
--config Show current configuration --config
--config-set K=V Set a config value --config-set offline_mode=true
--no-pull-prompt Skip interactive pull prompt --no-pull-prompt
--version Show version --version

Run ollama-scout --help for the full list. See docs/USAGE.md for detailed examples.


How Scoring Works

Fit Mode Meaning
Excellent GPU Model fits fully in single GPU VRAM
Excellent Multi-GPU Model distributed across multiple GPUs
Good CPU+GPU Partially offloaded to RAM — usable but slower
Possible CPU CPU-only inference with time estimate
(excluded) — Model too large for available memory

Apple Silicon: unified memory is treated as VRAM, so M1/M2/M3/M4 Macs get GPU-tier scoring with a 4GB system reserve.


Roadmap

  • Interactive guided mode
  • Real benchmark timing (ollama run)
  • Model comparison mode (--compare)
  • Multi-GPU support
  • XDG-compliant config paths
  • Model list caching (24hr TTL)
  • pip installable on PyPI
  • Live streaming benchmarks with progress bar
  • Model search and filter by keyword
  • Web UI version
  • Config profiles (work / gaming / minimal)

Contributing

Contributions welcome! See CONTRIBUTING.md for setup instructions and guidelines.

License

MIT — see LICENSE.


Generated by ollama-scout

Metadata

Release files for ollama-scout 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ollama-scout 0.3.0
File Size Uploaded
ollama_scout-0.3.0.tar.gz 46.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ollama-scout 0.3.0
File Interpreter ABI Platform
ollama_scout-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 81.7 kB

Release files / ollama_scout-0.3.0.tar.gz

Download URL ollama_scout-0.3.0.tar.gz
Size 46.5 kB
Tags Source
SHA-256 checksum
How to use checksums
1ceb626853022021d7341cffefe509f18fa70ca6d11c8cb991300ccf626df9da
BLAKE2b-256 checksum
How to use checksums
5c5f43bb6f9afac77b8c5dcbddb307a5e575e5d261d3a97442e4e22be941d948
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Feb 23, 2026.

Transparency log

Release files / ollama_scout-0.3.0-py3-none-any.whl

Download URL ollama_scout-0.3.0-py3-none-any.whl
Size 35.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a1e4f10a9ad160d077b4395dbeea59d258a2ade8fa63eed69c5c7cd19d875373
BLAKE2b-256 checksum
How to use checksums
913e8341763a980bddcac51fab2b2d6cfdec93686bbf363bf6fcc7c4ea6a7c4b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Feb 23, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page