Skip to main content

ollama_model_search

PyPI version Python License: MIT

A Python library to search and filter Ollama models by size, context length, input type, MLX support, and recency. Includes a GPU-aware system check that automatically detects your GPU VRAM and shows which models will fit.

Features

  • Scrape ollama.com — one command pulls down every model and every tag variant from ollama.com/search
  • Six-column terminal table — Tag, Size, Context, Input, MLX, and Updated, rendered with Rich
  • Powerful filtering — combine filters on size (GB), context window, input type, and MLX availability
  • GPU-aware matching — system_check() detects your GPU via GPUtil, reserves 500 MB overhead, then returns all non-MLX models that fit in usable VRAM
  • Sort by recency — newest-first or oldest-first using the relative timestamps from ollama.com
  • Programmatic access — every function returns raw Python data when table=False

Installation

pip install ollama_model_search

After installing the package, download the Chromium browser used by Playwright:

playwright install chromium

Quick Start

import ollama_model_search

# 1. Scrape the latest model catalogue from ollama.com
ollama_model_search.refresh()

# 2. List every model/tag in a table
ollama_model_search.list_all()

# 3. Find models that fit on your GPU
ollama_model_search.system_check()

API Reference

refresh()

Scrapes every public model from ollama.com/search and every tag variant from each model's tags page. Results are cached to ollama_models.json in the current working directory.

models, all_tags = ollama_model_search.refresh()
# ✓ 1600+ models discovered
# ✓ XXXX models · YYYY variants saved.

print(f"{len(models)} models, {len(all_tags)} tag variants scraped")

How it works internally:

  1. Launches a headless Chromium browser via Playwright
  2. Navigates to ollama.com/search and auto-scrolls until all models are loaded
  3. Concurrently visits every model's /tags page (25 at a time via asyncio.Semaphore)
  4. Extracts size, context, input type, MLX flag, digest, and update time from each tag row
  5. Saves everything to ollama_models.json

Returns: (models: list[str], all_tags: list[dict]) — a tuple of model names and parsed tag dicts.

You must call refresh() at least once before using any other function. All other functions read from the cached JSON.


list_all()

List all cached models in a Rich terminal table.

# Every model/tag
ollama_model_search.list_all()

# Top 10 most recently updated
ollama_model_search.list_all(recent=True, count=10)

# Return raw dicts — no table printed
results = ollama_model_search.list_all(table=False)
Parameter Type Default Description
recent bool or None None True = newest first, False = oldest first, None = leave unsorted
table bool True Print a Rich table to the terminal when True
count int or None None Limit results to the first N entries

Returns: list[dict] — the tag dicts, or [] if no cached data exists.


Filter models by any combination of size, context window, input type, and MLX support.

# Models ≤ 3.5 GB
ollama_model_search.search(size_gb=3.5)

# Models ≤ 7 GB, text-only input
ollama_model_search.search(size_gb=7, input_type="Text")

# Models with a 128K context window
ollama_model_search.search(context="128K")

# Exclude MLX models
ollama_model_search.search(is_mlx=False)

# Combine filters, sort newest-first, top 10
ollama_model_search.search(size_gb=3.5, is_mlx=False, recent=True, count=10)

# Return as raw dicts
results = ollama_model_search.search(size_gb=3.5, table=False)
Parameter Type Default Description
size_gb float or None None Maximum model size in GB (inclusive: <=)
context str or None None Case-insensitive substring match on the context window ("128K", "32K", etc.)
input_type str or None None Case-insensitive substring match on input type ("Text", "Image", etc.)
is_mlx bool or None None True = MLX only, False = exclude MLX, None = no filter
recent bool or None None True = newest first, False = oldest first, None = leave unsorted
table bool True Print a Rich table to the terminal when True
count int or None None Limit results to the first N entries

Filter behaviour: size_gb uses <= (inclusive). context and input_type use case-insensitive substring matching (e.g. "text" matches "Text & Image"). is_mlx uses exact equality.

Returns: list[dict] — the filtered tag dicts, or [] if no cached data exists.


system_check()

Detects your GPU (via GPUtil), calculates usable VRAM (total minus 500 MB overhead), and lists all non-MLX models that fit.

# Show all compatible models
ollama_model_search.system_check()

# Top 10 newest compatible models
ollama_model_search.system_check(recent=True, count=10)

# Return raw dicts
results = ollama_model_search.system_check(table=False)

Output example:

GPU detected  : NVIDIA GeForce RTX 3060
Total VRAM    : 12288 MB
Usable VRAM   : 11788 MB (11.51 GB after 500MB overhead)

[table of compatible models...]

✓ 247 compatible models found for your NVIDIA GeForce RTX 3060
Parameter Type Default Description
recent bool or None None Passed through to search() — True = newest first
table bool True Passed through to search()
count int or None None Passed through to search()

Internally calls: search(size_gb=usable_vram_gb, is_mlx=False, ...)

If no GPU is detected, a [red]No GPU detected.[/red] message is printed and an empty list is returned.


Data Format

Each tag dict returned by list_all(), search(), and system_check() contains:

Field Type Description
model_name str Parent model name (e.g. "llama3.2")
tag str Full pull tag (e.g. "llama3.2:3b-instruct-q4_K_M")
size str or None Human-readable size string (e.g. "3.5GB")
size_gb float or None Numeric size in GB — used by size_gb filter
context str or None Context window label (e.g. "128K", "32K")
input_type str or None Input modality (e.g. "Text", "Text & Image")
is_mlx bool Whether the tag supports Apple MLX
digest str or None SHA digest hash string
updated_at str or None Relative time string (e.g. "2 weeks", "3 months")
updated_at_score float Numeric recency score (lower = newer) — used internally for sorting

The cached JSON file (ollama_models.json) additionally contains:

{
  "last_updated": "2026-06-08T21:00:00.123456",
  "total_models": 1600,
  "total_tags": 12000,
  "models": ["llama3.2", "gemma3", ...],
  "tags": [ { ...tag dict... }, ... ]
}

Terminal Table Columns

When table=True (the default), the Rich table displays:

Column Source field Style
Tag tag white, no-wrap
Size size cyan, right-aligned
Context context green, centered
Input input_type magenta
MLX is_mlx green ✓ if MLX, empty otherwise
Updated updated_at yellow

Requirements

Dependency Purpose
Python 3.8+ Runtime
Playwright Headless browser automation (scraping)
BeautifulSoup4 HTML parsing
GPUtil GPU detection and VRAM querying
Rich Terminal table formatting, progress bars, and coloured output

Complete Function Summary

Function Signature Returns
refresh() no arguments tuple[list[str], list[dict]]
list_all() recent=None, table=True, count=None list[dict]
search() size_gb=None, context=None, input_type=None, is_mlx=None, recent=None, table=True, count=None list[dict]
system_check() recent=None, table=True, count=None list[dict]

License

MIT © Abhinav Abhi

Metadata

Release files for ollama-model-search 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ollama-model-search 0.1.1
File Size Uploaded
ollama_model_search-0.1.1.tar.gz 11.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ollama-model-search 0.1.1
File Interpreter ABI Platform
ollama_model_search-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 20.5 kB

Release files / ollama_model_search-0.1.1.tar.gz

Download URL ollama_model_search-0.1.1.tar.gz
Size 11.5 kB
Tags Source
SHA-256 checksum
How to use checksums
0f59fd3b33a91e3d822da8a731d6bf2955ae4c4ae2c965eeac848cb0619fd54e
BLAKE2b-256 checksum
How to use checksums
6c50637ecc8f72c7034ac2d84950c24ec3a7ecac6afead84d487ce3dcbd283ae
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.6

Release files / ollama_model_search-0.1.1-py3-none-any.whl

Download URL ollama_model_search-0.1.1-py3-none-any.whl
Size 9.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e2f92d46f3e5b7915ffe8e477ae8e97468a91ef7febc529fa47f60ef5f3dd507
BLAKE2b-256 checksum
How to use checksums
c6ffe18f5e18545f9b52556805e41ce6466ba5c5ccb24d7a25280e02e6b48da4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.13.6

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page