Skip to main content

owist-modelfile-lint

Static validation for Ollama Modelfiles. Catch broken FROM paths, invalid PARAMETER values, and missing TEMPLATEs before you run ollama create and get a cryptic Go error three minutes into a model build.

Built by Openwist AI, maker of the LimitAI open model family.

The problem

ollama create parses your Modelfile on the Go side and fails late:

Error: invalid file magic

That's it. No line number, no hint about which instruction caused it, and you find out only after Ollama has already started reading your (possibly multi-gigabyte) model file. A typo'd PARAMETER key gets silently ignored instead of erroring. A missing TEMPLATE on an unrecognized base model ships a model with no chat formatting at all, and you don't notice until it responds with garbage.

owist-modelfile-lint reads your Modelfile before any of that, the same way ruff or eslint check source code before you run it.

Install

pip install owist-modelfile-lint

Usage

CLI

modelfile-lint ./Modelfile
[ERROR]   line 1: FROM path './sophia-q4.gguf' does not exist  (FROM005)
[ERROR]   line 5: PARAMETER 'temprature' is not a recognized Ollama parameter (did you mean 'temperature'?)  (PARAM002)
[WARNING] line 7: PARAMETER 'temperature' value 5.7 is outside the typical range [0.0, 2.0] (default: 0.8) — this is valid syntax but likely unintentional  (PARAM007)
[WARNING] general: no TEMPLATE instruction found and FROM points to a local file/directory — Ollama's chat-template auto-detection may fail for unrecognized architectures, producing a model with no chat formatting at all. Consider adding an explicit TEMPLATE.  (TPL002)
[INFO]    line 1: estimated memory at context=8192: 4.10 GB weights + 1.00 GB KV cache + 0.75 GB overhead = ~5.85 GB total  (RES003)
[WARNING] line 1: estimated 5.85 GB exceeds detected free VRAM (4.0 GB on RTX 3060) — Ollama will offload some layers to CPU, reducing speed  (RES005)
[INFO]    line 1: estimated decode speed on RTX 3060: 43.2-75.6 tok/s (physics-based range, not a measured benchmark)  (RES008)

✗ ./Modelfile: 2 error(s), 4 warning(s)

Exit code is 0 when there are no errors (warnings don't fail the check), 1 otherwise — so it's a drop-in CI or pre-commit gate:

modelfile-lint ./Modelfile || exit 1

Other flags:

modelfile-lint ./Modelfile --quiet      # errors only, suppress warnings
modelfile-lint ./Modelfile --json       # machine-readable output
modelfile-lint ./Modelfile --no-color
modelfile-lint ./Modelfile --no-estimate  # skip memory/speed estimation (syntax-only, faster for CI)

Memory and speed estimation (new in v0.2.0)

When FROM points at a local .gguf file, owist-modelfile-lint now estimates whether the model will actually fit on your hardware and roughly how fast it'll run — before you wait through a multi-gigabyte ollama create to find out.

[INFO]    line 1: estimated memory at context=8192: 4.10 GB weights + 1.00 GB KV cache + 0.75 GB overhead = ~5.85 GB total  (RES003)
[WARNING] line 1: estimated 5.85 GB exceeds detected free VRAM (4.0 GB on RTX 3060) — Ollama will offload some layers to CPU, reducing speed  (RES005)
[INFO]    line 1: estimated decode speed on RTX 3060: 43.2-75.6 tok/s (physics-based range, not a measured benchmark)  (RES008)

Two very different confidence levels here, and the output says so explicitly:

  • Memory is a real calculation — weights size (on-disk, quantization already baked in) + KV cache (from the model's own architecture metadata and your PARAMETER num_ctx) + a documented compute-buffer allowance. Not a guess.
  • Decode speed is a physics-grounded range (memory-bandwidth-bound at batch size 1), not a fabricated precise number — no static tool can know your exact runtime conditions, and this one doesn't pretend to.

Detects NVIDIA GPUs (nvidia-smi) and Apple Silicon (sysctl) automatically; falls back to system RAM otherwise. Known gaps, stated plainly: AMD GPUs aren't detected yet, and pull-by-name models (FROM llama3.2) can't be inspected — only local .gguf file paths, since there's nothing on disk yet to read metadata from. See CHANGELOG.md for the full list.

RAM detection uses os.sysconf on Linux/macOS — no extra dependency. Windows needs one optional extra for this specific check:

pip install "owist-modelfile-lint[resources]"

Python API

from owist_modelfile_lint import lint

result = lint("Modelfile")

if not result.ok:
    for issue in result.issues:
        print(issue)
    raise SystemExit(1)
from owist_modelfile_lint import lint_text

# lint content that doesn't exist on disk yet, e.g. generated programmatically
result = lint_text("""
FROM llama3.2
PARAMETER temperature 0.7
SYSTEM You are a helpful assistant.
""")
print(result.ok)  # True

LintResult gives you .ok, .issues, .errors, .warnings, .infos, and is truthy/falsy based on .ok so if result: works too.

What it checks

Instruction Checks
FROM required and present exactly once; conventionally first; if it's a local path, the path exists; if it's a .gguf file, the magic bytes actually say GGUF; if it's a directory, it has .safetensors weights and a config.json
PARAMETER key is a real Ollama parameter (with "did you mean...?" suggestions for typos); value is the right type (int/float/string); value is in the typical sane range; duplicate non-repeatable parameters
TEMPLATE present when the base model isn't one Ollama can auto-detect; contains actual Go template variables ({{ .Prompt }}, {{ .Response }}) when present
SYSTEM not empty; warns on duplicates
ADAPTER requires a FROM; path exists; GGUF adapters are validated the same way as FROM
MESSAGE role is one of system / user / assistant; has content
structure unrecognized instructions, unterminated """ strings
resources (new in v0.2.0, local .gguf FROM targets only) estimated memory footprint at your configured context length; whether it fits detected VRAM/RAM; a physics-based decode-speed range

This is static analysis — it never loads model weights or runs Ollama. Every check, including the new resource estimation, reads only the GGUF header and metadata section — a few KB at most — never the multi-gigabyte tensor data itself, so it stays fast even on 70B-class files.

What it deliberately does not do

  • It does not validate that your TEMPLATE Go-template syntax is semantically correct for the model's actual chat format — that requires knowing what the base model expects, which is out of scope for a static linter.
  • It does not check model quality — see tinyeval (planned) for that.
  • It does not talk to the Ollama daemon or registry. Library model references like FROM llama3.2 are accepted as-is without checking whether that tag exists.
  • It does not benchmark or guarantee speed numbers — decode-speed estimates are a physics-based range (memory bandwidth ÷ model size), not a measurement. Real speed depends on kernel implementation, thermal state, and background load, none of which a static linter can see.

Why "owist"

Short for Openwist AI — we build the LimitAI open model family (Anan, Sophia) and got tired of debugging our own Modelfiles by trial and error.

License

MIT

Release files for owist-modelfile-lint 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for owist-modelfile-lint 0.2.1
File Size Uploaded
owist_modelfile_lint-0.2.1.tar.gz 27.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for owist-modelfile-lint 0.2.1
File Interpreter ABI Platform
owist_modelfile_lint-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 52.7 kB

Release files / owist_modelfile_lint-0.2.1.tar.gz

Download URL owist_modelfile_lint-0.2.1.tar.gz
Size 27.1 kB
Tags Source
SHA-256 checksum
How to use checksums
fa9db2ff9f201281a474af5453d065cca0f4c6f4ffe37b9ffc3724b7aecd814a
BLAKE2b-256 checksum
How to use checksums
a9e3a116f2b1b1edcb122c0e17f836eb1288cd9130a702d855588a00cc0e0718
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.10

Release files / owist_modelfile_lint-0.2.1-py3-none-any.whl

Download URL owist_modelfile_lint-0.2.1-py3-none-any.whl
Size 25.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5c98c6e827ccffe93798389cbb1267038a9019940e2418f9268f73ebea5f6bb5
BLAKE2b-256 checksum
How to use checksums
db3cfaef223ad5316ec9196543a421fea9d1324099627d8b8d7a0a9666f37b5a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.10

Release history Release notifications | RSS feed

This release

0.2.1 This release

2 release files

0.2.0

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page