Hardware-aware LLM orchestration — zero-config local model setup for any GPU
Project description
smartpull ⚡
Stop guessing which LLM fits your GPU — SmartPull picks the best model automatically.
SmartPull profiles your GPU in real-time and recommends the best Ollama model, quantization, and context window — so you get zero VRAM crashes and maximum performance on any hardware.
🎥 Demo
Record with
asciinema rec demo.cast→ upload withasciinema upload demo.cast→ replaceREPLACE_MEwith your ID.
😫 The Problem
Running local LLMs is frustrating:
- Models crash mid-session due to VRAM overflow
- No guidance on which quantization fits your GPU
- Trial and error wastes 30+ minutes every time a new model drops
- Wrong context window = performance degradation or OOM errors
✅ The Solution
SmartPull automatically:
- Detects your GPU and real-time free VRAM
- Applies a 15% safety buffer to prevent Windows DWM swap
- Recommends the best model + quantization for your exact hardware
- Expands your context window using leftover VRAM headroom
- Generates a ready-to-use Ollama Modelfile in one command
⚡ Quickstart
pip install smartpull
smartpull build
Then copy and run the 3 ollama commands SmartPull prints. Done.
📟 Sample Output
PS C:\Users\you> smartpull build
⚙️ Running SmartPull analysis...
==================================================
smartpull — Final Recommendation
==================================================
GPU : NVIDIA GeForce RTX 3050 Ti Laptop GPU
Total VRAM : 4.0 GB
Usable VRAM : 3.29 GB (after 15% buffer)
✅ Model : qwen2.5-coder:3b
✅ Quantization : Q4_K_S
✅ VRAM needed : 2.54 GB
✅ Context window : 14,336 tokens (expanded from 8,192)
✅ Headroom : 768 MB
✅ Swap risk : LOW
✅ ELO (approx) : 1290
Note: Strong coding model. Best quality in the 3-4.5GB window.
==================================================
📝 Generating Modelfile...
✅ Modelfile written to: C:\Users\you\Modelfile
==================================================
smartpull — Run These Commands
==================================================
# 1. Pull the model
ollama pull qwen2.5-coder:3b
# 2. Create optimized build
ollama create smartpull-qwen2.5-coder-3b -f C:\Users\you\Modelfile
# 3. Run it
ollama run smartpull-qwen2.5-coder-3b
==================================================
🚀 Getting Started
Prerequisites
- Python 3.10+
- Ollama installed and running
- NVIDIA GPU with drivers installed
Install
pip install smartpull
Install from source
git clone https://github.com/punvesh/smartpull.git
cd smartpull
pip install -e ".[dev]"
Full user flow
pip install smartpull
↓
smartpull build
↓
copy 3 ollama commands
↓
model running in under 5 mins
smartpull build— scans GPU, recommends model, writes Modelfileollama pull <model>— downloads the recommended model (~1–2 GB, one time)ollama create smartpull-<model> -f ./Modelfile— creates optimized buildollama run smartpull-<model>— start chatting
🖥️ Usage
| Command | Description |
|---|---|
smartpull scan |
Scan your GPU and print real-time hardware profile |
smartpull recommend |
Recommend the best model for your current usable VRAM |
smartpull build |
Generate a Modelfile and print the recommended Ollama commands |
smartpull matrix |
Print the full model-quantization reference matrix |
⚙️ How SmartPull Works
nvidia-smi
↓
Real-time free VRAM (not total)
↓
15% safety buffer applied
↓
Model matrix lookup (VRAM → model + quant)
↓
MoE active parameter check
↓
Dynamic context window scaling
↓
Jinja2 Modelfile generation
↓
ollama create → run
- Queries
nvidia-smifor current GPU memory and driver data - Converts raw VRAM into safe usable VRAM after 15% buffer
- Matches usable VRAM to best model + quant in the SmartPull matrix
- Adjusts for MoE active parameter behavior where needed
- Expands context window when extra headroom is available
- Generates an Ollama-compatible
Modelfile
📊 Recommended Model Matrix
| Usable VRAM | Model | Quant | Safe CTX | Notes |
|---|---|---|---|---|
| < 1.8 GB | gemma2:2b |
IQ4_XS | 2,048 | Lightweight local assistant |
| 1.8–2.8 GB | gemma4:e2b |
IQ4_XS | 4,096 | Best sub-3GB balance for code |
| 2.8–4.5 GB | qwen2.5-coder:3b |
Q4_K_S | 8,192 | Strong coding performance |
| 4.5–6 GB | qwen2.5-coder:7b |
Q4_K_S | 16,384 | High context for larger workloads |
| 6–9 GB | llama3.1:8b |
Q5_K_M | 32,768 | Balanced code + general tasks |
| 9 GB+ | qwen2.5-coder:14b |
Q5_K_M | 32,768 | Premium local model experience |
📈 Performance
Benchmarked on RTX 3050 Ti Laptop GPU (4 GB VRAM) running qwen2.5-coder:3b:
| Metric | Result |
|---|---|
| VRAM usage | 2,953 / 4,096 MB — stable, zero drift |
| GPU utilization | 88–90% during inference |
| CPU swap | None — all layers in VRAM |
| Context window | 14,336 tokens (expanded by SmartPull) |
| Output quality | Production-grade code with type hints + docstrings |
Live nvidia-smi during active inference:
# gpu fb sm mem
0 2953MB 88% 48% ← generating response
0 2953MB 90% 50% ← generating response
0 2953MB 0% 0% ← response complete, idle
VRAM locked solid at 2,953 MB throughout. No swap. No crashes.
🐳 Docker
docker build -t smartpull .
docker run --gpus all smartpull recommend
🛠️ Development
pip install -e ".[dev]"
pytest tests/ -v
black .
flake8 .
Project layout:
smartpull/
├── hardware.py ← nvidia-smi scraper
├── matrix.py ← model matrix lookup
├── smartpull.py ← core smart pull logic
├── modelfile_gen.py ← Jinja2 Modelfile generator
├── cli.py ← Click CLI
├── tests/ ← pytest unit tests
├── Dockerfile
├── docker-compose.yml
└── pyproject.toml
🔁 Release & CI
SmartPull uses GitHub Actions for:
- Formatting and lint checks (
black,flake8) - Multi-version test coverage (Python 3.10, 3.11, 3.12)
- Release-triggered PyPI publishing
When a v* tag is pushed, the publish workflow builds and uploads the package automatically.
🗺️ Roadmap
- NVIDIA GPU support
- Ollama Modelfile generation
- MoE model awareness
- Dynamic context window scaling
- PyPI package
- Docker support
- GitHub Actions CI/CD
- Apple Silicon (MLX) support
- AMD ROCm support
- Claude Code bridge (auto-set
ANTHROPIC_BASE_URL) - Interactive model picker TUI
- Auto model matrix updates from remote JSON
🤝 Contributing
Contributions are welcome. If you want to add support for new hardware, extend the model matrix, or improve the CLI experience, please open a PR.
Recommended process:
- Fork the repository
- Create a feature branch
- Run tests locally
- Submit a PR with a clear description
📝 Known Notes
- Ollama may pull
Q4_K_Mby default instead ofQ4_K_S— SmartPull's ctx and GPU settings still apply correctly either way - Windows users: if
smartpullis not recognized after install, addC:\Users\<you>\AppData\Roaming\Python\Python3xx\Scriptsto your PATH nvidia-smimust be accessible — requires NVIDIA drivers installed
⭐ If SmartPull saves you time, please star the repo — it helps others find it!
License
MIT © Punvesh
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file smartpull-0.1.2.tar.gz.
File metadata
- Download URL: smartpull-0.1.2.tar.gz
- Upload date:
- Size: 16.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.15
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
3367b4e363676c6ecc2163a656108b3aa8615a2369defa944a4dbd1a7e504fe4
|
|
| MD5 |
8ccd811127c67f1ad3f641caa5a14a1e
|
|
| BLAKE2b-256 |
ef5f503b0c47e9d60d5b08b84ae605e3fd669becffaee6e6d50c45e338b3e9f9
|
File details
Details for the file smartpull-0.1.2-py3-none-any.whl.
File metadata
- Download URL: smartpull-0.1.2-py3-none-any.whl
- Upload date:
- Size: 15.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.15
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8b98bf7e708ab61767cb13a905b97dd4a1907bc4c8d0f98a69e7707036eb2a32
|
|
| MD5 |
1ac11f25216b8f1ca1615fb6c1093d49
|
|
| BLAKE2b-256 |
5c66f52cf36395ef4bfcfff30c3e695f63000dbe27cf0f69f733dafb90ec01c6
|