notebook_llm
One interactive CLI to run LLMs on a Kaggle/Colab GPU with Ollama: detect GPUs, check whether a model fits, estimate tokens/sec, search the Ollama library, pull, load, benchmark, and open a public tunnel link.
Install (Kaggle notebook: Internet ON, Accelerator = GPU)
!pip install -q notebook-llm
From a local folder (copy it to a writable place first, /kaggle/input is read-only):
!pip install -q ./notebook_llm # or: pip install notebook_llm
import notebook_llm
notebook_llm.run() # opens the interactive menu (input boxes appear in the cell)
In a real terminal just run notebook-llm.
Note: !notebook-llm cannot take keyboard input in Kaggle, so use notebook_llm.run() there.
Menu
- Quick start - installs everything, picks the best model that fits, loads it, opens the tunnel
- Search Ollama library - live search of ollama.com/search (paged with
m, filters like/tools /vision /thinking /embedding /newest, or pastemodel:tag/ a library URL). Cloud-only models are hidden because they don't run on your GPU (/cloudshows them). Pick a model to see FIT verdict and ~tok/s for every size, then pull/load - Recommended models for my GPU
- Installed models - load / unload / benchmark (real tok/s) / delete
- Tunnel - open, show link, Continue config, close, new link
- Server - start / stop / restart / logs / GPU report
- Settings - context length, KV cache type, flash attention, keep-alive (saved)
- Benchmark the active model
Non-interactive
notebook-llm gpus # hardware + recommended table
notebook-llm search coder
notebook-llm check llama3.1:70b --ctx 8192
notebook-llm quickstart --model auto
notebook-llm status | stop
Python API
from notebook_llm import NotebookLLM
llm = NotebookLLM()
url = llm.quickstart() # install -> serve -> pull -> load -> tunnel
llm.chat("hello"); llm.benchmark(); llm.close_tunnel(); llm.stop_server()
How the estimates work
- Needs = weights (exact size from registry.ollama.ai, else params x bytes/param) + KV cache (scales with context and KV type) + ~0.7 GB per GPU.
- FIT: FITS (<=92% of total VRAM), TIGHT (<=100%), SLOW (spills to RAM), TOO BIG.
- tok/s = memory bandwidth / bytes read per token, using a built-in GPU bandwidth table (T4, P100, V100, A100, L4, RTX...). MoE models only read their active experts (e.g. qwen3-coder:30b ~3.3B active) so they are much faster than dense models of the same size. Multi-GPU is layer-split, so it adds capacity, not speed.
- Estimates are +-30%. Use Benchmark for the real number.
Notes
- The tunnel link has no authentication; anyone with it can use your GPU.
- NVIDIA GPUs only. Library search scrapes ollama.com; if it is unreachable a built-in list is used.
Metadata
Release files for notebook-llm-cli 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| notebook_llm_cli-0.1.0.tar.gz | 25.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| notebook_llm_cli-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 53.1 kB
Release files / notebook_llm_cli-0.1.0.tar.gz
| Download URL | notebook_llm_cli-0.1.0.tar.gz |
|---|---|
| Size | 25.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4062a415d0c489b801a822bb5cf44e7772d86e95afce02445fc1bcdcb9619cae
|
|
BLAKE2b-256 checksum How to use checksums |
515f8450403142b73e24cbfd1487d589b661fe9a8a2dbfe18fc761dafc61fb25
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|
Release files / notebook_llm_cli-0.1.0-py3-none-any.whl
| Download URL | notebook_llm_cli-0.1.0-py3-none-any.whl |
|---|---|
| Size | 27.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2126d35b4501626a37bf6c8f78a6251d5a46a6a535d98c3057ab21f12a3886cd
|
|
BLAKE2b-256 checksum How to use checksums |
7a55207012454fc77b56c4fbe83fe3bc36dcbd24cd92316454943d5ea2301239
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.7
|