Skip to main content

notebook_llm

One interactive CLI to run LLMs on a Kaggle/Colab GPU with Ollama: detect GPUs, check whether a model fits, estimate tokens/sec, search the Ollama library, pull, load, benchmark, and open a public tunnel link.

Install (Kaggle notebook: Internet ON, Accelerator = GPU)

!pip install -q notebook-llm-cli

From a local folder (copy it to a writable place first, /kaggle/input is read-only):

!pip install -q ./notebook_llm     # or: pip install notebook-llm-cli
import notebook_llm
notebook_llm.run()                 # opens the interactive menu (input boxes appear in the cell)

In a real terminal just run notebook-llm. Note: !notebook-llm cannot take keyboard input in Kaggle, so use notebook_llm.run() there.

Menu

  1. Quick start - installs everything, picks the best model that fits, loads it, opens the tunnel
  2. Search Ollama library - live search of ollama.com/search (paged with m, filters like /tools /vision /thinking /embedding /newest, or paste model:tag / a library URL). Cloud-only models are hidden because they don't run on your GPU (/cloud shows them). Pick a model to see FIT verdict and ~tok/s for every size, then pull/load
  3. Recommended models for my GPU
  4. Installed models - load / unload / benchmark (real tok/s) / delete
  5. Tunnel - open, show link, Continue config, close, new link
  6. Server - start / stop / restart / logs / GPU report
  7. Settings - context length, KV cache type, flash attention, keep-alive (saved)
  8. Benchmark the active model

Non-interactive

notebook-llm gpus                    # hardware + recommended table
notebook-llm search coder
notebook-llm check llama3.1:70b --ctx 8192
notebook-llm quickstart --model auto
notebook-llm status | stop

Python API

from notebook_llm import NotebookLLM
llm = NotebookLLM()
url = llm.quickstart()               # install -> serve -> pull -> load -> tunnel
llm.chat("hello"); llm.benchmark(); llm.close_tunnel(); llm.stop_server()

How the estimates work

  • Needs = weights (exact size from registry.ollama.ai, else params x bytes/param) + KV cache (scales with context and KV type) + ~0.7 GB per GPU.
  • FIT: FITS (<=92% of total VRAM), TIGHT (<=100%), SLOW (spills to RAM), TOO BIG.
  • tok/s = memory bandwidth / bytes read per token, using a built-in GPU bandwidth table (T4, P100, V100, A100, L4, RTX...). MoE models only read their active experts (e.g. qwen3-coder:30b ~3.3B active) so they are much faster than dense models of the same size. Multi-GPU is layer-split, so it adds capacity, not speed.
  • Estimates are +-30%. Use Benchmark for the real number.

Notes

  • The tunnel link has no authentication; anyone with it can use your GPU.
  • NVIDIA GPUs only. Library search scrapes ollama.com; if it is unreachable a built-in list is used.

Metadata

Release files for notebook-llm-cli 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for notebook-llm-cli 0.1.1
File Size Uploaded
notebook_llm_cli-0.1.1.tar.gz 25.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for notebook-llm-cli 0.1.1
File Interpreter ABI Platform
notebook_llm_cli-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 53.1 kB

Release files / notebook_llm_cli-0.1.1.tar.gz

Download URL notebook_llm_cli-0.1.1.tar.gz
Size 25.5 kB
Tags Source
SHA-256 checksum
How to use checksums
bc7869924aac8b10ac679726ac83c99d9c0cd2b2bc1609dffdac394f6cd024fb
BLAKE2b-256 checksum
How to use checksums
c03735fda7d19f33a58b65b9ed64ae752736a65018e9b9094af8bd815d72a0d5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release files / notebook_llm_cli-0.1.1-py3-none-any.whl

Download URL notebook_llm_cli-0.1.1-py3-none-any.whl
Size 27.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bda9c3c9d38cca27762c6493b1f0eb7944a541b9778562351e747228eddac045
BLAKE2b-256 checksum
How to use checksums
8b789e39c1f2138a67f757e92067926d940968b2d3d1df5fd250c389c2b97a05
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page