Skip to main content

CRANE-LLM

Predict whether a ML notebook cell will crash — before you run it.

CRANE-LLM is a JupyterLab 4 / Notebook 7 extension. Select a cell, click a button, and it tells you whether executing that cell is likely to raise. What makes the prediction useful is that it does not read your code alone: it also inspects the live state of your kernel, so it knows the actual shape of the array, the dtype of the column and the length of the list the cell is about to use. The target cell is never executed.

This is the tool from the paper CRANE-LLM: Runtime-Augmented LLMs for Crash Prediction and Diagnosis in ML Notebooks.

CRANE-LLM works in two ways:

  • In JupyterLab, as a toolbar button with a sidebar. This is the full experience.
  • Anywhere else a notebook runs Python — Kaggle, Colab, VS Code, classic Notebook — as the %%crane_llm cell magic. See On Kaggle and Colab below.

Install

pip install "crane-llm[lab]"
jupyter lab

The [lab] part installs JupyterLab 4 alongside; leave it out if the environment already has JupyterLab 4: pip install "crane-llm". Check it registered:

jupyter labextension list        # expect: crane-llm-jlab <version> enabled ok

Set up a model

You need an API key for whichever model you want to use. One line in any notebook cell, remembered across sessions:

import crane_llm
crane_llm.set_api_key("sk-...")

That writes ~/.crane_llm/config.json. On Kaggle and Colab, store the key as a notebook secret instead, as described under On Kaggle and Colab below.

Settings are resolved in this order, and the first one found wins:

  1. an explicit argument, e.g. %%crane_llm gpt-5-mini
  2. the CRANE_LLM_API_KEY, CRANE_LLM_MODEL and CRANE_LLM_BASE_URL environment variables
  3. the provider's own variables, OPENAI_API_KEY and OPENAI_BASE_URL
  4. ~/.crane_llm/config.json
  5. on Kaggle and Colab, a notebook secret named CRANE_LLM_API_KEY or OPENAI_API_KEY

A .env file is not a step of its own. It is read first, from the directory you started Jupyter in or any directory above it, and fills in whichever of those variables the environment does not already define.

Using a provider other than OpenAI

Any OpenAI-compatible endpoint works by adding a base_url and a model name, which covers Claude and Gemini through a gateway, and local models with no account at all:

Provider base_url Example model
OpenAI (none needed) gpt-5
OpenRouter (Claude, Gemini, Llama, …) https://openrouter.ai/api/v1 anthropic/claude-sonnet-4.5
Google Gemini https://generativelanguage.googleapis.com/v1beta/openai/ gemini-2.5-flash
Groq https://api.groq.com/openai/v1 qwen/qwen3-32b
Ollama, local, no key needed http://localhost:11434/v1 qwen2.5-coder:32b
import crane_llm
crane_llm.set_api_key(
    "sk-or-...",
    base_url="https://openrouter.ai/api/v1",
    model="anthropic/claude-sonnet-4.5",
)

Or entirely locally:

crane_llm.set_api_key(base_url="http://localhost:11434/v1", model="qwen2.5-coder:32b")

With no base_url the OpenAI Responses API is used, which is what the paper's experiments used. Setting a base_url switches to Chat Completions, because almost no compatible server implements /responses.

Use it

  1. Run some cells as usual and let them finish.
  2. Select the cell you want to check.
  3. Click CRANE-LLM in the toolbar, or run Run CRANE-LLM from the command palette.

The verdict appears under the cell, and the right sidebar shows the full prompt and the raw response.

Colour Meaning
red a crash is certain or predicted
green no crash is predicted
amber the response could not be read as a verdict
grey the prediction has gone stale

A badge next to the verdict says where it came from:

  • Built-in check · certain: CRANE-LLM found the crash itself, from the live kernel state, and did not call the model. The cell will raise when it runs.
  • LLM prediction · model: the model judged the cell. This is a prediction, and it can be wrong.

The notebook must be idle. Reading the kernel namespace means running code in the kernel, and a kernel serves requests in order, so with cells still running the answer would describe the state afterwards rather than the one you asked about. A prediction goes stale when you edit the cell or run another one.

Two switches sit on the toolbar button (hover over it), in the sidebar and in the command palette:

  • Use the LLM. On by default. Turned off, only the built-in checks run and nothing is sent to any model, so no API key is needed. When the checks find no certain crash, the verdict says so in blue: that is not a "no crash" prediction, since the checks only report crashes they are certain of.
  • Include runtime information. Turned off, the model predicts from the code alone. That is the comparison the approach is built against, and it is also worth trying when a prompt gets too large. With it off, the built-in checks are skipped too, because they read the live kernel state. It has no effect while the LLM is off.

Built-in checks

Before calling the model, CRANE-LLM checks the cell against the live kernel state for crashes that can be detected for certain. When one of these is found, you get the answer immediately, without waiting for the model or spending tokens on it:

The cell... Raises
uses a name that is not defined NameError
reads an attribute that does not exist, including on None (df = df.dropna(inplace=True) leaves df as None) and APIs a library has removed, such as DataFrame.append or np.float AttributeError
selects, drops, groups, sorts or indexes by a DataFrame column that does not exist KeyError
reads a missing dict key, or a list, tuple or array index out of range KeyError, IndexError
assigns a list of the wrong length as a DataFrame column ValueError
multiplies, adds, stacks or reshapes NumPy arrays whose shapes do not fit ValueError
calls predict on a scikit-learn model that has not been fitted, or passes it a different number of features than it was fitted on NotFittedError, ValueError
fits a scikit-learn model, or calls train_test_split, with X and y of different lengths ValueError
divides by zero, adds incompatible types such as None + 1, unpacks the wrong number of values, or cannot be parsed ZeroDivisionError, TypeError, ValueError, SyntaxError

A check reports only crashes that are certain. It reads the cell from the top in the order Python runs it. The moment the cell would run code whose effect it cannot know, such as a call to one of your own functions, a loop, an if or a try block, it stops and leaves the cell to the model. That is why the checks catch df.head() followed by df['nope'], but not the same line inside a for loop. A check never declares a cell safe: when nothing certain is found, the model is asked exactly as before.

While a check runs, the sidebar and a box under the cell list each step as it happens: the built-in checker, then, if it found nothing certain, building the prompt and waiting for the model. A verdict from the model also says that the checker ran first and found nothing.

The checks run inside your kernel and send nothing anywhere.

Origin of the crash

The cell that crashes is rarely where the mistake was made. When a crash is found or predicted, CRANE-LLM lists under Origin of the crash the cells that gave the variables involved their current state:

  • the cell that last assigned the variable, and the line that did it,
  • every cell that modified it since, for example by dropping columns in place or fitting a model,
  • for a variable that does not exist, the cells of the notebook that define it and why they did not: not run yet, or raised before reaching the assignment. This one needs the whole notebook, so it appears in JupyterLab only, not with the cell magic.

Click an entry to jump to its cell. The cells themselves are outlined with a dashed violet line and carry a short note saying what they did, until the verdict goes stale.

This works from the order the cells actually ran in, which the notebook file does not record. CRANE-LLM records it from the moment the kernel starts, or from %load_ext crane_llm with the cell magic. Cells run before that are known from their code only, and are marked as such.

What the runtime information contains

Only the variables the target cell uses are included: every name it reads that exists in the kernel, plus the attributes and methods it uses on them.

Object What the prompt includes
int, float, str, bool the value itself, including the full text of a string
list, tuple, set length. A flat list also gets the value summary below
dict length, and depending on the contents: the metric names and epoch count of a Keras training history; the keys and data/target shapes of a scikit-learn dataset; the keys of a dict of numbers; otherwise the first 5 entries, with each value shown up to 50 characters
NumPy array shape, dtype, whether it contains NaN, minimum and maximum
pandas Series dtype, length, whether it contains NaN
pandas DataFrame shape, whether it contains NaN, and for each of the first 20 columns: dtype, number of distinct values, and either the minimum and maximum (numeric columns) or up to 5 values, each shortened to 20 characters (other columns)
PyTorch tensor shape, dtype, device, requires_grad, whether it contains NaN
PyTorch DataLoader and Subset number of batches and examples, batch size, the dataset's fields, and the shapes of its first 10 samples and of a batch built from them
TensorFlow tf.data dataset its element spec
Keras DirectoryIterator and DataFrameIterator number of samples and classes, batch size, image shape
scikit-learn estimator class, whether it has been fitted, and once fitted, the number of input features and outputs. A fitted LabelEncoder adds its number of classes
functions, methods, classes, modules the type only

Value summary:

A 1-D array, a Series or a flat list is also described by its values: binary with the two values, categorical with the number of distinct values (listed when there are 5 or fewer), or continuous with its minimum and maximum.

Some of this is your data itself: whole strings, a few values per column, dictionary entries. It is sent to the model provider along with your code, so turn runtime information off for notebooks whose data must not leave your machine. Collecting it does not change your variables.

The cell magic

Where the toolbar button is not available, put the code you want to check in a cell under %%crane_llm:

%load_ext crane_llm
%%crane_llm
model.fit(x_train, y_train)

The cell body is analysed, not executed. The verdict appears as the cell's output, in the same colours and with the same badge as above, followed by the origin of the crash and, folded under Prompt and raw response, what was sent to the model. It turns grey once you run any other cell, because the kernel state it was based on may have changed; checking another cell with %%crane_llm does not count, since nothing is executed.

%%crane_llm --no-runinfo turns runtime information off, %%crane_llm --no-llm runs only the built-in checks, and a model name overrides the configured one, as in %%crane_llm gpt-5-mini.

On Kaggle and Colab

Hosted notebooks cannot load JupyterLab extensions, so there is no toolbar button or sidebar; the cell magic does the same job.

1. Store your API key as a secret, once per account, so that it never appears in the notebook:

  • Kaggle: in the notebook editor, Add-ons → Secrets → Add a new secret, with the label CRANE_LLM_API_KEY and your key as the value. Tick the checkbox next to it in each notebook that should use it.
  • Colab: the key icon in the left sidebar, a secret named CRANE_LLM_API_KEY, with Notebook access switched on.

CRANE-LLM reads the secret itself; there is no setup cell to write.

2. Install and load it in the notebook. On Kaggle, first switch Internet on in the notebook's settings panel, which requires a phone-verified account.

%pip install crane-llm
%load_ext crane_llm

3. Run your notebook as usual, then check a cell by copying its code under %%crane_llm:

%%crane_llm
# your code (in the target cell)

Hosted sessions start from a fresh image each time, so the %pip install cell has to be run again in every new session. Competitions that require Internet to be off cannot use CRANE-LLM, since both the install and the model call need it.

Requirements and what is not included

Python 3.9+. The toolbar button needs JupyterLab 4, which crane-llm[lab] installs; the cell magic needs only IPython. Two things are deliberately left out:

  • Notebook 7. Add pip install notebook if you want that interface rather than JupyterLab.
  • The ML stack. Runtime summarisation understands pandas, numpy, torch, scikit-learn, TensorFlow objects when those packages are present in your kernel, and quietly skips them when they are not. Your own notebooks will already have brought whichever ones they use.

License

BSD 3-Clause.

Metadata

Release files for crane-llm 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for crane-llm 0.1.0
File Size Uploaded
crane_llm-0.1.0.tar.gz 173.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for crane-llm 0.1.0
File Interpreter ABI Platform
crane_llm-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 357.2 kB

Release files / crane_llm-0.1.0.tar.gz

Download URL crane_llm-0.1.0.tar.gz
Size 173.7 kB
Tags Source
SHA-256 checksum
How to use checksums
a673fa1cf70fdbada2f2d99836188df64cc30e19c265bc8b435a64ed34ff5b4b
BLAKE2b-256 checksum
How to use checksums
8ed1ac8d1751c1bf7e0667162eaecf7214ea0836feb9db44c77025751c5ced73
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release files / crane_llm-0.1.0-py3-none-any.whl

Download URL crane_llm-0.1.0-py3-none-any.whl
Size 183.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ee4065f85dfa8cecc4a5c9030923a78cec14c7f3ffe2084e0fa3f66446ef2c40
BLAKE2b-256 checksum
How to use checksums
54a35873250b5b3595dec9f86511331570a8bdc383c417f7507d37f4a09734e5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release history Release notifications | RSS feed

0.1.1

2 release files

This release

0.1.0 This release

2 release files

0.0.2

2 release files

0.0.1

2 release files

0.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page