Skip to main content

CRANE-LLM: Runtime-Augmented LLMs for Crash Prediction and Diagnosis in ML Notebooks

This is the official repository for our paper "CRANE-LLM: Runtime-Augmented LLMs for Crash Prediction and Diagnosis in ML Notebooks". In this paper, we propose CRANE-LLM, a novel approach that prompts LLMs with static code and runtime information extracted from the notebook kernel state to enhance their prediction and explanation of ML notebook crashes.

Using the CRANE-LLM notebook extension

CRANE-LLM ships as a JupyterLab 4 / Notebook 7 extension. You do not need to clone this repository or build anything to use it: the published wheel already contains the compiled frontend.

Installing

pip install crane-llm
jupyter lab

Or, to install a specific release directly from GitHub without PyPI. Take the version from the releases page and substitute it in both places:

pip install https://github.com/yarinamomo/crane_llm/releases/download/v0.0.0/crane_llm-0.0.0-py3-none-any.whl

Check that it registered:

jupyter labextension list        # expect: crane-llm-jlab <version> enabled ok

Then open a notebook, run a few cells, select the cell you want to check, and click CRANE-LLM in the toolbar. The full walkthrough is in Using the extension.

Two things the install deliberately leaves out. Add pip install notebook if you want the Notebook 7 interface rather than JupyterLab, and note that the ML stack is not pulled in: runtime summarisation covers pandas, numpy, torch, sklearn and TensorFlow objects when those packages are present in your kernel, and quietly skips them when they are not.

To work on the extension rather than use it, see the extension README, which covers the source build. To publish a new version of it, see RELEASING.md.

Setting up a model

You need an API key for whichever model you want to use. The quickest way, which is remembered across sessions, is one line in any notebook cell:

import crane_llm
crane_llm.set_api_key("sk-...")

That writes ~/.crane_llm/config.json. The extension looks for each setting in this order, and takes the first one it finds:

  1. an explicit argument, e.g. %%crane_llm gpt-5-mini
  2. the CRANE_LLM_API_KEY, CRANE_LLM_MODEL and CRANE_LLM_BASE_URL environment variables
  3. the provider's own variables, OPENAI_API_KEY and OPENAI_BASE_URL
  4. ~/.crane_llm/config.json
  5. for the model only, the default in config_llms.py

A .env file is not a step of its own. Before the lookup runs, any .env in the directory you started Jupyter from, or in a directory above it, is read and used to fill in whichever of those variables the environment does not already define. Its values are then found at step 2 or step 3, under whatever variable name they were written with. Two consequences:

  • a variable already present in the environment beats the same name in .env, because the file never overwrites something already set. This includes a persistent variable, such as a Windows user environment variable, which a Jupyter kernel inherits without your having exported anything. Keeping the same key in both places is fine; just remember that editing only the .env copy will appear to do nothing
  • OPENAI_API_KEY=... in .env beats a key in ~/.crane_llm/config.json, because it is read at step 3 and the file is step 4

Using a model other than OpenAI

Any OpenAI-compatible endpoint works by adding a base_url and a model name. This covers most providers, including Claude and Gemini through a gateway:

Provider base_url Example model
OpenAI (none needed) gpt-5
OpenRouter (Claude, Gemini, Llama, …) https://openrouter.ai/api/v1 anthropic/claude-sonnet-4.5
Google Gemini https://generativelanguage.googleapis.com/v1beta/openai/ gemini-2.5-flash
Groq https://api.groq.com/openai/v1 qwen/qwen3-32b
Ollama, local, no key needed http://localhost:11434/v1 qwen2.5-coder:32b

For example, to use Claude through OpenRouter:

import crane_llm
crane_llm.set_api_key(
    "sk-or-...",
    base_url="https://openrouter.ai/api/v1",
    model="anthropic/claude-sonnet-4.5",
)

Or a local model with no account at all:

crane_llm.set_api_key(base_url="http://localhost:11434/v1", model="qwen2.5-coder:32b")

With no base_url, CRANE-LLM calls OpenAI's Responses API, which is what the experiments in the paper used. As soon as a base_url is set it switches to Chat Completions, because almost no compatible server implements /responses. Azure OpenAI is the exception: it needs a base_url and the Responses API, so set CRANE_LLM_API_STYLE=responses there.

Paper Artefacts: Repository structure and reproducibility details

Dataset:

We use Junobench dataset in our experiments.

LLMs:

LLMs include Gemini (Gemini-2.5-Flash), Qwen (Qwen-2.5-Coder-32B-Instruct), GPT-5.

Repository structure:

All importable Python code lives under the single package crane_llm/, which is also what the installable wheel contains. The scripts at the repository root drive the experiments and are run from a checkout rather than installed. Generated experiment inputs and outputs stay in llms/ at the root, beside the code that produces them.

  • crane_llm/: the installable package
  • crane-llm.py: script to run CRANE-LLM given a target notebook
  • main_LLM.py: script to run the experiment pipeline, including batch run all notebooks in the dataset for a specific task and experimental setting
  • main.py: script to run result compilation and analysis
  • llms: LLM experiment inputs and outputs
    • llms_inputs/: generated inputs (executed code cells only, executed code cells with runtime information) to the LLMs
    • llms_outputs/: generated outputs by the LLMs
      • results_raw/: LLM outputs organized into: prefix_[LLM]_[experimental setup]_[runtime category ablation setup or API grounding]
      • ground_truth_crash_prediction.xlsx: ground truth labels used for evaluating LLM outputs as well as downstream analysis, provided by JunoBench
  • results: evaluated result outcomes and compiled statistics
    • results_parsed_detection_and_diagnosis.xlsx: CRANE-LLM performance on the joint crash prediction and diagnosis task
      • Sheet "Final_evaluation": Detailed outcomes per LLM per experimental setup. Settings include:
        • code: -RT
        • runinfo: +RT (CRANE-LLM)
      • Sheet "Results_summary": Compiled results and statistics on crash prediction and diagnosis performance of CRANE-LLM
    • results_parsed_detection_only.xlsx: CRANE-LLM performance on the crash prediction-only task
      • Sheet "Final_evaluation": Detailed outcomes per LLM per experimental setup including runtime information category ablation study, and API documentation grounding study. All settings include:
        • code: -RT
        • runinfo: +RT (CRANE-LLM)
        • runinfo_r_v: +RT-S (CRANE-LLM - S), ablated structural runtime information
        • runinfo_s_r: +RT-V (CRANE-LLM - V), ablated value semantics runtime information
        • runinfo_s_v: +RT-R (CRANE-LLM - R), ablated type-level (representation and type semantics) runtime information
        • runinfo_full_doc: +RT+doc (CRANE-LLM + doc), full runtime information with additional API documentation information
      • Sheet "Results_summary": Compiled results and statistics on crash prediction-only performance of CRANE-LLM, including runtime information category ablation study and API documentation grounding study results (and token analysis results)
    • runtime_doc_token_analysis.txt: statistics of tokens of additional API documentation information, results gained by running script token_analysis.py
    • pairwise_significance_detection_and_diagnosis.json and pairwise_significance_detection_only.json: statistical test results of the joint crash prediction and diagnosis task and the crash prediction-only task, the statistics tests are ran by script statistical_test.py
    • cohens_kappa_human_validation.txt: statistics of human evaluation on crash diagnosis outputs
    • runtime_recording/: statistics of runtime for prior cell executions and querying CRANE-LLM (when using GPT-5).

Environment

To ensure full reproducibility, we provide a docker image (digest: sha256:ecb5753d1cdfc9f0d5dfeb59818cde5be5be2f79541c5facf99761393919e171):

docker pull yarinamomo/crane_env:latest

Then run the docker container:

docker run -v [volumn_mount_windows_path]:/cranellm_env -w /cranellm_env -p 8888:8888 -it yarinamomo/crane_env:latest /bin/bash

Then you can attach this environment to VS Code "Dev Containers: Attach to Running Container..."

For the commercial LLMs used in the experiments (Gemini and GPT-5), please ensure that the API keys are properly set up before running the scripts (for example, set as global environment variable or config in .env). Open-source LLMs (Qwen) can be run directly; however, note that execution may take longer depending on the computational resources available.

License

This project is licensed under the terms of the BSD 3-Clause License.

Metadata

Release files for crane-llm 0.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for crane-llm 0.0.0
File Size Uploaded
crane_llm-0.0.0.tar.gz 118.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for crane-llm 0.0.0
File Interpreter ABI Platform
crane_llm-0.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 245.4 kB

Release files / crane_llm-0.0.0.tar.gz

Download URL crane_llm-0.0.0.tar.gz
Size 118.9 kB
Tags Source
SHA-256 checksum
How to use checksums
e003b3218dfe13488339c3a09211f5ba8af42f151d9c50c1f7d3952d080d61a1
BLAKE2b-256 checksum
How to use checksums
3e6d15295df014119cd85ced5ab0a626499dd8923b5e976e3c59a9fa185441a7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.

Transparency log

Release files / crane_llm-0.0.0-py3-none-any.whl

Download URL crane_llm-0.0.0-py3-none-any.whl
Size 126.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9be1dbabf68e022c4c1ee90ff4e179c0ffe57d139092040470baea985f715ee7
BLAKE2b-256 checksum
How to use checksums
6ee7527af61db28b2554ec875d3616d3c0a28921e3b5dd33efa5f1d664707c34
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.

Transparency log

Release history Release notifications | RSS feed

0.1.1

2 release files

0.1.0

2 release files

0.0.2

2 release files

0.0.1

2 release files

This release

0.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page