CRANE-LLM: Runtime-Augmented LLMs for Crash Prediction and Diagnosis in ML Notebooks
This is the official repository for our paper "CRANE-LLM: Runtime-Augmented LLMs for Crash Prediction and Diagnosis in ML Notebooks". In this paper, we propose CRANE-LLM, a novel approach that prompts LLMs with static code and runtime information extracted from the notebook kernel state to enhance their prediction and explanation of ML notebook crashes.
Using the CRANE-LLM notebook extension
CRANE-LLM ships as a JupyterLab 4 / Notebook 7 extension. You do not need to clone this repository or build anything to use it: the published wheel already contains the compiled frontend.
Installing
pip install crane-llm
jupyter lab
Or, to install a specific release directly from GitHub without PyPI. Take the version from the releases page and substitute it in both places:
pip install https://github.com/yarinamomo/crane_llm/releases/download/v0.0.0/crane_llm-0.0.0-py3-none-any.whl
Check that it registered:
jupyter labextension list # expect: crane-llm-jlab <version> enabled ok
Then open a notebook, run a few cells, select the cell you want to check, and click CRANE-LLM in the toolbar. The full walkthrough is in Using the extension.
Two things the install deliberately leaves out. Add pip install notebook if
you want the Notebook 7 interface rather than JupyterLab, and note that the ML
stack is not pulled in: runtime summarisation covers pandas, numpy, torch,
sklearn and TensorFlow objects when those packages are present in your kernel,
and quietly skips them when they are not.
To work on the extension rather than use it, see the extension README, which covers the source build. To publish a new version of it, see RELEASING.md.
Setting up a model
You need an API key for whichever model you want to use. The quickest way, which is remembered across sessions, is one line in any notebook cell:
import crane_llm
crane_llm.set_api_key("sk-...")
That writes ~/.crane_llm/config.json. The extension looks for each setting in
this order, and takes the first one it finds:
- an explicit argument, e.g.
%%crane_llm gpt-5-mini - the
CRANE_LLM_API_KEY,CRANE_LLM_MODELandCRANE_LLM_BASE_URLenvironment variables - the provider's own variables,
OPENAI_API_KEYandOPENAI_BASE_URL ~/.crane_llm/config.json- for the model only, the default in
config_llms.py
A .env file is not a step of its own. Before the lookup runs, any .env
in the directory you started Jupyter from, or in a directory above it, is read
and used to fill in whichever of those variables the environment does not
already define. Its values are then found at step 2 or step 3, under whatever
variable name they were written with. Two consequences:
- a variable already present in the environment beats the same name in
.env, because the file never overwrites something already set. This includes a persistent variable, such as a Windows user environment variable, which a Jupyter kernel inherits without your having exported anything. Keeping the same key in both places is fine; just remember that editing only the.envcopy will appear to do nothing OPENAI_API_KEY=...in.envbeats a key in~/.crane_llm/config.json, because it is read at step 3 and the file is step 4
Using a model other than OpenAI
Any OpenAI-compatible endpoint works by adding a base_url and a model name.
This covers most providers, including Claude and Gemini through a gateway:
| Provider | base_url |
Example model |
|---|---|---|
| OpenAI | (none needed) | gpt-5 |
| OpenRouter (Claude, Gemini, Llama, …) | https://openrouter.ai/api/v1 |
anthropic/claude-sonnet-4.5 |
| Google Gemini | https://generativelanguage.googleapis.com/v1beta/openai/ |
gemini-2.5-flash |
| Groq | https://api.groq.com/openai/v1 |
qwen/qwen3-32b |
| Ollama, local, no key needed | http://localhost:11434/v1 |
qwen2.5-coder:32b |
For example, to use Claude through OpenRouter:
import crane_llm
crane_llm.set_api_key(
"sk-or-...",
base_url="https://openrouter.ai/api/v1",
model="anthropic/claude-sonnet-4.5",
)
Or a local model with no account at all:
crane_llm.set_api_key(base_url="http://localhost:11434/v1", model="qwen2.5-coder:32b")
With no base_url, CRANE-LLM calls OpenAI's Responses API, which is what the
experiments in the paper used. As soon as a base_url is set it switches to
Chat Completions, because almost no compatible server implements /responses.
Azure OpenAI is the exception: it needs a base_url and the Responses API,
so set CRANE_LLM_API_STYLE=responses there.
Paper Artefacts: Repository structure and reproducibility details
Dataset:
We use Junobench dataset in our experiments.
LLMs:
LLMs include Gemini (Gemini-2.5-Flash), Qwen (Qwen-2.5-Coder-32B-Instruct), GPT-5.
Repository structure:
All importable Python code lives under the single package
crane_llm/, which is also what the installable wheel contains.
The scripts at the repository root drive the experiments and are run from a
checkout rather than installed. Generated experiment inputs and outputs stay in
llms/ at the root, beside the code that produces them.
crane_llm/: the installable packagenb_extension/: the JupyterLab extension, documented in its own READMEruninfo_parser/: scripts and configuration for runtime information extractionconfig_llms.py: configuration for experiments, including prompts, LLM configs, and related input and output pathsprompt_extractor.py: script for constructing prompts (i.e.,llms_inputs/)llm_executor.py: script for querying LLMs to generate outputs inllms_outputs/
crane-llm.py: script to run CRANE-LLM given a target notebookmain_LLM.py: script to run the experiment pipeline, including batch run all notebooks in the dataset for a specific task and experimental settingmain.py: script to run result compilation and analysisllms: LLM experiment inputs and outputsllms_inputs/: generated inputs (executed code cells only, executed code cells with runtime information) to the LLMsllms_outputs/: generated outputs by the LLMsresults_raw/: LLM outputs organized into: prefix_[LLM]_[experimental setup]_[runtime category ablation setup or API grounding]ground_truth_crash_prediction.xlsx: ground truth labels used for evaluating LLM outputs as well as downstream analysis, provided by JunoBench
results: evaluated result outcomes and compiled statisticsresults_parsed_detection_and_diagnosis.xlsx: CRANE-LLM performance on the joint crash prediction and diagnosis task- Sheet "Final_evaluation": Detailed outcomes per LLM per experimental setup. Settings include:
- code: -RT
- runinfo: +RT (CRANE-LLM)
- Sheet "Results_summary": Compiled results and statistics on crash prediction and diagnosis performance of CRANE-LLM
- Sheet "Final_evaluation": Detailed outcomes per LLM per experimental setup. Settings include:
results_parsed_detection_only.xlsx: CRANE-LLM performance on the crash prediction-only task- Sheet "Final_evaluation": Detailed outcomes per LLM per experimental setup including runtime information category ablation study, and API documentation grounding study. All settings include:
- code: -RT
- runinfo: +RT (CRANE-LLM)
- runinfo_r_v: +RT-S (CRANE-LLM - S), ablated structural runtime information
- runinfo_s_r: +RT-V (CRANE-LLM - V), ablated value semantics runtime information
- runinfo_s_v: +RT-R (CRANE-LLM - R), ablated type-level (representation and type semantics) runtime information
- runinfo_full_doc: +RT+doc (CRANE-LLM + doc), full runtime information with additional API documentation information
- Sheet "Results_summary": Compiled results and statistics on crash prediction-only performance of CRANE-LLM, including runtime information category ablation study and API documentation grounding study results (and token analysis results)
- Sheet "Final_evaluation": Detailed outcomes per LLM per experimental setup including runtime information category ablation study, and API documentation grounding study. All settings include:
runtime_doc_token_analysis.txt: statistics of tokens of additional API documentation information, results gained by running scripttoken_analysis.pypairwise_significance_detection_and_diagnosis.jsonandpairwise_significance_detection_only.json: statistical test results of the joint crash prediction and diagnosis task and the crash prediction-only task, the statistics tests are ran by script statistical_test.pycohens_kappa_human_validation.txt: statistics of human evaluation on crash diagnosis outputsruntime_recording/: statistics of runtime for prior cell executions and querying CRANE-LLM (when using GPT-5).
Environment
To ensure full reproducibility, we provide a docker image (digest: sha256:ecb5753d1cdfc9f0d5dfeb59818cde5be5be2f79541c5facf99761393919e171):
docker pull yarinamomo/crane_env:latest
Then run the docker container:
docker run -v [volumn_mount_windows_path]:/cranellm_env -w /cranellm_env -p 8888:8888 -it yarinamomo/crane_env:latest /bin/bash
Then you can attach this environment to VS Code "Dev Containers: Attach to Running Container..."
For the commercial LLMs used in the experiments (Gemini and GPT-5), please ensure that the API keys are properly set up before running the scripts (for example, set as global environment variable or config in .env). Open-source LLMs (Qwen) can be run directly; however, note that execution may take longer depending on the computational resources available.
License
This project is licensed under the terms of the BSD 3-Clause License.
Metadata
Release files for crane-llm 0.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| crane_llm-0.0.0.tar.gz | 118.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| crane_llm-0.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 245.4 kB
Release files / crane_llm-0.0.0.tar.gz
| Download URL | crane_llm-0.0.0.tar.gz |
|---|---|
| Size | 118.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e003b3218dfe13488339c3a09211f5ba8af42f151d9c50c1f7d3952d080d61a1
|
|
BLAKE2b-256 checksum How to use checksums |
3e6d15295df014119cd85ced5ab0a626499dd8923b5e976e3c59a9fa185441a7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.
Transparency logRelease files / crane_llm-0.0.0-py3-none-any.whl
| Download URL | crane_llm-0.0.0-py3-none-any.whl |
|---|---|
| Size | 126.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
9be1dbabf68e022c4c1ee90ff4e179c0ffe57d139092040470baea985f715ee7
|
|
BLAKE2b-256 checksum How to use checksums |
6ee7527af61db28b2554ec875d3616d3c0a28921e3b5dd33efa5f1d664707c34
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.
Transparency log