Skip to main content

Audit pipeline that detects unsafe exception handling in LLM‑generated Python code

Project description

Silent Killers

An Exploratory Audit of Exception‑Handling in LLM‑Generated Python

CI license

tl;dr We show that large‑language models often add try/except blocks that silently swallow errors. Our AST‑based metric pipeline lets anyone quantify that risk across thousands of generated scripts in seconds.


1  Scope of this study

Modern LLMs can write Python that “runs”, but how it fails matters. A bare except: or a blanket except Exception: with no re‑raise can mask fatal bugs, leading to silent data corruption or debugging nightmares—these are the silent killers.

We collected 5 seeds × 8 models × 3 prompts (easy → hard rewrite tasks) and asked:

  • How often do models inject try/except at all?
  • Of those, how many are “bad” under a strict re‑raise rule?
  • Does difficulty exacerbate the problem?

The full paper is in docs/ (LaTeX source) and the main plots live in data/figures/.


2  Repository layout

repo-root/
├─ src/
│   └─ llm_exception_audit/        ← **reusable package**
│        ├─ __init__.py
│        ├─ metrics.py             (AST visitors & regex metrics)
│        └─ cli/
│             ├─ process_files.py
│             └─ post_processing.py
│
├─ data/                           ← study‑specific artefacts
│   ├─ propagation_prompt/
│   ├─ calibration_prompt/
│   ├─ optimization_prompt{,2}/
│   └─ figures/
├─ tests/
│   └─ test_exception_labels.py
├─ pyproject.toml
└─ README.md

Everything under src/llm_exception_audit/ is published to PyPI; data/ stays in the repo (or Git LFS) but is not shipped inside the wheel.


3  Installation

git clone https://github.com/your‑org/llm-exception-audit.git
cd llm-exception-audit
python -m pip install --upgrade pip
pip install -e .[dev]          # runtime + pytest + ruff

Requires Python ≥ 3.9
Runtime deps: pandas, numpy, matplotlib


4  Quick start

4.1  Generate metrics CSVs

process_files --base-dir data/propagation_prompt
process_files --base-dir data/calibration_prompt
process_files --base-dir data/optimization_prompt
process_files --base-dir data/optimization_prompt2

Each run creates

data/<prompt_dir>/
    llm_code_metrics.csv
    llm_response_metrics.csv

4.2  Plots & summary tables

post_processing --root data

Creates:

plots_grid_refactored/
    grid_status_3color.png
    grid_loc_continuous.png
    grid_bad_exception_rate.png
    grid_bad_exception_count.png
    bar_parsed_ok_by_difficulty.png
    summary_by_model.csv
    summary_by_difficulty.csv
Example output
code‑status bad‑rate heatmap

4.3  Library usage

from llm_exception_audit import code_metrics

python_code = "try:\n    1/0\nexcept Exception:\n    pass"
for metric in code_metrics(python_code):
    print(metric.name, metric.value)

5  Metrics at a glance

metric description
exception_handling_blocks count of except clauses
bad_exception_blocks bare except: or except Exception: without raise
bad_exception_rate bad / total, 2 dp
uses_traceback calls traceback.print_exc() / .format_exc()
see src/llm_exception_audit/metrics.py

6  Key pilot finding

When a model adds any error handling, 50–100 % of those handlers are unsafe.
Inclusive bad‑rates look tame (0 – 0.6) but conditional bad‑rates (only_with_try) spike to 1.0 for several models on simple prompts.


7  Development

ruff check .          # lint
pytest                # run unit tests
coverage run -m pytest && coverage html

CI runs on GitHub Actions across Python 3.9‑3.11 (see .github/workflows/ci.yml).


8  Roadmap

  • 🚧 dynamic execution traces (runtime errors, coverage)
  • 🚧 extend to other unsafe patterns (weak crypto, insecure I/O)
  • 🚧 publish TestPyPI wheel

PRs & issues welcome!


9  License & citation

MIT License.
If you use the metrics or figures, please cite:

@misc{Quick2025SilentKillers,
  title  = {Silent Killers: An Exploratory Audit of Exception‑Handling in LLM‑Generated Python},
  author = {Julian Quick},
  year   = {2025},
  url    = {https://github.com/your‑org/llm-exception-audit}
}

Happy auditing – don’t let silent errors slip through!


Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

silent_killers-0.1.0.tar.gz (6.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

silent_killers-0.1.0-py3-none-any.whl (5.4 kB view details)

Uploaded Python 3

File details

Details for the file silent_killers-0.1.0.tar.gz.

File metadata

  • Download URL: silent_killers-0.1.0.tar.gz
  • Upload date:
  • Size: 6.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.11.11

File hashes

Hashes for silent_killers-0.1.0.tar.gz
Algorithm Hash digest
SHA256 576499d614d6f5edc1bc467b32dea7547c5f39dafce7f998719c561fae6ed169
MD5 d80b39c98fa4baf1886e0c7e77eb5584
BLAKE2b-256 661e6afa1cb489168fca6b4e4d7f02cccfb1029e1505d1a4f544fc708259ec64

See more details on using hashes here.

File details

Details for the file silent_killers-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: silent_killers-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 5.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.11.11

File hashes

Hashes for silent_killers-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 b6079e642ccf477c00120e0fdcfb2a535fb01e7817996cee2b6053bd1606d8d9
MD5 f0b0c85b8f794c753db5ea8595ab3bcc
BLAKE2b-256 51ae3a50ca2bd7181f4f03a05156fcf309ee11957bf154753405e298e1ae1812

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page