Databricks logging framework for Python notebooks, including support for Spark SQL execution, with optional SFTP upload of log and HTML artifacts.
Project description
Databricks Notebook Logger & Audit Trail System
A logging system for Databricks notebooks that creates SAS-style audit trails for regulatory compliance and validation. This tool generates detailed execution logs and HTML snapshots of your notebook, then transfers them securely via SFTP to any remote directory you specify. See demo notebook and demo log file for a full example of this utilized in a notebook with further explanations.
Why This Tool?
In regulated industries (biotechnology, finance, healthcare), regulatory organizations require proof that analytical results were generated from validated code. Traditional SAS programs provide this through detailed log files that are automatically generated when the code runs. However, SAS is expensive, proprietary, and many organizations are moving toward open-source alternatives.
Databricks is traditionally used for exploratory analysis and interactive development, therefore it lacks the automatic audit logging that regulatory work requires. This tool brings SAS-style audit trails to Databricks, making it a viable platform for validated analytic workflows with regulatory compliance.
Key Benefits:
- Regulatory Compliance: Creates a complete audit trail showing exactly what code was executed, when, and what results were produced
- Reproducibility: Captures all metadata (Python/Spark/package versions, cluster info, timestamps) needed to reproduce results
- Secure Artifact Management: Automatically transfers log files and HTML notebook snapshots to your secure project repository via SFTP
- SAS-style Documentation: Mimics familiar SAS log format that regulatory reviewers expect to see
What Gets Captured
The logging system creates a comprehensive record of your notebook execution:
- Metadata: User, timestamp, cluster configuration, Python/Spark versions
- All executed Python code: Every code cell that runs in Python and it's output
- SQL queries: All queries and their results, captured from both
%sqlcells andspark.sql() - Complete outputs: Results from
print(),df.show(), and table displays - Warnings and errors: Full error tracebacks and warning messages
- Performance metrics: Runtime per cell and total execution time
- Package versions: Versions of all packages imported during the run, recorded at the end of the log
Secure Delivery: Both the text log and HTML notebook snapshot are automatically uploaded to your specified location via SFTP
Note: for a visual overview of how the logging system works on the backend, refer to databricks_log_diagram.pdf. The PDF may not render clearly in GitHub, so you will need to download it to zoom in
Quick Start
Follow these steps to enable logging in your notebook:
1. Set Notebook Language to Python
Before adding any code, ensure your notebook is set to use Python as the default language — this can be changed in the language dropdown at the top left of the notebook.
2. Import the Logging Module
Add this code to the first cell of your notebook:
%pip install nb-audit-logger
from nb_audit_logger import *
3. Start Logging
In the second cell, initialize the logger:
start_logging(output_directory="path/to/output/directory")
The output_directory filepath is the location where your audit trail artifacts (log file and HTML notebook snapshot) will be uploaded via SFTP.
When prompted:
- Enter your SFTP host and port
- Enter your SFTP credentials
- Optionally save your credentials for the cluster session duration
This establishes the secure connection and begins capturing all cells for your audit trail. If you do not specify an output_directory, SFTP upload will be skipped.
4. Add Your Code
Place all your notebook code in the cells following start_logging(). Everything will be automatically logged, including:
- Python code
- SQL queries and their results using
%sqlcells (automatically captured) - SQL queries executed using
spark.sql("INSERT SQL QUERY HERE") - Output from SQL queries displayed using
spark.sql("INSERT SQL QUERY HERE").show(truncate=False) - Spark tables displayed using
log_df()(see examples of this in the Logging SQL... section below) - Output from
print()statements - DataFrame displays using
df.show() - Warnings and errors
5. Stop Logging
In the last cell of your notebook, finalize the log:
stop_logging()
This will:
- Generate a complete audit log file in your Workspace directory (e.g.,
notebook_name.log) - Securely transfer both the log file and HTML notebook snapshot to your designated storage location via SFTP
- Display an execution summary including total runtime, warnings, and confirmation of successful artifact delivery
6. Run Your Notebook
Click Run all to execute the entire notebook. After completion, your audit trail artifacts will be available in:
- Your Workspace directory (same location as your notebook)
- Your secure storage location (uploaded via SFTP)
Note: If your notebook encounters an error during execution, the error is automatically recorded in the log file, and all artifacts are still securely transferred to ensure a complete audit trail even for failed runs.
Logging SQL Queries, SQL Output, and Spark Tables
Since the notebook language is set to Python, you need to indicate when you want to run SQL code. There are two ways to do this: %sql cells and spark.sql(). Both are supported by the logging solution and will be captured in your log. If you are just running typical SQL queries, either way works. %sql is simpler for standalone queries, while spark.sql() is better when you want to integrate with Python.
Option 1: %sql cells
Write SQL directly in a %sql cell. The query and its results are both captured in the log automatically:
%sql
SELECT * FROM schema.table_name
WHERE condition = 'value'
Option 2: spark.sql()
Use spark.sql() when you want to run SQL inside a Python cell (e.g., inside a loop or function). The query code is always captured in the log. To also capture the output, call .show():
Query only (code logged, no output in log):
spark.sql("""
SELECT * FROM schema.table_name
WHERE condition = 'value'
""")
Query + output (code and results logged):
spark.sql("""
SELECT * FROM schema.table_name
WHERE condition = 'value'
""").show(truncate=False)
Note: it is good practice to include truncate=False to ensure none of the output values are cut off.
Helper function: log_df()
Display the contents of a Spark table in both the log file and the notebook UI:
log_df("schema.table_name", limit=10)
Note: see the demo notebook for examples.
Genie Code Skill
Feel free to skip this section! It is entirely optional, and just a way to use AI to assist with using this package. Genie Code (the Databricks AI assistant) supports skills, which are files you can create to give the assistant context/instructions. A sample Genie Code skill is included in the genie_code_skill/ folder. Copy the nb_audit_logger/ subfolder into .assistant/skills/ in your Databricks workspace and Genie Code will auto-discover it. You can then ask the assistant to set up logging in your notebook or answer questions about the package.
Limitations
- Markdown cells: Cannot be captured in the log (not executed by the Python kernel)
- Cell titles: Not included in the log (part of the notebook interface, not executable code)
- Data processing and quality: Unlike SAS, Python does not natively log dataset-level operations (e.g., row/column counts, data flow between steps) or flag data quality issues (e.g., missing values, merge inconsistencies). See docs/logging_overview.md for more details
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file nb_audit_logger-0.2.0.tar.gz.
File metadata
- Download URL: nb_audit_logger-0.2.0.tar.gz
- Upload date:
- Size: 26.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
61d0d61f536976054204ab4338b94325d8bb80a7a94b2e633dda84091f9ef74e
|
|
| MD5 |
ff999d7bcb0e8077cae83366a0db1be3
|
|
| BLAKE2b-256 |
9797c5d6f88760cdb6966afd0e46a44957b2c3c9fda8f9d7cb7367b1bce9d4fb
|
File details
Details for the file nb_audit_logger-0.2.0-py3-none-any.whl.
File metadata
- Download URL: nb_audit_logger-0.2.0-py3-none-any.whl
- Upload date:
- Size: 22.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.4
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
61fd6d8c467cb14969ae1dc933f1be0f2cb0497790c6eb183155249b6b730a05
|
|
| MD5 |
e365f52a6ab995e5539d08af83f26a80
|
|
| BLAKE2b-256 |
9fa0cc1b54f845c47c1858d2d70add084ec08030d778e72acca92edf63033fa1
|