Skip to main content

🧼 claude-code-scrubber🫧

Scrub API keys, secrets, and personal information from Claude Code transcripts before publishing them to public repos.

Inspired by Simon Willison's claude-code-transcripts — this tool sits between your raw session data and your public GitHub Pages, giving you peace of mind that you're not leaking credentials.

What it catches

Severity What's detected
🔴 High Anthropic, OpenAI, GitHub, AWS, Google, Slack, Stripe, HuggingFace, SendGrid, Twilio, npm, PyPI, Vercel, Supabase, Netlify API keys & tokens. Bearer/Basic auth headers. JWT tokens. SSH private keys. Database connection strings. Generic SECRET/TOKEN/PASSWORD assignments.
🟡 Medium Email addresses. Private IP addresses (10.x, 192.168.x, 172.16-31.x).
🔵 Low OS usernames in file paths (/Users/you/, /home/you/). Encoded Claude project paths. Shell prompts with user@host.local.

Supported formats

  • JSONL — Local Claude Code sessions from ~/.claude/projects/
  • JSON — Claude Code for web session exports
  • HTML — Output from claude-code-transcripts (preserves HTML structure)

Install

# With pip
pip install claude-code-scrubber

# With uv (no install needed)
uvx claude-code-scrubber scan session.jsonl

# From source
git clone https://github.com/yanndebray/claude-code-scrubber
cd claude-code-scrubber
pip install -e .

Quick start

# 1. Scan first (dry-run) — see what would be scrubbed
claude-code-scrubber scan my-session.jsonl -u $(whoami)

# 2. Scrub and write clean output
claude-code-scrubber scrub my-session.jsonl -u $(whoami)
# → writes my-session.scrubbed.jsonl

# 3. Scrub HTML transcripts into an output directory
claude-code-scrubber scrub transcripts/*.html -u $(whoami) -o clean/

Usage

scan — Dry-run detection

# Scan with verbose output showing every match
claude-code-scrubber scan session.jsonl -u myuser -v

# Scan only high-severity items
claude-code-scrubber scan session.jsonl -s high

# Scan multiple files
claude-code-scrubber scan *.jsonl *.html

Example output:

📄 Scanning session.jsonl ...
  🔴 [HIGH] Anthropic API key  at line 1.message.content  → sk-ant-***…
  🔴 [HIGH] AWS access key     at line 3.message.content  → AKIAIOSFO…
  🟡 [MEDIUM] Email address    at line 3.message.content  → yann.dup…
  🔵 [LOW] Home directory path  at line 2.message.content  → /Users/yan…

🧼 Found 23 item(s) to scrub across 1 file(s):
  🔴 High:   19  (API keys, tokens, passwords)
  🟡 Medium: 2   (emails, private IPs)
  🔵 Low:    2   (usernames in paths)

The scan command exits with code 1 if findings exist — useful in CI.

scrub — Redact and write clean files

# Write to a suffixed file (default: .scrubbed)
claude-code-scrubber scrub session.jsonl -u myuser
# → session.scrubbed.jsonl

# Write to an output directory
claude-code-scrubber scrub session.jsonl -o clean/

# Overwrite originals (careful!)
claude-code-scrubber scrub session.jsonl --in-place

# Custom suffix
claude-code-scrubber scrub session.jsonl --suffix .clean

init — Create a config file

claude-code-scrubber init           # creates .claude-code-scrubber.json
claude-code-scrubber init -f toml   # creates .claude-code-scrubber.toml

Configuration

Create a .claude-code-scrubber.json (or .toml) in your project root:

{
    "username": "yann",
    "severity": ["high", "medium", "low"],
    "output_suffix": ".scrubbed",
    "allowlist": [
        "sk-ant-this-is-a-dummy-key-for-docs"
    ],
    "string_replacements": {
        "MyCompanyName": "ACME",
        "my-secret-project": "project-x"
    },
    "patterns": [
        {
            "name": "Internal ticket ID",
            "regex": "PROJ-[0-9]{4,}",
            "replacement": "PROJ-XXXX",
            "severity": "medium"
        }
    ]
}
Field Description
username Your OS username, used to redact paths like /Users/you/
severity Which levels to scrub. Default: all three.
allowlist Strings that should never be redacted (false positive prevention)
string_replacements Exact string → replacement pairs, applied after regex
patterns Extra regex patterns with name, regex, replacement, severity
output_suffix Suffix for output files. Default: .scrubbed

Config files are auto-discovered by walking up from the current directory.

Workflow: Claude Code → Public GitHub

# 1. Generate HTML transcripts with Simon's tool
claude-code-transcripts local -o transcripts/

# 2. Scrub them
claude-code-scrubber scrub transcripts/*.html -u $(whoami) -o docs/

# 3. Push to GitHub Pages
git add docs/
git commit -m "Add scrubbed transcripts"
git push

CI gate (GitHub Actions)

- name: Check transcripts for secrets
  run: |
    pip install claude-code-scrubber
    claude-code-scrubber scan docs/**/*.html -u runner -s high

How it works

  1. Pattern matching — A curated set of 30+ regex patterns detect common API key formats, PII, and secrets
  2. Format-aware parsing — JSONL/JSON files are parsed and scrubbed recursively at the value level (preserving valid JSON structure). HTML is split on tags to avoid breaking markup.
  3. Layered severity — You control what gets scrubbed. Need to keep email addresses but redact API keys? Use -s high
  4. Allowlisting — Known-safe strings (like dummy keys in docs) are never touched

License

MIT

Metadata

Release files for claude-code-scrubber 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for claude-code-scrubber 0.1.0
File Size Uploaded
claude_code_scrubber-0.1.0.tar.gz 17.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for claude-code-scrubber 0.1.0
File Interpreter ABI Platform
claude_code_scrubber-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 34.7 kB

Release files / claude_code_scrubber-0.1.0.tar.gz

Download URL claude_code_scrubber-0.1.0.tar.gz
Size 17.4 kB
Tags Source
SHA-256 checksum
How to use checksums
7cf100a4becd78aada72198b6a3c8f979190b5ee8fa845d2df4e95dbca39e911
BLAKE2b-256 checksum
How to use checksums
f9b6d3fb0837b9cc94f00c342e0414e97bd6a56aee126365c1d0b4d536fa0306
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.10.6 {"installer":{"name":"uv","version":"0.10.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / claude_code_scrubber-0.1.0-py3-none-any.whl

Download URL claude_code_scrubber-0.1.0-py3-none-any.whl
Size 17.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7ff21a90826d54bdcf6b52657751cc5b56ca46f158940364ca6abc15a9dc1257
BLAKE2b-256 checksum
How to use checksums
dda5947f41ac375524f56ae9cac69528035707739662b3baccb88cdfa797d2ce
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.10.6 {"installer":{"name":"uv","version":"0.10.6","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page