Skip to main content

🕵️ HCI

Hardcoded Credential Investigator

Find leaked secrets before attackers do — and get told exactly what to do about each one.

CI Self-scan License: MIT Python 3.10+ Rules


For every finding: the credential type, file and line, and a concrete remediation step. Opt-in verification checks supported credential types against their providers and reports whether they are live, invalid, or could not be verified.

hci scan running in a real terminal against a demo project with fake AWS keys, a committed Terraform state file, and a leaked Slack webhook — every finding shown with masked value, location, and remediation, then exiting non-zero to gate a CI pipeline

A real terminal recording, not a mockup — regenerate with vhs scripts/demo.tape. Static screenshot: docs/screenshot.svg.

Table of contents

Why another secrets scanner?

Mature tools like gitleaks, trufflehog, detect-secrets, and whispers already exist and are good — this project borrows deliberately from them rather than pretending they don't exist (their fingerprint/baseline/plugin designs directly shaped .hciignore, hci baseline, and the rules engine here). HCI isn't trying to out-detect them — it's built around the two things that determine whether a scanner actually gets used:

  1. Every finding tells you what to do next, not just that something matched a pattern — the masked value, the exact remediation steps for that credential type, and (opt-in) whether it's still live right now.
  2. The false-positive problem is treated as the product. Known placeholders and lockfiles are filtered by default; findings in test fixtures are lowered in severity. A per-project .hciignore and a one-command hci baseline let you silence known findings permanently instead of re-triaging them every run.

A scanner that cries wolf gets turned off — which is worse than no scanner, because the team now thinks they're covered.

Install

Install from source with Python 3.10+ and Git. PyPI publishing remains on the roadmap, so use the repository installation below:

git clone https://github.com/taylannuhogluofficial-png/hci-scan.git
cd hci-scan
python3 -m venv .venv
source .venv/bin/activate  # Windows PowerShell: .venv\Scripts\Activate.ps1
python -m pip install -e .
hci --version

Quick start

# Scan the current directory
hci scan .

# Machine-readable output for CI/tooling
hci scan . --json --fail-on high

# Also scan full git history (a secret deleted last commit is still in history)
hci scan . --history

# Check whether found credentials are still live (opt-in, read-only API calls,
# including a real SigV4-signed check for paired AWS access/secret keys)
hci scan . --verify

# Only scan what changed relative to a base branch — PR-sized, not repo-sized
hci scan . --diff main

# Counts only, no per-finding detail
hci scan . --summary

# See why a file/finding was skipped (ignored path, placeholder, baseline hit)
hci scan . --verbose

# SARIF for GitHub Code Scanning / Azure DevOps / other SARIF dashboards
hci scan . --sarif > results.sarif

# Turn findings into a ranked response plan (see below)
hci respond . --report incident.md

# See every active detection rule (content + filename-based)
hci list-rules

Exit codes: 0 means no finding at or above --fail-on (default low), 1 means the severity threshold was reached, and 2 means invalid input or a scan error. Both 1 and 2 should fail a CI gate.

The "I already committed it" workflow

Most scanners tell you a secret exists in your working tree. The moment that actually causes panic is realizing it's been sitting in git history for three months. HCI treats that as a first-class case, not an afterthought:

hci scan . --history

walks every commit's diff and flags secrets that were ever added, even if a later commit deleted them — because deleting the file doesn't remove it from history. Findings from history get commit hash, author, and date, plus a reminder that rotation (not just a new commit) is the fix.

This isn't limited to content-regex matches: a Terraform state file or default-named SSH key that existed for one commit and was deleted still gets caught by the filename rules, and a Kubernetes Secret manifest that was added and later removed still gets its data: block decoded and rescanned — each against the full file content as of that specific commit, since a diff hunk alone may not include both the kind: Secret line and the data: entry it gives meaning to.

hci respond: from "found it" to "handled it"

Finding the leak is the start of the work. hci respond turns findings into a response plan, one incident per leaked credential, however many files, commits and rules it shows up in. Incidents are ranked by urgency:

Status Meaning First move
live --verify confirmed it still authenticates revoke now
on a remote a commit containing it is reachable from a remote-tracking branch rotate, treat as public, then purge
local history only committed, but not on any remote ref rewrite history before pushing
working tree only never committed remove it, no history rewrite needed
hci respond .                              # scan workdir + full history, triage
hci respond . --verify                     # ...and check which credentials still work
hci respond . --report incident.md         # Markdown checklist for the ticket/postmortem
hci respond . --purge-plan                 # write a git filter-repo script (nothing is run)
hci respond . --import gitleaks.json       # respond to gitleaks / trufflehog findings instead

For each incident you get:

  • Exposure: the commits that added it and when it was first committed (so how many days it has been exposed), plus which branches, remote branches and tags still contain it. Run git fetch --all first, because remote reachability is based on your last fetch.
  • A provider playbook: where to revoke it, the rotation steps, and where to look for signs it was misused (CloudTrail, GitHub security log, Stripe request logs, npm view <pkg> time, and so on). The playbooks are data in hci/config/playbooks.yaml.
  • A purge plan: purge.sh plus a replacements.txt for git filter-repo. Run it with bash purge.sh [remote-url].
    • It works on a fresh mirror clone, re-scans the rewritten history with HCI, and leaves the force-push commented out.
    • Files that are leaks by themselves (Terraform state, SSH/PuTTY key files) are removed by path.
    • Private key blocks inside other files are replaced with a marker, so the rest of the file is kept. Keys are purged by their material, taken only from the keys HCI reported (keys under ignored paths, like a vendored library's test keys, are never touched): the base64 body is replaced wherever it appears, even without its BEGIN/END lines (an env var, a config value); so are the base64 of the whole key file (base64 < key.pem) and a Kubernetes Secret's encoded form; files holding the key's DER bytes are removed. Keys imported from gitleaks or trufflehog, and keys pasted into commit messages, get the same handling. Other encodings (hex dumps, byte arrays in source code) are not derived; a copy that HCI can't rewrite makes the purge stop rather than report clean only if HCI still detects it.
    • Every other value is replaced with the constant marker HCI-REDACTED, in file contents and commit messages. For a value decoded from a Kubernetes Secret, its base64 form is replaced too. The marker never depends on the secret, because the rewritten history is often public.
    • Every replacement is an ordered entry (longest value first), so a short value can't break the replacement of a longer one that contains it.
    • filter-repo can't edit binary files, so a binary file that contains a secret (or a private key) is removed from history (the script lists it, so you can restore a clean copy). Files are removed by exact path, never by prefix.
    • Before anything can be pushed, the script checks every object of the rewritten repository (file contents, binary or not, and commit messages) for every raw value — and for each half of it, so a fragment left by two overlapping values also blocks the push — and for any piece of a key's material, even re-wrapped at a different line length. This check runs as hci check-purge (not grep, whose implementations differ) and prints counts, never values. Then it re-scans the history with HCI, using the same ignore rules the plan was made with.
    • The replacements and check files necessarily hold the raw secrets. They are created exclusively (never reusing, truncating or following an existing file), owner-only (0600), in a fresh private directory (0700) that HCI creates itself, even inside a --purge-dir you give it. HCI refuses to write them inside the repository, and removes them again if anything fails part-way. The script's temporary clone is deleted on exit (KEEP_WORK=1 keeps it).
    • Paths from the repository are never trusted as script text, so a crafted filename can't inject commands into purge.sh.

hci respond always covers the whole repository, even when given a subdirectory. With --max-commits, a secret not found in the scanned commits is reported as "exposure unknown", not "working tree only". Credentials that --verify confirms are already rejected by the provider go to the bottom of the list. For those, the remaining work is to confirm the revocation and clean up history.

Already using gitleaks or trufflehog? --import reads their JSON reports. You keep your detector and add the response workflow on top. Supported credentials imported this way can still be checked with --verify.

Cutting down false positives

  • Path ignores — lockfiles, node_modules/, vendor/, binaries, and other near-always-noise paths are skipped by default (hci/config/default_ignore.yaml).
  • Placeholder detection — values containing markers like changeme, your_api_key_here, example, xxxxxxxx are filtered automatically.
  • Per-project .hciignore — add path globs or specific fingerprint:<id> lines (every finding prints its fingerprint) to silence a match for good. See .hciignore.example.
  • Keyed fingerprints — a fingerprint is an HMAC of the file, rule and secret, not a plain hash: a plain hash would let anyone holding a report or CI log test guesses of a weak password offline. hci baseline records the key in .hciignore (# hci-fingerprint-key:), so every clone and CI run, shallow or not, computes the same fingerprints. Anyone who can read the repository can read that key, so on a public repository it protects nothing: set HCI_FINGERPRINT_KEY instead (locally and as a CI secret), and hci baseline won't write the key into the file. Without any key, fingerprints are random per run and HCI tells you to run hci baseline. Fingerprints are only shown in an interactive terminal report, and only included in --json output, when they can't help someone who reads the repository: in a terminal (not a CI log) or with HCI_FINGERPRINT_KEY set. (A CI job that allocates a terminal, e.g. docker run -t, counts as interactive.) For the same reason hci baseline only fingerprints values that are committed: not findings in files git doesn't track (like .env), nor uncommitted edits. Those values never entered git, and a fingerprint next to a readable key would let anyone test guesses. Ignore such paths with a glob line, commit or remove the value, or set the key. Baselines from before 0.8 keep working; hci baseline upgrades them, warns about lines that match nothing, and --prune-stale removes those.
  • hci baseline . — accept every current finding as reviewed when you first adopt HCI on an existing repo, so day two only shows new secrets. It also records which rule set produced the baseline; if the active rules later drift from that (a rule added, removed, or its regex changed) a plain scan prints a one-line, non-blocking note suggesting you re-baseline — the same problem detect-secrets solves by recording plugin versions in its baseline file.
  • Context-aware severity — a finding inside a path that looks like tests/docs/examples/fixtures (tests/, test_*.py, docs/, fixtures/, etc.) gets its severity stepped down one notch automatically, with a note explaining why, so a fake key in a test fixture doesn't compete for attention with one in production code.
  • Grouped output — the same secret copy-pasted into five files shows up as one entry with five locations, not five repeats of identical remediation text (--no-group to disable).

Detection: regex first, entropy on top

23 built-in content rules cover AWS, GitHub, GitLab, Slack, Stripe, Google, Twilio, SendGrid, Mailgun, npm, Heroku, JWTs, private key blocks, and generic password/secret assignments — both quoted (password = "x") and the unquoted style YAML/docker-compose/Kubernetes/CI configs conventionally use (password: x), which a plain quote-requiring regex misses entirely; this is the one piece of whispers' structured-format-awareness we've adopted, done as a second regex rule rather than a full YAML/JSON parser. Rules are data, not code — hci/config/rules.yaml is a plain YAML file; add a detector by adding a name + regex + remediation entry, no Python required. Point HCI at your own rule pack with --rules your-rules.yaml to extend or override the built-ins.

Filename rules are a second, independent detector layer (hci/config/filename_rules.yaml): some files are a leak just by being committed, regardless of what any regex matches inside them — a terraform.tfstate file records full resource attributes (often plaintext database passwords, generated API keys) as plain JSON with quoted keys, a shape the quote-requiring generic rules don't parse; a .tfstate.backup; an SSH key at its default name (id_rsa/id_dsa/id_ecdsa/id_ed25519); a .ppk (PuTTY) key, whose header format the private-key-block content rule doesn't recognize at all. These fire on the filename alone and still go through the same baseline/ context-severity pipeline as content findings. See them with hci list-rules.

Pass --entropy to also flag high-Shannon-entropy strings that match no known format — catches custom tokens and random passwords. Scoring is alphabet-aware (hex vs base64 vs base58 have different "how random is random" baselines) rather than one flat threshold, which cuts down on flagging structured-but-wide-charset text; it's still opt-in and passes through the same placeholder/allowlist filters as everything else.

Blocking secrets before they're committed

hci install-hook .

installs a pre-commit git hook that scans staged files and blocks the commit if anything at medium severity or above is found. HCI is also available as a pre-commit framework hook — add to .pre-commit-config.yaml:

repos:
  - repo: https://github.com/taylannuhogluofficial-png/hci-scan
    rev: v0.8.0  # a release tag, or a reviewed commit SHA
    hooks:
      - id: hci

Run without installing Python (Docker)

docker build -t hci .
docker run --rm -v "$PWD":/scan:ro hci scan . --fail-on high

The read-only mount is suitable for scanning. Commands that write to the repository, such as baseline and install-hook, need a writable mount.

Security notes

Scanning an untrusted repository (a PR from an outside contributor, a suspicious open-source project) means processing attacker-controlled file paths, commit metadata, and file content. Two classes of issue have already been found and fixed here, and are worth knowing about if you're evaluating the tool or extending it:

  • File paths, commit author/date, and matched-secret previews are escaped before being handed to Rich's markup-parsing terminal renderer. A filename can otherwise be crafted to render as fake formatting or a spoofed clickable link, hiding or misdirecting attention away from a real finding — this was found and fixed by construction, not left to review.
  • hci baseline sanitizes control characters (including newlines, which are legal in POSIX filenames) out of any repo-derived text before writing it into .hciignore. Without this, a single crafted filename could inject arbitrary lines into the ignore file — including a bare ** glob that silently suppresses every future finding.

If you find another instance of untrusted repo content reaching a terminal, config file, or the --verify network layer without going through this kind of sanitization, please open an issue.

What HCI never reveals

Every output — terminal, --json, --sarif, respond --report, logs and error messages — is built so that holding it doesn't help anyone recover a secret:

  • No raw value and no source line, ever (outputs use an explicit allowlist of fields).
  • Masked previews hide short values completely: nothing under 16 characters, 2+2 characters up to 23, 4+4 beyond that — for provider tokens that is mostly the public type prefix (ghp_, AKIA). The length is never shown.
  • Fingerprints are keyed (see above). Incident IDs and SARIF fingerprints don't depend on the secret at all: they're derived from the rule, file and commit.
  • Passwords reveal nothing at any length.
  • Commit authors are reported by name only, never email (imported gitleaks/trufflehog reports included).
  • respond --report is written owner-only (0600) to a new file that is renamed into place, so a symlink or hard link at that path is replaced, never written through.
  • Error tracebacks never include local variables.

Known limitations: a secret that is part of a file name is printed as part of the path, and a history purge can't remove it by replacing text — rename the file in a new commit and purge with git filter-repo --path-rename. Secrets are detected in commit messages, but not in branch or tag names, not inside binary files in the working tree (history scans read every file as text, and a purge handles binary files that contain an already-found value), and not in the part of a commit message after a NUL byte (git itself stops displaying it there). History scans read every added line as text, whatever .gitattributes says and however large or binary-looking the file: any skip rule would let whoever writes a file hide a secret behind it. Files stored with Git LFS are recognised: history holds only a pointer, so hci respond reports such a finding as "committed via Git LFS" and tells you to rotate it and purge it on the LFS server, which git-filter-repo can't do.

A note on --verify

--verify makes real, read-only API calls using the credential HCI just found (a "who am I" call for each provider — sts:GetCallerIdentity for AWS, GET /user for GitHub, auth.test for Slack, etc.). It never mutates anything and never sends the secret anywhere except that provider's own API. It's opt-in for a reason: it's the one feature that reaches the network, so don't run it against a scan you don't trust, and be aware it costs a few seconds per verifiable finding.

Requests identify themselves as hci-secrets-scanner/<version>, refuse redirects (a credential is only ever sent to the provider's fixed address), and honour HTTPS_PROXY and system proxy settings. The TLS session stays end to end with the provider, so a normal proxy only sees the host name; only a proxy that intercepts TLS with a certificate authority you trust could see the credential. If the local Python has no CA certificates (python.org builds on macOS until Install Certificates.command is run), checks report "unknown" and HCI prints a warning saying why.

CI / GitHub Actions

Create .github/workflows/secrets.yml in the repository to scan:

name: Secrets scan
on: [push, pull_request]
permissions:
  contents: read
jobs:
  scan:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0  # Required to inspect the full available history.
      - uses: taylannuhogluofficial-png/hci-scan@v0.8.0
        with:
          fail-on: high
          history: "true"

Pin a release tag (as above) or, for the strictest supply-chain hygiene, a reviewed commit SHA. Check out the target repository before running the action; use fetch-depth: 0 when enabling history. The job log shows only finding counts, never values, paths or rule names: logs of a public repository are world-readable. Set upload-results: "true" to also keep the JSON findings as a workflow artifact (7 days). It is off by default because on a public repository anyone signed in to GitHub can download artifacts. With sarif: "true", results go only to the Security tab, which is access-controlled.

See action.yml for all inputs.

Native GitHub Code Scanning (SARIF)

Beyond the pass/fail build gate above, findings can show up as real annotations on a PR diff and alerts in the repo's Security → Code scanning tab — the same place CodeQL results land — instead of only existing in a log a human has to go read:

permissions:
  contents: read           # required to check out the repository
  security-events: write   # required to upload SARIF
  actions: read            # required on a private repo (upload-sarif checks workflow-run status)

steps:
  - uses: actions/checkout@v4
  - uses: taylannuhogluofficial-png/hci-scan@v0.8.0
    with:
      sarif: "true"
      fail-on: high

Or directly: hci scan . --sarif > results.sarif, then github/codeql-action/upload-sarif. No raw secret value is ever written to the SARIF file — same masked preview as everywhere else, since a SARIF file is itself a stored artifact, not just terminal output. See it live: this repo's own self-scan.yml uploads its own SARIF on every push — and, on a daily schedule, even with zero pushes. Both this and ci.yml (which reruns pip-audit daily too) carry a schedule: trigger for exactly this reason: a CVE can be disclosed for an already-pinned dependency at any time, and a push-only trigger would leave that unnoticed until someone happens to next touch the repo.

Kubernetes Secret decoding

A kind: Secret manifest's data: block is base64-encoded by Kubernetes convention — meaning every content rule here is otherwise blind to it, no matter how good the regex is, since none of them can match inside a base64 blob. HCI detects kind: Secret YAML documents, decodes each data: entry, and re-runs the same rule engine against the decoded content — so a tls.key entry gets caught by the private-key-block rule, a db-password entry by the password rules, using k8s's own field name for context. (stringData: entries need no special handling — they're already plaintext and caught by the normal per-line scan.)

apiVersion: v1
kind: Secret
metadata:
  name: db-credentials
data:
  db-password: TXlDMG1w…                  # -> flagged, decoded value shown masked
  tls.key: LS0tLS1CRUdJTi...              # -> flagged as a private key, not a blob

A ConfigMap or any other kind: is left alone — this only activates inside documents that are actually kind: Secret.

Architecture

input (files / staged / git history)
        │
        ▼
   rules engine  ──  rules.yaml (data, not code)
        │
        ▼
     scanner      ──  regex per line, + optional entropy pass
        │
        ▼
  filter layer     ──  path ignores, placeholders, .hciignore baseline
        │
        ▼
 verify (opt-in)   ──  live read-only API check per credential type
        │
        ▼
     reporter      ──  Rich terminal output, or --json

Each layer is independent, which is what let git-history scanning, entropy detection, and live verification all get added without touching the others — and is why a contributor can add a new rule or a new verifier without touching the scanner at all.

Roadmap

  • CLI: directory + git-history + --diff/--staged scanning, 23 content rules, entropy mode
  • False-positive controls: ignore paths, placeholders, baseline, context-aware severity
  • Pre-commit hook
  • GitHub Action + Dockerfile
  • Opt-in live-credential verification: GitHub, Slack, Stripe, SendGrid, npm, and real SigV4-signed AWS key-pair verification
  • Dependency vulnerability scanning in CI (pip-audit)
  • Baseline rule-set versioning (drift note when rules change since baseline)
  • Filename-based structural rules (Terraform state, default-named SSH/PuTTY keys)
  • SARIF output + native GitHub Code Scanning integration
  • Kubernetes kind: Secret data: block base64 decoding + rescanning
  • Filename rules and Kubernetes Secret decoding applied to --history, not just working-tree/staged/diff — a deleted .tfstate file or a Secret manifest removed in a later commit is still caught
  • hci respond: per-credential incidents, exposure analysis, provider playbooks, filter-repo purge plans, Markdown incident reports, gitleaks/trufflehog import
  • More verifiers (Google, Twilio, Mailgun, GitLab)
  • Multi-line secret detection (values split across lines in JSON/YAML)
  • Config-file support (hci.toml) instead of only CLI flags
  • Publish to PyPI and create versioned GitHub Action releases
  • Team dashboard / hosted monitoring (the commercial layer, later)

Contributing

Adding a detection rule is the easiest way to contribute: add an entry to hci/config/rules.yaml with a name, regex, severity, description, and remediation, plus a test. Build fake credentials from pieces at runtime (see tests/fakes.py) so the repository itself never contains a scannable one.

After the source installation above, install the development tools, the pre-commit hook (once per clone), and run every check:

python -m pip install -e '.[dev]'
make install-hooks   # every commit is scanned; a staged secret blocks it
make check           # tests, self-scan (tree + history), audit, package scan

CI runs the same checks on every push.

Use synthetic credentials in regression fixtures. The test suite checks detection, filtering, history/staging behavior, masked reports, and verifier handling without needing live credentials.

License

MIT — see LICENSE.


If a scanner cries wolf, people turn it off. That's the whole design brief.

Metadata

Release files for hci-scan 0.8.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for hci-scan 0.8.0
File Size Uploaded
hci_scan-0.8.0.tar.gz 128.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for hci-scan 0.8.0
File Interpreter ABI Platform
hci_scan-0.8.0-py3-none-any.whl Python 3 none any Details

Total release size: 217.8 kB

Release files / hci_scan-0.8.0.tar.gz

Download URL hci_scan-0.8.0.tar.gz
Size 128.8 kB
Tags Source
SHA-256 checksum
How to use checksums
9257a4a0bd8bcf6a6479cf877a635bb902098e88554b368095df51743a36de3c
BLAKE2b-256 checksum
How to use checksums
27612e6b5ea92a948be4a09f1f05c8656e592264bf4760425158e5ec770f95c8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release files / hci_scan-0.8.0-py3-none-any.whl

Download URL hci_scan-0.8.0-py3-none-any.whl
Size 89.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bbcb42f6e0f77641f960b3863c1dfa0dc398b218a855e07f9579393254ec6a6d
BLAKE2b-256 checksum
How to use checksums
45af843277eb1ec8384dc6b6570b8aa6c637e8d992f94d205b857e26106fb6ae
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.8.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page