🕵️ HCI
Hardcoded Credential Investigator
Find leaked secrets before attackers do — and get told exactly what to do about each one.
For every finding: the credential type, file and line, and a concrete remediation step. Opt-in verification checks supported credential types against their providers and reports whether they are live, invalid, or could not be verified.
A real terminal recording, not a mockup — regenerate with vhs scripts/demo.tape.
Static screenshot: docs/screenshot.svg.
Table of contents
- Why another secrets scanner?
- Install
- Quick start
- The "I already committed it" workflow
- Cutting down false positives
- Detection: regex first, entropy on top
- Blocking secrets before they're committed
- Run without installing Python (Docker)
- Security notes
- A note on
--verify - CI / GitHub Actions
- Kubernetes Secret decoding
- Architecture
- Roadmap
- Contributing
- License
Why another secrets scanner?
Mature tools like gitleaks,
trufflehog,
detect-secrets, and
whispers already exist and are
good — this project borrows deliberately from them rather than pretending
they don't exist (their fingerprint/baseline/plugin designs directly shaped
.hciignore, hci baseline, and the rules engine here). HCI isn't trying
to out-detect them — it's built around the two
things that determine whether a scanner actually gets used:
- Every finding tells you what to do next, not just that something matched a pattern — the masked value, the exact remediation steps for that credential type, and (opt-in) whether it's still live right now.
- The false-positive problem is treated as the product. Known placeholders
and lockfiles are filtered by default; findings in test fixtures are
lowered in severity. A per-project
.hciignoreand a one-commandhci baselinelet you silence known findings permanently instead of re-triaging them every run.
A scanner that cries wolf gets turned off — which is worse than no scanner, because the team now thinks they're covered.
Install
Install from source with Python 3.10+ and Git. PyPI publishing remains on the roadmap, so use the repository installation below:
git clone https://github.com/taylannuhogluofficial-png/hci-scan.git
cd hci-scan
python3 -m venv .venv
source .venv/bin/activate # Windows PowerShell: .venv\Scripts\Activate.ps1
python -m pip install -e .
hci --version
Quick start
# Scan the current directory
hci scan .
# Machine-readable output for CI/tooling
hci scan . --json --fail-on high
# Also scan full git history (a secret deleted last commit is still in history)
hci scan . --history
# Check whether found credentials are still live (opt-in, read-only API calls,
# including a real SigV4-signed check for paired AWS access/secret keys)
hci scan . --verify
# Only scan what changed relative to a base branch — PR-sized, not repo-sized
hci scan . --diff main
# Counts only, no per-finding detail
hci scan . --summary
# See why a file/finding was skipped (ignored path, placeholder, baseline hit)
hci scan . --verbose
# SARIF for GitHub Code Scanning / Azure DevOps / other SARIF dashboards
hci scan . --sarif > results.sarif
# Turn findings into a ranked response plan (see below)
hci respond . --report incident.md
# See every active detection rule (content + filename-based)
hci list-rules
Exit codes: 0 means no finding at or above --fail-on (default low),
1 means the severity threshold was reached, and 2 means invalid input
or a scan error. Both 1 and 2 should fail a CI gate.
The "I already committed it" workflow
Most scanners tell you a secret exists in your working tree. The moment that actually causes panic is realizing it's been sitting in git history for three months. HCI treats that as a first-class case, not an afterthought:
hci scan . --history
walks every commit's diff and flags secrets that were ever added, even if a later commit deleted them — because deleting the file doesn't remove it from history. Findings from history get commit hash, author, and date, plus a reminder that rotation (not just a new commit) is the fix.
This isn't limited to content-regex matches: a Terraform state file or
default-named SSH key that existed for one commit and was deleted still
gets caught by the filename rules, and a Kubernetes Secret manifest that
was added and later removed still gets its data: block decoded and
rescanned — each against the full file content as of that specific commit,
since a diff hunk alone may not include both the kind: Secret line and
the data: entry it gives meaning to.
hci respond: from "found it" to "handled it"
Finding the leak is the start of the work. hci respond turns findings
into a response plan, one incident per leaked credential, however many
files, commits and rules it shows up in. Incidents are ranked by urgency:
| Status | Meaning | First move |
|---|---|---|
| live | --verify confirmed it still authenticates |
revoke now |
| on a remote | a commit containing it is reachable from a remote-tracking branch | rotate, treat as public, then purge |
| local history only | committed, but not on any remote ref | rewrite history before pushing |
| working tree only | never committed | remove it, no history rewrite needed |
hci respond . # scan workdir + full history, triage
hci respond . --verify # ...and check which credentials still work
hci respond . --report incident.md # Markdown checklist for the ticket/postmortem
hci respond . --purge-plan # write a git filter-repo script (nothing is run)
hci respond . --import gitleaks.json # respond to gitleaks / trufflehog findings instead
For each incident you get:
- Exposure: the commits that added it and when it was first committed
(so how many days it has been exposed), plus which branches, remote
branches and tags still contain it. Run
git fetch --allfirst, because remote reachability is based on your last fetch. - A provider playbook: where to revoke it, the rotation steps, and
where to look for signs it was misused (CloudTrail, GitHub security log,
Stripe request logs,
npm view <pkg> time, and so on). The playbooks are data inhci/config/playbooks.yaml. - A purge plan:
purge.shplus areplacements.txtfor git filter-repo. Run it withbash purge.sh [remote-url].- It works on a fresh mirror clone, re-scans the rewritten history with HCI, and leaves the force-push commented out.
- Files that are leaks by themselves (Terraform state, SSH/PuTTY key files) are removed by path.
- Private key blocks inside other files are replaced with a marker, so the
rest of the file is kept. Keys are purged by their material, taken
only from the keys HCI reported (keys under ignored paths, like a
vendored library's test keys, are never touched): the base64 body is
replaced wherever it appears, even without its BEGIN/END lines (an env
var, a config value); so are the base64 of the whole key file
(
base64 < key.pem) and a Kubernetes Secret's encoded form; files holding the key's DER bytes are removed. Keys imported from gitleaks or trufflehog, and keys pasted into commit messages, get the same handling. Other encodings (hex dumps, byte arrays in source code) are not derived; a copy that HCI can't rewrite makes the purge stop rather than report clean only if HCI still detects it. - Every other value is replaced with the constant marker
HCI-REDACTED, in file contents and commit messages. For a value decoded from a Kubernetes Secret, its base64 form is replaced too. The marker never depends on the secret, because the rewritten history is often public. - Every replacement is an ordered entry (longest value first), so a short value can't break the replacement of a longer one that contains it.
- filter-repo can't edit binary files, so a binary file that contains a secret (or a private key) is removed from history (the script lists it, so you can restore a clean copy). Files are removed by exact path, never by prefix.
- Before anything can be pushed, the script checks every object of
the rewritten repository (file contents, binary or not, and commit
messages) for every raw value — and for each half of it, so a fragment
left by two overlapping values also blocks the push — and for any piece
of a key's material, even re-wrapped at a different line length. This
check runs as
hci check-purge(not grep, whose implementations differ) and prints counts, never values. Then it re-scans the history with HCI, using the same ignore rules the plan was made with. - The replacements and check files necessarily hold the raw secrets. They
are created exclusively (never reusing, truncating or following an
existing file), owner-only (0600), in a fresh private directory (0700)
that HCI creates itself, even inside a
--purge-diryou give it. HCI refuses to write them inside the repository, and removes them again if anything fails part-way. The script's temporary clone is deleted on exit (KEEP_WORK=1keeps it). - Paths from the repository are never trusted as script text, so a crafted
filename can't inject commands into
purge.sh.
hci respond always covers the whole repository, even when given a
subdirectory. With --max-commits, a secret not found in the scanned
commits is reported as "exposure unknown", not "working tree only".
Credentials that --verify confirms are already rejected by the provider
go to the bottom of the list. For those, the remaining work is to confirm
the revocation and clean up history.
Already using gitleaks or trufflehog? --import reads their JSON reports.
You keep your detector and add the response workflow on top. Supported
credentials imported this way can still be checked with --verify.
Cutting down false positives
- Path ignores — lockfiles,
node_modules/,vendor/, binaries, and other near-always-noise paths are skipped by default (hci/config/default_ignore.yaml). - Placeholder detection — values containing markers like
changeme,your_api_key_here,example,xxxxxxxxare filtered automatically. - Per-project
.hciignore— add path globs or specificfingerprint:<id>lines (every finding prints its fingerprint) to silence a match for good. See.hciignore.example. - Keyed fingerprints — a fingerprint is an HMAC of the file, rule and
secret, not a plain hash: a plain hash would let anyone holding a report
or CI log test guesses of a weak password offline.
hci baselinerecords the key in.hciignore(# hci-fingerprint-key:), so every clone and CI run, shallow or not, computes the same fingerprints. Anyone who can read the repository can read that key, so on a public repository it protects nothing: setHCI_FINGERPRINT_KEYinstead (locally and as a CI secret), andhci baselinewon't write the key into the file. Without any key, fingerprints are random per run and HCI tells you to runhci baseline. Fingerprints are only shown in an interactive terminal report, and only included in--jsonoutput, when they can't help someone who reads the repository: in a terminal (not a CI log) or withHCI_FINGERPRINT_KEYset. (A CI job that allocates a terminal, e.g.docker run -t, counts as interactive.) For the same reasonhci baselineonly fingerprints values that are committed: not findings in files git doesn't track (like.env), nor uncommitted edits. Those values never entered git, and a fingerprint next to a readable key would let anyone test guesses. Ignore such paths with a glob line, commit or remove the value, or set the key. Baselines from before 0.8 keep working;hci baselineupgrades them, warns about lines that match nothing, and--prune-staleremoves those. hci baseline .— accept every current finding as reviewed when you first adopt HCI on an existing repo, so day two only shows new secrets. It also records which rule set produced the baseline; if the active rules later drift from that (a rule added, removed, or its regex changed) a plain scan prints a one-line, non-blocking note suggesting you re-baseline — the same problem detect-secrets solves by recording plugin versions in its baseline file.- Context-aware severity — a finding inside a path that looks like
tests/docs/examples/fixtures (
tests/,test_*.py,docs/,fixtures/, etc.) gets its severity stepped down one notch automatically, with a note explaining why, so a fake key in a test fixture doesn't compete for attention with one in production code. - Grouped output — the same secret copy-pasted into five files shows up
as one entry with five locations, not five repeats of identical
remediation text (
--no-groupto disable).
Detection: regex first, entropy on top
23 built-in content rules cover AWS, GitHub, GitLab, Slack, Stripe, Google, Twilio,
SendGrid, Mailgun, npm, Heroku, JWTs, private key blocks, and generic
password/secret assignments — both quoted (password = "x") and the
unquoted style YAML/docker-compose/Kubernetes/CI configs conventionally use
(password: x), which a plain quote-requiring regex misses entirely; this
is the one piece of whispers'
structured-format-awareness we've adopted, done as a second regex rule
rather than a full YAML/JSON parser. Rules are data, not code —
hci/config/rules.yaml is a plain YAML file; add a
detector by adding a name + regex + remediation entry, no Python
required. Point HCI at your own rule pack with --rules your-rules.yaml to
extend or override the built-ins.
Filename rules are a second, independent detector layer
(hci/config/filename_rules.yaml): some
files are a leak just by being committed, regardless of what any regex
matches inside them — a terraform.tfstate file records full resource
attributes (often plaintext database passwords, generated API keys) as
plain JSON with quoted keys, a shape the quote-requiring generic rules
don't parse; a .tfstate.backup; an SSH key at its default name
(id_rsa/id_dsa/id_ecdsa/id_ed25519); a .ppk (PuTTY) key, whose
header format the private-key-block content rule doesn't recognize at all.
These fire on the filename alone and still go through the same baseline/
context-severity pipeline as content findings. See them with
hci list-rules.
Pass --entropy to also flag high-Shannon-entropy strings that match no
known format — catches custom tokens and random passwords. Scoring is
alphabet-aware (hex vs base64 vs base58 have different "how random is
random" baselines) rather than one flat threshold, which cuts down on
flagging structured-but-wide-charset text; it's still opt-in and passes
through the same placeholder/allowlist filters as everything else.
Blocking secrets before they're committed
hci install-hook .
installs a pre-commit git hook that scans staged files and blocks the
commit if anything at medium severity or above is found. HCI is also
available as a pre-commit framework hook — add to
.pre-commit-config.yaml:
repos:
- repo: https://github.com/taylannuhogluofficial-png/hci-scan
rev: v0.8.0 # a release tag, or a reviewed commit SHA
hooks:
- id: hci
Run without installing Python (Docker)
docker build -t hci .
docker run --rm -v "$PWD":/scan:ro hci scan . --fail-on high
The read-only mount is suitable for scanning. Commands that write to the
repository, such as baseline and install-hook, need a writable mount.
Security notes
Scanning an untrusted repository (a PR from an outside contributor, a suspicious open-source project) means processing attacker-controlled file paths, commit metadata, and file content. Two classes of issue have already been found and fixed here, and are worth knowing about if you're evaluating the tool or extending it:
- File paths, commit author/date, and matched-secret previews are escaped before being handed to Rich's markup-parsing terminal renderer. A filename can otherwise be crafted to render as fake formatting or a spoofed clickable link, hiding or misdirecting attention away from a real finding — this was found and fixed by construction, not left to review.
hci baselinesanitizes control characters (including newlines, which are legal in POSIX filenames) out of any repo-derived text before writing it into.hciignore. Without this, a single crafted filename could inject arbitrary lines into the ignore file — including a bare**glob that silently suppresses every future finding.
If you find another instance of untrusted repo content reaching a terminal,
config file, or the --verify network layer without going through this
kind of sanitization, please open an issue.
What HCI never reveals
Every output — terminal, --json, --sarif, respond --report, logs and
error messages — is built so that holding it doesn't help anyone recover a
secret:
- No raw value and no source line, ever (outputs use an explicit allowlist of fields).
- Masked previews hide short values completely: nothing under 16
characters, 2+2 characters up to 23, 4+4 beyond that — for provider
tokens that is mostly the public type prefix (
ghp_,AKIA). The length is never shown. - Fingerprints are keyed (see above). Incident IDs and SARIF fingerprints don't depend on the secret at all: they're derived from the rule, file and commit.
- Passwords reveal nothing at any length.
- Commit authors are reported by name only, never email (imported gitleaks/trufflehog reports included).
respond --reportis written owner-only (0600) to a new file that is renamed into place, so a symlink or hard link at that path is replaced, never written through.- Error tracebacks never include local variables.
Known limitations: a secret that is part of a file name is printed as
part of the path, and a history purge can't remove it by replacing text —
rename the file in a new commit and purge with git filter-repo --path-rename. Secrets are detected in commit messages, but not in branch
or tag names, not inside binary files in the working tree (history scans
read every file as text, and a purge handles binary files that contain an
already-found value), and not in the part of a commit
message after a NUL byte (git itself stops displaying it there). History
scans read every added line as text, whatever .gitattributes says and
however large or binary-looking the file: any skip rule would let whoever
writes a file hide a secret behind it. Files stored with Git LFS are
recognised: history holds only a pointer, so hci respond reports such a
finding as "committed via Git LFS" and tells you to rotate it and purge it
on the LFS server, which git-filter-repo can't do.
A note on --verify
--verify makes real, read-only API calls using the credential HCI just
found (a "who am I" call for each provider — sts:GetCallerIdentity for AWS,
GET /user for GitHub, auth.test for Slack, etc.). It never mutates
anything and never sends the secret anywhere except that provider's own API.
It's opt-in for a reason: it's the one feature that reaches the network, so
don't run it against a scan you don't trust, and be aware it costs a few
seconds per verifiable finding.
Requests identify themselves as hci-secrets-scanner/<version>, refuse
redirects (a credential is only ever sent to the provider's fixed address),
and honour HTTPS_PROXY and system proxy settings. The TLS session stays
end to end with the provider, so a normal proxy only sees the host name;
only a proxy that intercepts TLS with a certificate authority you trust
could see the credential. If the local Python has no CA certificates
(python.org builds on macOS until Install Certificates.command is run),
checks report "unknown" and HCI prints a warning saying why.
CI / GitHub Actions
Create .github/workflows/secrets.yml in the repository to scan:
name: Secrets scan
on: [push, pull_request]
permissions:
contents: read
jobs:
scan:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0 # Required to inspect the full available history.
- uses: taylannuhogluofficial-png/hci-scan@v0.8.0
with:
fail-on: high
history: "true"
Pin a release tag (as above) or, for the strictest supply-chain hygiene,
a reviewed commit SHA. Check out the target repository
before running the action; use fetch-depth: 0 when enabling history.
The job log shows only finding counts, never values, paths or rule names:
logs of a public repository are world-readable. Set upload-results: "true"
to also keep the JSON findings as a workflow artifact (7 days). It is off by
default because on a public repository anyone signed in to GitHub can
download artifacts. With sarif: "true", results go only to the Security
tab, which is access-controlled.
See action.yml for all inputs.
Native GitHub Code Scanning (SARIF)
Beyond the pass/fail build gate above, findings can show up as real annotations on a PR diff and alerts in the repo's Security → Code scanning tab — the same place CodeQL results land — instead of only existing in a log a human has to go read:
permissions:
contents: read # required to check out the repository
security-events: write # required to upload SARIF
actions: read # required on a private repo (upload-sarif checks workflow-run status)
steps:
- uses: actions/checkout@v4
- uses: taylannuhogluofficial-png/hci-scan@v0.8.0
with:
sarif: "true"
fail-on: high
Or directly: hci scan . --sarif > results.sarif, then
github/codeql-action/upload-sarif.
No raw secret value is ever written to the SARIF file — same masked
preview as everywhere else, since a SARIF file is itself a stored artifact,
not just terminal output. See it live: this repo's own
self-scan.yml uploads its own SARIF on
every push — and, on a daily schedule, even with zero pushes. Both this and
ci.yml (which reruns pip-audit daily too)
carry a schedule: trigger for exactly this reason: a CVE can be disclosed
for an already-pinned dependency at any time, and a push-only trigger would
leave that unnoticed until someone happens to next touch the repo.
Kubernetes Secret decoding
A kind: Secret manifest's data: block is base64-encoded by Kubernetes
convention — meaning every content rule here is otherwise blind to it, no
matter how good the regex is, since none of them can match inside a base64
blob. HCI detects kind: Secret YAML documents, decodes each data:
entry, and re-runs the same rule engine against the decoded content —
so a tls.key entry gets caught by the private-key-block rule, a
db-password entry by the password rules, using k8s's own field name for
context. (stringData: entries need no special handling — they're already
plaintext and caught by the normal per-line scan.)
apiVersion: v1
kind: Secret
metadata:
name: db-credentials
data:
db-password: TXlDMG1w… # -> flagged, decoded value shown masked
tls.key: LS0tLS1CRUdJTi... # -> flagged as a private key, not a blob
A ConfigMap or any other kind: is left alone — this only activates
inside documents that are actually kind: Secret.
Architecture
input (files / staged / git history)
│
▼
rules engine ── rules.yaml (data, not code)
│
▼
scanner ── regex per line, + optional entropy pass
│
▼
filter layer ── path ignores, placeholders, .hciignore baseline
│
▼
verify (opt-in) ── live read-only API check per credential type
│
▼
reporter ── Rich terminal output, or --json
Each layer is independent, which is what let git-history scanning, entropy detection, and live verification all get added without touching the others — and is why a contributor can add a new rule or a new verifier without touching the scanner at all.
Roadmap
- CLI: directory + git-history +
--diff/--stagedscanning, 23 content rules, entropy mode - False-positive controls: ignore paths, placeholders, baseline, context-aware severity
- Pre-commit hook
- GitHub Action + Dockerfile
- Opt-in live-credential verification: GitHub, Slack, Stripe, SendGrid, npm, and real SigV4-signed AWS key-pair verification
- Dependency vulnerability scanning in CI (
pip-audit) - Baseline rule-set versioning (drift note when rules change since baseline)
- Filename-based structural rules (Terraform state, default-named SSH/PuTTY keys)
- SARIF output + native GitHub Code Scanning integration
- Kubernetes
kind: Secretdata:block base64 decoding + rescanning - Filename rules and Kubernetes Secret decoding applied to
--history, not just working-tree/staged/diff — a deleted.tfstatefile or a Secret manifest removed in a later commit is still caught -
hci respond: per-credential incidents, exposure analysis, provider playbooks, filter-repo purge plans, Markdown incident reports, gitleaks/trufflehog import - More verifiers (Google, Twilio, Mailgun, GitLab)
- Multi-line secret detection (values split across lines in JSON/YAML)
- Config-file support (
hci.toml) instead of only CLI flags - Publish to PyPI and create versioned GitHub Action releases
- Team dashboard / hosted monitoring (the commercial layer, later)
Contributing
Adding a detection rule is the easiest way to contribute: add an entry to
hci/config/rules.yaml with a name, regex, severity,
description, and remediation, plus a test. Build fake credentials from pieces
at runtime (see tests/fakes.py) so the repository itself
never contains a scannable one.
After the source installation above, install the development tools, the pre-commit hook (once per clone), and run every check:
python -m pip install -e '.[dev]'
make install-hooks # every commit is scanned; a staged secret blocks it
make check # tests, self-scan (tree + history), audit, package scan
CI runs the same checks on every push.
Use synthetic credentials in regression fixtures. The test suite checks detection, filtering, history/staging behavior, masked reports, and verifier handling without needing live credentials.
License
MIT — see LICENSE.
Metadata
Release files for hci-scan 0.8.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| hci_scan-0.8.0.tar.gz | 128.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| hci_scan-0.8.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 217.8 kB
Release files / hci_scan-0.8.0.tar.gz
| Download URL | hci_scan-0.8.0.tar.gz |
|---|---|
| Size | 128.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
9257a4a0bd8bcf6a6479cf877a635bb902098e88554b368095df51743a36de3c
|
|
BLAKE2b-256 checksum How to use checksums |
27612e6b5ea92a948be4a09f1f05c8656e592264bf4760425158e5ec770f95c8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.
Transparency logRelease files / hci_scan-0.8.0-py3-none-any.whl
| Download URL | hci_scan-0.8.0-py3-none-any.whl |
|---|---|
| Size | 89.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
bbcb42f6e0f77641f960b3863c1dfa0dc398b218a855e07f9579393254ec6a6d
|
|
BLAKE2b-256 checksum How to use checksums |
45af843277eb1ec8384dc6b6570b8aa6c637e8d992f94d205b857e26106fb6ae
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.
Transparency log