SAM Doctor
SAM Doctor reads a failed sam deploy, cdk deploy, or
aws cloudformation deploy log locally and reports the first supported failure
pattern it finds: a short diagnosis, redacted evidence lines, safe verification
commands, and a link to the relevant official documentation.
It does not access AWS, upload logs, change resources, or claim an authoritative root cause. It matches known patterns in text you provide. When nothing matches, it says so instead of guessing.
Project page | GitHub Marketplace | Report a bad diagnosis | Request a rule
Try it
python -m pip install sam-doctor
sam-doctor demo
The bundled demo needs no AWS credentials and makes no network calls. Other install paths:
pipx install sam-doctor # isolated global CLI
uvx sam-doctor demo # run without installing
To install from a tagged source release instead of PyPI, use
pip install "sam-doctor @ git+https://github.com/jakegold1647/sam-doctor.git@<tag>"
with a tag from the releases page.
If your shell cannot find sam-doctor after installing, use
python -m sam_doctor instead.
The demo diagnoses a bundled GitHub Actions OIDC failure:
SAM Doctor found 1 possible issue(s) in oidc-assume-role-failure.txt.
1. GitHub Actions cannot assume the configured AWS role through OIDC (high confidence)
Matched on line: 2
The workflow reached AWS STS but the role trust relationship did not accept
the GitHub-issued OIDC token. ...
Evidence:
- Error: Not authorized to perform: sts:AssumeRoleWithWebIdentity
- An error occurred (AccessDenied) when calling the AssumeRoleWithWebIdentity
operation: Not authorized to perform sts:AssumeRoleWithWebIdentity
Verify:
- Confirm the workflow or job permissions include `id-token: write`.
- Check that the role trust policy accepts `token.actions.githubusercontent.com:aud`
equal to `sts.amazonaws.com`.
Docs: https://docs.github.com/actions/how-tos/secure-your-work/security-harden-deployments/oidc-in-aws
The description line is truncated here; the CLI prints it in full, along with a
third trust-policy sub check. Output is deterministic for the same input, and
evidence is redacted before display.
Who this is for
Use it as a fast local first pass when a SAM, CloudFormation, or GitHub Actions deployment fails and the useful error line is buried under rollback noise. It works best on logs with explicit error lines and rollback context.
Skip it if you need account-state inspection, drift analysis, quota checks, or automatic fixes. See When not to use this.
Usage
Diagnose a log file:
sam-doctor diagnose deployment.log
Pick the output format for where the report is going:
| Situation | Command |
|---|---|
| Handoff in a ticket or thread | sam-doctor diagnose deployment.log --format markdown |
| Machine-readable output for CI | sam-doctor diagnose deployment.log --format json --output diagnosis.json |
| GitHub workflow annotations | sam-doctor diagnose deployment.log --format github |
| Code scanning / SARIF consumers | sam-doctor diagnose deployment.log --format sarif --output sam-doctor.sarif |
| Pasted excerpt, no file | printf '%s\n' "...error excerpt..." | sam-doctor diagnose - |
| A workflow that saves a log | The GitHub Action below |
All formats include the first matching line number and the matched evidence, not the full input log. Standard input works anywhere a file path does, so you can pipe from other tools:
kubectl logs deploy/my-api | sam-doctor diagnose -
Multiple findings
When several supported patterns appear, findings are ordered by their first
matching log line, which puts the root failure before the rollback it caused.
sam-doctor demo --scenario cloudformation reproduces this shape:
SAM Doctor found 2 possible issue(s) in cloudformation-resource-failure.txt.
1. CloudFormation resource creation or update failed (high confidence)
Matched on line: 1
Evidence:
- ... MyApiDeployment CREATE_FAILED Resource handler returned message: "Invalid
request provided: API Gateway deployment cannot be created because the stage
already exists." ...
Verify:
- Identify the failed logical resource ID and preserve its exact status reason.
- Check the underlying service event or API error named in that status reason.
- Fix the resource-level cause before retrying the stack operation.
2. CloudFormation stack entered rollback after an earlier resource failure (medium confidence)
Matched on line: 2
Verify:
- Inspect stack events in chronological order and locate the first
`CREATE_FAILED` or `UPDATE_FAILED` resource.
Other bundled scenarios: sam-doctor demo --scenario capabilities,
api-gateway, esbuild, python-pip.
Batch mode
Diagnose many logs in one run:
sam-doctor batch logs/*.log logs/*.txt --format json --output batch-results.json
Add --fail-on-findings to exit 1 when any file has a supported finding;
the full batch report is still written first. With --format github, batch
mode emits one annotation per finding and skips successful inputs.
Exit codes
| Status | Meaning |
|---|---|
0 |
Command completed with no enforced fail gate hit. |
1 |
--fail-on-findings found one or more supported findings (diagnose or batch). |
2 |
CLI usage or error-path failure (missing inputs, invalid arguments). |
Details and examples: docs/cli-exit-and-action-exit-codes.md.
JSON schemas
The JSON payload shapes are documented in checked-in schemas:
docs/schemas/diagnose-report.schema.jsondocs/schemas/batch-report.schema.jsondocs/schemas/rules-report.schema.json
sam-doctor schemas prints the schema URLs. The contracts are
additive-compatible: new top-level fields may appear, but removing or renaming
a documented required field is a breaking change and gets a coordinated
version bump.
GitHub Actions
Add a diagnostics step after any step that saves a deployment log:
- name: Deploy
shell: bash
run: |
set -o pipefail
sam deploy --no-confirm-changeset 2>&1 | tee deployment.log
- name: Diagnose deployment log
if: always()
id: sam-doctor
uses: jakegold1647/sam-doctor@v0
with:
log-file: deployment.log
summary: true
# Uncomment to fail this job when a supported finding is detected.
# fail-on-findings: true
Keep if: always(); otherwise GitHub Actions skips the step exactly when the
deployment fails. The Markdown job summary is opt-in and contains only matched,
redacted evidence. The action also adds redacted workflow annotations for each
finding by default; set annotations: "false" to disable them.
sam-doctor init generates this workflow for you:
sam-doctor init --deploy-command "sam deploy --no-confirm-changeset" --summary --annotations
By default the generated workflow only runs on workflow_dispatch (the
"Run workflow" button in the Actions tab), so an init you ran to try
things out can't quietly turn into a deployment on your next push. Add
--on-push when you're ready for the workflow to deploy automatically on
pushes to main:
sam-doctor init --deploy-command "sam deploy --no-confirm-changeset" --on-push --summary --annotations
A rollout pattern that works: run non-blocking for 3-5 stable runs, then add
--fail-on-findings --force to regenerate with strict gating. If you would
rather not regenerate per mode, the
two-phase starter workflow
stays non-blocking by default and enforces only on a manual
workflow_dispatch with rollout-mode: strict.
Action exit codes and outputs
0: no enforced failure (findings may still exist).1: findings present andfail-on-findings: true.2: runtime or precondition failure (invalid boolean inputs, missing Python).
The action exposes finding-count and has-findings outputs for non-blocking
routing:
- name: Route to dedicated triage when action reports findings
if: steps.sam-doctor.outputs.has-findings == 'true'
run: |
echo "Routing failure with ${{ steps.sam-doctor.outputs.finding-count }} findings to a higher-signal runbook."
Action batch mode
For CI setups that write many logs per run, set batch: true and point
log-file at a directory or glob:
- name: Diagnose logs in batch
if: always()
id: sam-doctor-batch
uses: jakegold1647/sam-doctor@v0
with:
log-file: logs/
batch: true
summary: true
Starter workflows
Pick the template that matches your deploy command:
- SAM deploy:
examples/github-actions-workflow.yml - SAM sync:
examples/github-actions-workflow-sam-sync.yml - CloudFormation package/deploy:
examples/github-actions-workflow-cf-pipeline.yml - CDK deploy:
examples/github-actions-workflow-cdk.yml - Batched logs in one run:
examples/github-actions-workflow-batch-logs.yml
The CI command matrix maps exact deploy commands
to templates, and examples/README.md indexes everything.
Other CI systems
- GitLab:
examples/gitlab-ci-sam-doctor.yml - CircleCI:
examples/circleci-sam-doctor.yml - Azure Pipelines:
examples/azure-pipelines-sam-doctor.yml - Bitbucket Pipelines:
examples/bitbucket-pipelines-sam-doctor.yml
What it detects
Run sam-doctor rules (or rules --format json) for the current
machine-readable catalog. Each rule triggers on an explicit error signal in the
log, not on template inspection or AWS account access, and carries a stable id
(iam.deny.explicit, and so on) that CI tooling can match on across releases -
see docs/stability.md. The current set:
- GitHub Actions OIDC errors: missing
id-token: write, audience mismatch, trust-policy/subject mismatch, andAssumeRoleWithWebIdentityfailures - IAM
AccessDeniedfailures, with explicit denies (including service control policies) distinguished from missing-policy denials - Expired AWS credentials and runner clock skew (
ExpiredToken,Signature expired) - CloudFormation API throttling (
Rate exceeded) - CloudFormation failed-resource events and rollback states
- Another operation already in progress on the stack
(
OperationInProgressException,*_IN_PROGRESS state and can not be updated) - Empty change sets (
No changes to deployin CI) - Resources that fail to stabilize, with the nested handler message surfaced first
- Exports that cannot change because another stack imports them
- Lambda deployment packages over a per-function size limit, and the regional
code storage quota (
CodeStorageExceededException) - Blocked stack deletion:
DELETE_FAILEDblockers and termination protection - ECR push authentication failures from the CI runner (missing login, expired
token, denied
ecr:GetAuthorizationToken) - CloudFormation capability acknowledgement errors (
InsufficientCapabilities) - Lambda container-image failures caused by missing ECR image access
- API Gateway deployments created before methods exist
- API Gateway CORS preflight conflicts
- SAM deployment/configuration errors, including conflicting artifact-bucket
settings and missing
esbuilddependencies - SAM build errors where Docker is unavailable for
sam build --use-container - Python dependency resolution or validation errors in SAM/Python builds
- Interactive changeset prompts that stall non-interactive CI
- Template failures: SAM/CloudFormation schema validation
(
InvalidSamDocumentException, unsupported properties), invalid properties for a resource type, and templates over a CloudFormation size or count quota - S3 naming failures: invalid bucket names and globally taken names
(
BucketAlreadyExists,BucketAlreadyOwnedByYou) - Artifact-path failures: a
CodeUrithat was never built, a deployment bucket that denies access to the packaged artifacts, and Lambda layer artifacts CloudFormation cannot read back - IAM trust-policy shape errors and Lambda code-signing conflicts
If a deployment error you hit is not covered, open a
rule request
with a sanitized 5-15 line excerpt and the command you ran.
sam-doctor request-packet deployment.log writes that excerpt for you: a
redacted context window around the first likely error, never the full log.
What a report includes
- A likely failure category and confidence level.
- Up to three matched log lines, redacted before output.
- Safe checks to validate the diagnosis before changing a policy or stack.
- A link to the relevant official documentation.
Reports redact AWS account IDs, ARNs, email addresses, common AWS access key IDs, bare STS session tokens, secret assignments, bearer tokens, JWT-style tokens, and common GitHub token formats before matched evidence is shown. This is a guardrail, not a secret scanner: review a report before sharing it.
To package a diagnosis for handoff, sam-doctor packet deployment.log writes
diagnosis.md and diagnosis.json; the
evidence packet template describes what
to share alongside them, and RESEARCHER_OVERVIEW.md
is the summary to hand a reviewer or researcher.
How it compares
- vs. reading the log yourself. For a failure you have seen before, just read the log. SAM Doctor helps when the useful line is buried under rollback noise, or when the error text (OIDC trust-policy mismatches especially) does not say what to check next. It finds the first supported failure signal and pairs it with the verification steps and the official doc page.
- vs. pasting the log into an LLM. An LLM can reason about failures SAM Doctor has no rule for, and that is sometimes the right call. The trade-offs: you upload the log (deployment logs routinely contain account IDs, ARNs, and role names), the answer varies run to run, and it may be confidently wrong. SAM Doctor is deterministic, runs offline, and redacts by default - and when it has no matching rule, it says so instead of guessing. Using it first and an LLM for the leftovers is a reasonable workflow.
- vs. AWS Support. Support can see your account state; SAM Doctor cannot and does not try. It is the two-minute local check you run before deciding whether a ticket is worth opening, and its redacted report is a safer artifact to paste into one.
When not to use this
- The failure is in application runtime behavior, not the deployment itself - this reads deployment logs, not CloudWatch application logs.
- You need account-state inspection (drift, quotas, existing resources). SAM Doctor never calls AWS, by design.
- Your failure is outside the supported rules - you get an honest "no supported pattern found", not a guess.
- You want an automatic fix. Every report is a prompt to verify, not a change to apply.
Scope and safety
Run this only on logs you are authorized to inspect. Review every suggested command and policy change before applying it. SAM Doctor is diagnostic help, not security, legal, or production-operations advice.
Guides
- Add SAM Doctor to an existing GitHub Actions deployment
- On-call playbook: triage sequence, handoff template, escalation threshold
- Fix "Not authorized to perform: sts:AssumeRoleWithWebIdentity" in GitHub Actions
- Find the first useful error in a CloudFormation ROLLBACK_COMPLETE
- Fix "InsufficientCapabilitiesException" in an AWS SAM deployment
- Worked examples (incident-to-action workflows)
- Rolling out SAM Doctor on a team (commands by role)
- Create a reproducible evidence packet for collaboration
Contributing
New contributors are welcome, and the best first changes are small: a
documentation correction, a reproducible false positive or missed diagnostic,
or one new diagnostic rule with a positive and a nearby-negative test. Start
with the contributor setup, pick a fully specified rule
from the rule roadmap or an issue labeled
good first issue,
and run python scripts/check-pr.py before opening the PR — it is the same
gate CI runs. Before sharing any log excerpt, remove account IDs, ARNs,
credentials, tokens, and customer data.
When you report a wrong or unclear diagnosis, include the SAM Doctor version, the exact command, a sanitized excerpt, and what you expected. Small reproducible reports get fixed fastest.
Development
python -m pip install -e ".[dev]"
python -m pytest -q
python -m build
python scripts/run-smoke.py runs the packaged demo and a sample diagnosis,
then checks that the JSON output is well-formed and contains findings.
See CHANGELOG.md for release history, SECURITY.md for vulnerability reporting, SUPPORT.md for help boundaries, and docs/pypi-publishing.md for the stable-release publishing setup.
Related projects
- Portfolio: jacobgoldstein.dev
- Historical text tooling: aktreader and aktreader-research
- Records corpus: congress-poland-registers
Release files for sam-doctor 0.9.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| sam_doctor-0.9.0.tar.gz | 82.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| sam_doctor-0.9.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 127.1 kB
Release files / sam_doctor-0.9.0.tar.gz
| Download URL | sam_doctor-0.9.0.tar.gz |
|---|---|
| Size | 82.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3715537425002bc88a41c868895c268c916432e19d90b579dd2a8c36186bcd03
|
|
BLAKE2b-256 checksum How to use checksums |
e461660d3ac6931bb1e0018cca491b2b7cf915ea9bc3d5a18181a53876242668
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 8, 2026.
Transparency logRelease files / sam_doctor-0.9.0-py3-none-any.whl
| Download URL | sam_doctor-0.9.0-py3-none-any.whl |
|---|---|
| Size | 44.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8394928b83e6e5c0177e74d2b8552faf524a3d7f7f1b55f956963dba4464e480
|
|
BLAKE2b-256 checksum How to use checksums |
01146e30da2839e66234bac5876c542818220d368e53a2e94850bd5624dbc460
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 8, 2026.
Transparency log