Skip to main content

litscan

License: MIT Version Python PyPI

A small CLI tool that scans a codebase for string and numeric literals, helping you quickly spot hard-coded values in source files.

Prerequisites

  • Python 3.14+

Installation

pip install litscan

Usage

After installation, litscan is available as a console script:

litscan <path> [options]

What is detected

The scanner uses tree-sitter to parse each file and extracts literal nodes per language:

Language Extensions Literal types detected
Python .py .pyi strings, integers, floats
JavaScript .js .mjs .cjs strings, numbers, template strings
TypeScript .ts .tsx strings, numbers, template strings
Java .java string literals, text blocks, integer literals, floating-point literals
Go .go interpreted strings, raw strings, integer literals, float literals
Gosu .gs .gsx string literals, integer literals, floating-point literals
C .c .h string literals, number literals, char literals
C++ .cpp .cc .cxx .hpp .hxx string literals, number literals, char literals
C# .cs string literals, integer literals, real literals
Rust .rs string literals, integer literals, float literals
Kotlin .kt .kts string literals, number literals, float literals
Swift .swift string literals, integer literals, real literals
Scala .scala strings, integer literals, floating-point literals
Groovy .groovy .gradle string literals, integer literals, floating-point literals

Files with extensions not in the table above are skipped with a warning.

Results are grouped by unique literal value and sorted by occurrence count (highest first).

Arguments

Argument Description
path Target directory or file to scan. Multiple paths can be specified, separated by a semicolon (e.g. src;lib;tests).

Options

Option Default Description
--ext <exts> (all files) Comma-separated extensions to include (e.g. py,js,ts)
--output <name> litscan-output Base name (without extension) for output file(s)
--output-dir <dir> reports Directory where output file(s) will be written
--format <fmt> json Output format: json, html, or all
--workers <n> min(32, cpu_count + 4) Number of parallel worker threads used during scanning
--db <path> <system-temp>/litscan.db Path to the SQLite scratch database that stores occurrences during a scan run. Session records are removed after the report is written.
--functions-only (off) Scan only literals that appear inside function or method implementations. Supported languages: Python, JavaScript, TypeScript, Java, Go, Gosu, C, C++, C#, Rust, Kotlin, Swift, Scala, Groovy.
--min <count> 0 Minimum occurrence count a literal must have to be included in the report. 0 means no filtering.
--mode <mode> both Literal category to scan: string, number, or both.
--literals <values> (all) Semicolon-separated target literal values to restrict the report to (e.g. foo;bar). Matched against the decoded, single-line literal value; multi-line literals are never matched.
--target-list (off) Treat path as a single existing file listing target paths (files and/or directories), one per line, instead of a semicolon-separated path list. Blank lines and lines starting with # are skipped.
--version Print the version and exit.

Examples

Scan all files in the current directory and produce a JSON report:

litscan .

Scan only Python and JavaScript files in src/:

litscan src --ext py,js

Generate both JSON and HTML reports in a custom directory:

litscan . --format all --output-dir my-reports

Scan a Java source tree with a custom output name:

litscan src/main/java --ext java --format all --output-dir reports

Scan only literals inside functions and methods:

litscan src --functions-only

Only report string literals that occur at least 3 times:

litscan src --mode string --min 3

Restrict the report to specific target literal values:

litscan src --literals "TODO;FIXME"

Scan the targets listed in a file, one path per line:

litscan targets.txt --target-list

Configuration

Environment variable Description
LITSCAN_CONFIG_DIR Directory where logging.ini, lit_ignore, .litscanignore, and config.ini are seeded on first run and read from. When unset, the bundled copies inside the package are used directly.

Ignore patterns

The lit_ignore file (seeded into LITSCAN_CONFIG_DIR on first run) contains one regex pattern per line. Any literal whose value matches a pattern is excluded from scan results. Edit the file to suppress noise such as common stop-words or numeric constants you do not care about.

Ignored files and directories

The .litscanignore file (also seeded into LITSCAN_CONFIG_DIR on first run) uses gitignore syntax to exclude entire files or directories from being scanned in the first place — matching directories are pruned during traversal, so their contents are never read. It ships with sensible defaults (.git/, node_modules/, dist/, build/, __pycache__/, .venv/); edit the file to add project-specific paths to skip.

Overriding the ignore filename

config.ini (also seeded into LITSCAN_CONFIG_DIR on first run) contains an [override] section with an ignore-file key, which names the file used in place of .litscanignore, resolved relative to LITSCAN_CONFIG_DIR:

[override]
ignore-file = .litscanignore

Point ignore-file at a different filename to use an alternate ignore file (also placed inside LITSCAN_CONFIG_DIR). If the configured file is missing, litscan logs a warning and falls back to the bundled .litscanignore.

Report metadata

Every JSON/HTML report records the run's inputs alongside the findings: the resolved --path entries (paths-scanned), the --min threshold (min-count), the --mode used, and any --literals targets applied.

Development

Prerequisites

  • Poetry 2.2+

Installation

poetry install

Architecture

flowchart TD
    CLI["cli.py\n(entry point)"] --> logenrich["setup_logger()\nlogenrich"]
    CLI --> discover["discover_files()"]
    discover --> pathignore[".litscanignore\n(braincraft.IgnoreFile)"]
    pathignore --> config["config.ini\n([override] ignore-file)"]
    discover --> concurrent["ThreadPoolExecutor\n(parallel scan)"]
    concurrent --> scan["scan_file()\nscanner.py"]
    scan --> parser["parser.py\n(tree-sitter)"]
    parser --> ts["Language-specific\ngrammar packages"]
    scan --> litignore["lit_ignore\n(exclude patterns)"]
    scan --> store["SessionStore\nstore.py (SQLite)"]
    store --> report["write_outputs()\nreporter.py"]
    report --> JSON["JSON report"]
    report --> HTML["HTML report"]
Module Responsibility
cli.py Argument parsing, file discovery, orchestration
config.py Config — reads config.ini overrides (e.g. [override] ignore-file)
parser.py Tree-sitter language loading (LRU-cached) and source parsing
scanner.py AST-based literal extraction; LiteralOccurrence / LiteralGroup types
store.py SessionStore — thread-safe SQLite scratch store; one UUID per scan run
reporter.py write_outputs() — renders JSON and/or HTML reports
logenrich External library that provides setup_logger() — logging config seeded from logging.ini

Test with coverage

poetry run pytest --cov=litscan tests --cov-report html

Format and lint

poetry run black litscan; poetry run pylint litscan

Quality gates

  • Coverage ≥ 90%
  • Pylint score 10/10

Example

Scan the test fixtures and produce both JSON and HTML reports:

poetry run litscan tests\fixtures --format all

Publishing to PyPI

Prerequisites

  • A PyPI account with an API token.

Configure the token

poetry config pypi-token.pypi <your-token>

Build and publish

poetry publish --build

This builds the source distribution and wheel, then uploads them to PyPI in one step.

Note: PyPI releases are immutable. Once a version is published, it cannot be overwritten.
To fix a mistake, yank the release via the PyPI web UI and publish a new version.

Changelog

License

MIT

Author

Ron Webb <ron@ronella.xyz>

Release files for litscan 2.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for litscan 2.2.1
File Size Uploaded
litscan-2.2.1.tar.gz 27.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for litscan 2.2.1
File Interpreter ABI Platform
litscan-2.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 51.2 kB

Release files / litscan-2.2.1.tar.gz

Download URL litscan-2.2.1.tar.gz
Size 27.0 kB
Tags Source
SHA-256 checksum
How to use checksums
2e2c23e5cea6d7e4c92074cad52cf669e98d9becf78ad6cb815ae99e33ea576e
BLAKE2b-256 checksum
How to use checksums
5aebbed2be71da4c52aef7609b89335537dd8dd567716c57317fd78d023d7587
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/2.2.0 CPython/3.14.7 Linux/6.17.0-1022-azure

Release files / litscan-2.2.1-py3-none-any.whl

Download URL litscan-2.2.1-py3-none-any.whl
Size 24.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
379d02b23cfa25503b682de63ff5f77a28abf930bd5cadfa8029b11ba4d4be8f
BLAKE2b-256 checksum
How to use checksums
a16bffb63a2904cad26ac2661f331b859b662e71d1499b64632bd388a820c53d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/2.2.0 CPython/3.14.7 Linux/6.17.0-1022-azure

Release history Release notifications | RSS feed

This release

2.2.1 This release

2 release files

2.2.0

2 release files

2.1.1

2 release files

2.1.0

2 release files

2.0.1

2 release files

2.0.0

2 release files

1.4.0

2 release files

1.3.1

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.0

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page