Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

Arid banner

Status: WIP Scope: Python-only Purpose: Duplicate-code detection Config: tool.arid License: MIT OR Apache-2.0

Arid icon Arid

Fast Python duplicate-code checker written in Rust. A focused replacement for Pylint R0801 that complements Ruff.

What is Arid? · Project status · Goals · Usage · Output · Configuration · Architecture · License


Project status

[!IMPORTANT] Arid is a work in progress. The project is under active development and is not yet production-ready. Interfaces, behavior, defaults, and packaging details may change as the implementation matures.

Arid is currently being built as a small, focused CLI for one job:

Detect duplicated Python source code quickly and accurately.


What is Arid?

Arid is a Python-specific CLI for duplicate-code detection.

It is designed to replace the duplicate-code functionality of Pylint R0801 / symilar without turning into another general-purpose linter. Arid is intentionally narrow in scope and is meant to run alongside Ruff, not compete with it.

Ruff
├── linting
├── formatting
├── imports
├── modernization
└── general code quality

Arid
└── duplicate-code detection

Why Arid? Because duplicated code isn't DRY.


Why not just use Pylint?

Pylint's R0801 checker provides useful Python-aware duplicate-code detection, but duplicate analysis can become very slow on larger codebases.

Arid aims to preserve the useful behavior of R0801 while using a Rust-native architecture designed specifically for duplicate detection.

The goal is not bug-for-bug compatibility. Where Pylint relies on textual heuristics, Arid prefers correct Python syntax interpretation.


Why not just use jscpd?

jscpd is a capable multi-language copy/paste detector, and its current implementation is also written in Rust.

Arid occupies a narrower niche:

  • Python only
  • focused on Pylint-style duplicate-code semantics
  • Python-aware filtering for comments, docstrings, imports, and signatures
  • designed to fit naturally into modern Python workflows
  • intentionally minimal in scope

Arid is not intended to replace jscpd for multi-language repositories.


Goals

Arid v1 is being designed to:

  • detect duplicated Python source blocks across files
  • detect duplicated blocks within the same file
  • ignore comments, docstrings, imports, and function signatures when configured
  • preserve accurate original source locations
  • report concise DUP001 diagnostics
  • describe duplicate findings using Python structural context
  • provide deterministic duplication metrics
  • support pyproject.toml configuration via [tool.arid]
  • provide machine-readable JSON output
  • run substantially faster than Pylint's duplicate-code checker
  • require no Python runtime to analyze Python source

Non-goals

Arid is intentionally not a general-purpose linter.

It does not aim to provide:

  • formatting
  • import sorting
  • type checking
  • dead-code detection
  • complexity analysis
  • security scanning
  • semantic clone detection
  • structural clone matching
  • fuzzy AST similarity
  • multi-language duplicate detection

If a feature belongs naturally in Ruff, it does not belong in Arid.


Usage

Arid is intended to fit naturally into a Python quality workflow:

ruff check .
arid .

Scan specific files or directories:

arid src tests

Require a larger duplicate before reporting it:

arid . --min-lines 8

Override normalization behavior for a single scan:

arid . --no-ignore-docstrings

Configurable boolean options support both positive and negative forms:

--ignore-comments       --no-ignore-comments
--ignore-docstrings     --no-ignore-docstrings
--ignore-imports        --no-ignore-imports
--ignore-signatures     --no-ignore-signatures
--same-file             --no-same-file
--hidden                --no-hidden

This allows command-line arguments to explicitly override either value from pyproject.toml.

Hidden files and directories are skipped by default during directory discovery. Include them when needed with:

arid . --hidden

This allows Arid to scan Python files under hidden directories such as .github/ while still honoring .gitignore and configured exclude patterns.

Exclude matching paths:

arid . --exclude 'generated/**'

--exclude may be repeated:

arid . \
    --exclude 'generated/**' \
    --exclude 'vendor/**'

Include the original source in each finding:

arid . --show-source

Emit machine-readable JSON:

arid . --json

Example diagnostic:

DUP001 4 duplicated lines
Context: declarative
Scope: class
Occurrences: 2 across 2 files (cross-file)

  src/models/user.py:12-15
  src/models/account.py:20-23

Found 1 duplicate group.
4 duplicate lines (2.31%).

Understanding Arid's output

Arid separates two questions:

Detection answers: "Is this code duplicated?"
Context helps answer: "What kind of code is duplicated?"

Arid deliberately does not assign a severity or decide whether duplication should be removed. Duplicate code can be intentional, harmless, framework-driven, or worth refactoring.

The structural metadata exists to help you make that decision.

Consider:

DUP001 4 duplicated lines
Context: declarative
Scope: class
Occurrences: 2 across 2 files (cross-file)

  src/models/user.py:12-15
  src/models/account.py:20-23

Found 1 duplicate group.
4 duplicate lines (2.31%).

DUP001 4 duplicated lines

DUP001 is Arid's duplicate-code diagnostic.

4 duplicated lines means the matching region contains four effective normalized lines that satisfy the configured duplicate threshold.

Arid compares source after its configured Python-aware normalization. Depending on configuration, this can remove constructs such as:

  • comments
  • docstrings
  • imports
  • function signatures

Blank lines do not count toward min-lines, and lines containing only non-substantive punctuation do not increase the effective-line count.

Because of that, a finding reported as four duplicated lines may span more than four physical source lines.

Context

Context describes the structural kind of Python code involved in the duplicate.

Possible values are:

Context Meaning
declarative The duplicate consists of declarations or definitions, such as direct module/class assignments or definitions.
executable The duplicate consists of executable statements, control flow, or function-body logic.
mixed The duplicate contains or occurs across more than one structural context.

For example:

Context: declarative

often appears for repeated class or module definitions.

Context: executable

often appears for repeated application logic inside functions.

[!NOTE] Context is descriptive, not a severity. declarative does not mean "safe to ignore," and executable does not mean "must refactor."

Arid describes Python structure without attempting to infer framework semantics or developer intent.

It therefore does not label findings as "ORM boilerplate," "configuration noise," "safe duplication," or similar framework-specific categories.

Scope

Scope describes where the duplicated code occurs structurally.

Possible values are:

Scope Meaning
module Module-level code.
class Code structurally associated with a class.
function Code structurally associated with a function or method.
mixed The duplicate spans or occurs across more than one scope.

For example:

Context: executable
Scope: function

indicates repeated executable logic within functions or methods.

By contrast:

Context: declarative
Scope: class

indicates repeated declarative code associated with classes.

Again, scope describes where the duplicate exists, not whether it is a problem.

Occurrences

The occurrence line tells you how widely the duplicate appears.

Occurrences: 2 across 2 files (cross-file)

contains three pieces of information:

  • the number of duplicate occurrences
  • the number of distinct files containing them
  • how those occurrences are distributed

Distribution values are:

Distribution Meaning
same-file All occurrences are contained in one file.
cross-file Occurrences are spread across multiple files, with one occurrence in each involved file.
mixed Multiple files are involved and at least one file contains multiple occurrences.

Examples:

Occurrences: 2 across 1 file (same-file)

means the same block appears twice in one file.

Occurrences: 3 across 3 files (cross-file)

means one occurrence appears in each of three files.

Occurrences: 4 across 3 files (mixed)

means the duplicate spans multiple files and at least one of those files contains more than one occurrence.

Source locations

Locations such as:

src/models/user.py:12-15

always refer to the original physical Python source, not Arid's internal normalized representation.

This remains true even when ignored comments, imports, signatures, docstrings, or blank lines appear within the physical range.

Use:

arid . --show-source

to include the original source text alongside each location.

Duplicate groups

When the same block appears more than twice, Arid reports it as one duplicate group rather than generating every possible pair.

For example, a block appearing in:

a.py
b.py
c.py

is one finding with three occurrences, not three separate pairwise findings.

Duplicate lines and duplication percentage

The final summary:

4 duplicate lines (2.31%).

measures redundant effective lines, not every line participating in a duplicate.

One occurrence of each duplicate group is treated as canonical. Only redundant copies beyond that canonical occurrence contribute duplicate lines.

For example:

10-line block × 2 occurrences

contributes:

10 duplicate lines

not 20.

A 10-line block appearing three times contributes:

20 duplicate lines

because two of the three copies are redundant.

Overlapping redundant regions are not counted repeatedly.

The duplication percentage is:

duplicate effective lines
───────────────────────── × 100
 analyzed effective lines

This makes the metric represent how much analyzed code is redundant rather than how much code merely participates in a duplicated region.

How to interpret findings

There is no universal rule for which duplicate should be refactored first, but Arid's metadata can help you triage a large report.

A practical review order is often:

  1. Look at larger duplicate regions before very short ones.
  2. Review executable / function findings for repeated application logic.
  3. Look at findings with many occurrences to identify patterns repeated broadly through the codebase.
  4. Use same-file, cross-file, and mixed to distinguish localized repetition from code repeated across modules.
  5. Review declarative findings in context. Repeated declarations may be intentional, generated by a common coding pattern, or candidates for consolidation depending on the project.

Arid intentionally stops short of saying:

high severity
low value
safe to ignore
must refactor

Those are project-specific judgments.

Its job is to provide accurate duplicate detection and enough objective structural information for the developer to make them.


Configuration

Arid uses [tool.arid] in pyproject.toml:

[tool.arid]
min-lines = 4
ignore-comments = true
ignore-docstrings = true
ignore-imports = true
ignore-signatures = true
same-file = true
hidden = false
exclude = [
    "generated/**",
    "vendor/**",
]

Current defaults are:

Option Default Meaning
min-lines 4 Minimum effective normalized lines required for a duplicate.
ignore-comments true Ignore Python comments during matching.
ignore-docstrings true Ignore structural Python docstrings.
ignore-imports true Ignore import statements.
ignore-signatures true Ignore function and method declaration signatures.
same-file true Detect non-overlapping duplicate regions within the same file.
hidden false Include hidden files and directories during directory discovery.
exclude [] Path patterns excluded from discovery.

Configuration precedence is:

CLI arguments
    ↓
pyproject.toml
    ↓
built-in defaults

For example:

[tool.arid]
min-lines = 6
ignore-docstrings = true
same-file = true
hidden = false

can be overridden for one scan with:

arid . \
    --min-lines 10 \
    --no-ignore-docstrings \
    --no-same-file \
    --hidden

Each configurable boolean has both an enabling and disabling CLI form. This matters when the project configuration differs from the built-in default. For example, if the project contains:

[tool.arid]
ignore-comments = false

then:

arid . --ignore-comments

explicitly enables comment filtering for that scan.

Likewise:

arid . --no-ignore-comments

explicitly disables it.

Supplying one or more --exclude options on the command line overrides the configured exclude list for that scan:

arid . \
    --exclude 'build/**' \
    --exclude 'generated/**'

--json and --show-source control report output and are CLI-only options.


Detection model

Arid is focused on exact duplicate source blocks after configurable Python-aware normalization.

For example, with comments and function signatures ignored:

def first():
    # explanation
    value = calculate_value()
    save_value(value)

and:

def second():
    # different explanation
    value = calculate_value()
    save_value(value)

can be considered duplicates.

Arid v1 does not normalize identifiers, so these are intentionally different:

value = calculate_value()
save_value(value)
result = calculate_value()
save_value(result)

Arid can attach structural context such as declarative, executable, class, or function to a duplicate that it has already detected.

That does not make Arid a structural clone detector. Two pieces of code that are merely structurally similar but do not become identical after normalization are not considered duplicates.

Semantic clone detection, identifier-renaming clone detection, and fuzzy AST similarity remain outside the v1 scope.


Suppressing intentional duplication

Arid supports source-level suppression regions:

# arid: disable

# intentionally duplicated code

# arid: enable

Code inside a disabled region does not participate in duplicate detection.

Suppression regions also create matching boundaries, so Arid does not construct a duplicate across disabled source.

Use suppression for duplication that is intentionally accepted by the project rather than expecting Arid to infer whether a particular framework pattern or coding convention should be ignored.


Architecture

The v1 architecture is intentionally small:

discover
   ↓
parse
   ↓
normalize
   ↓
intern lines
   ↓
suffix array
   ↓
LCP
   ↓
maximal repeats
   ↓
DUP001

Arid analyzes Python source entirely in Rust and never imports or executes the project being scanned.

Duplicate detection operates on Arid's normalized source representation. Structural context is derived from Python syntax and attached as reporting metadata; it does not alter whether two normalized regions match.


Installation

[!WARNING] Arid is currently in early alpha. The CLI, configuration, and output format may change before 1.0.

uv

Install Arid as an isolated command-line tool:

uv tool install arid

pip

Install Arid from PyPI:

python -m pip install --pre arid

Verify the installation:

arid --version

Scan the current project:

arid .

Exit codes

Arid uses predictable exit codes for CLI and CI usage:

Exit code Meaning
0 Scan completed successfully and no duplicate findings failed the scan.
1 Duplicate-code findings were reported.
2 Invocation, configuration, parsing, or internal error.

A finding exit status is therefore distinct from an Arid execution failure.


License

Licensed under either of:

  • Apache License, Version 2.0
  • MIT License

at your option.


Contributing

Arid is in early development. Contribution guidelines will be added once the initial architecture and v1 behavior are established.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

arid-0.1.0a3-py3-none-win_amd64.whl (2.0 MB view details)

Uploaded Python 3Windows x86-64

arid-0.1.0a3-py3-none-manylinux_2_34_x86_64.whl (2.3 MB view details)

Uploaded Python 3manylinux: glibc 2.34+ x86-64

arid-0.1.0a3-py3-none-macosx_11_0_arm64.whl (2.1 MB view details)

Uploaded Python 3macOS 11.0+ ARM64

arid-0.1.0a3-py3-none-macosx_10_12_x86_64.whl (2.2 MB view details)

Uploaded Python 3macOS 10.12+ x86-64

File details

Details for the file arid-0.1.0a3-py3-none-win_amd64.whl.

File metadata

  • Download URL: arid-0.1.0a3-py3-none-win_amd64.whl
  • Upload date:
  • Size: 2.0 MB
  • Tags: Python 3, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for arid-0.1.0a3-py3-none-win_amd64.whl
Algorithm Hash digest
SHA256 7b0ddec594b4b6bd26a344f55aafa116d77cd7662beff506c12d433df4a4655c
MD5 4629676ac92f0b87c9b53e371597dd75
BLAKE2b-256 1a136927d2a95fc0fd4de37f0113907bef9ce81cac4d91cf61ec1b64456a79df

See more details on using hashes here.

Provenance

The following attestation bundles were made for arid-0.1.0a3-py3-none-win_amd64.whl:

Publisher: release.yml on sponge-b0b/arid

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file arid-0.1.0a3-py3-none-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for arid-0.1.0a3-py3-none-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 75b17dc6378a3381f3fe9c177b49a0198e91e6ff53afe77797759f73fc4116c6
MD5 2a79e765401ce828e6dd2e091a17b518
BLAKE2b-256 38c909ef62be8afb35090c87dc90ebeb5649d6ec332b2d1c8fd77cf328c6ebe0

See more details on using hashes here.

Provenance

The following attestation bundles were made for arid-0.1.0a3-py3-none-manylinux_2_34_x86_64.whl:

Publisher: release.yml on sponge-b0b/arid

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file arid-0.1.0a3-py3-none-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for arid-0.1.0a3-py3-none-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 d69c5f1a4619d24bd7f5f0e20ed9c7ed40eecb61d3a67e96a7e6c2d4b4c9e10c
MD5 a23be68461c443602b55350befb50859
BLAKE2b-256 26d488892487ddc29d26882ee78724e49a350e99616b5f99939d6db511031253

See more details on using hashes here.

Provenance

The following attestation bundles were made for arid-0.1.0a3-py3-none-macosx_11_0_arm64.whl:

Publisher: release.yml on sponge-b0b/arid

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file arid-0.1.0a3-py3-none-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for arid-0.1.0a3-py3-none-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 57d49080110842d06c4fe51d6b611831d7e4e6211f7c1868ab78eb37b1f80784
MD5 248e126120018eef7c978b34a0e4203d
BLAKE2b-256 8c0bea641e1d5dcf296f940edcd1e459e457170fc9964b8fcac047bd21f926d2

See more details on using hashes here.

Provenance

The following attestation bundles were made for arid-0.1.0a3-py3-none-macosx_10_12_x86_64.whl:

Publisher: release.yml on sponge-b0b/arid

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page