Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

Arid banner

Status: Release Candidate Scope: Python-only Purpose: Duplicate-code detection Config: tool.arid License: MIT OR Apache-2.0

Arid icon Arid

Fast Python duplicate-code checker written in Rust. A focused replacement for Pylint R0801 that complements Ruff.

What is Arid? · Project status · Goals · Usage · Output · Configuration · Pre-commit · Architecture · License


Project status

[!IMPORTANT] Arid is currently a release candidate. The release feature set and core interfaces are frozen, and the current build is believed ready for stable release without product-code changes.

Arid is a small, focused CLI for one job:

Detect duplicated Python source code quickly and accurately.


What is Arid?

Arid is a Python-specific CLI for duplicate-code detection.

It is designed to replace the duplicate-code functionality of Pylint R0801 / symilar without turning into another general-purpose linter. Arid is intentionally narrow in scope and is meant to run alongside Ruff, not compete with it.

Ruff
├── linting
├── formatting
├── imports
├── modernization
└── general code quality

Arid
└── duplicate-code detection

Why Arid? Because duplicated code isn't DRY.


Why not just use Pylint?

Pylint's R0801 checker provides useful Python-aware duplicate-code detection, but duplicate analysis can become very slow on larger codebases.

Arid aims to preserve the useful behavior of R0801 while using a Rust-native architecture designed specifically for duplicate detection.

The goal is not bug-for-bug compatibility. Where Pylint relies on textual heuristics, Arid prefers correct Python syntax interpretation.


Why not just use jscpd?

jscpd is a capable multi-language copy/paste detector, and its current implementation is also written in Rust.

Arid occupies a narrower niche:

  • Python only
  • focused on Pylint-style duplicate-code semantics
  • Python-aware filtering for comments, docstrings, imports, and signatures
  • designed to fit naturally into modern Python workflows
  • intentionally minimal in scope

Arid is not intended to replace jscpd for multi-language repositories.


Goals

Arid v1 is designed to:

  • detect duplicated Python source blocks across files
  • detect duplicated blocks within the same file
  • ignore comments, docstrings, imports, and function signatures when configured
  • preserve accurate original source locations
  • report concise DUP001 diagnostics
  • describe duplicate findings using Python structural context
  • provide deterministic duplication metrics
  • support pyproject.toml configuration via [tool.arid]
  • provide deterministic text, JSON, Markdown, and SARIF output
  • support baseline-based incremental adoption
  • integrate with pre-commit while preserving whole-project detection
  • run substantially faster than Pylint's duplicate-code checker
  • require no Python runtime to analyze Python source

Non-goals

Arid is intentionally not a general-purpose linter.

It does not aim to provide:

  • formatting
  • import sorting
  • type checking
  • dead-code detection
  • complexity analysis
  • security scanning
  • semantic clone detection
  • structural clone matching
  • fuzzy AST similarity
  • multi-language duplicate detection

If a feature belongs naturally in Ruff, it does not belong in Arid.


Usage

Arid is intended to fit naturally into a Python quality workflow:

ruff check .
arid .

Scan specific files or directories:

arid src tests

Require a larger duplicate before reporting it:

arid . --min-lines 8

Override normalization behavior for a single scan:

arid . --no-ignore-docstrings

Configurable boolean options support both positive and negative forms:

--ignore-comments       --no-ignore-comments
--ignore-docstrings     --no-ignore-docstrings
--ignore-imports        --no-ignore-imports
--ignore-signatures     --no-ignore-signatures
--same-file             --no-same-file
--hidden                --no-hidden

This allows command-line arguments to explicitly override either value from pyproject.toml.

Hidden files and directories are skipped by default during directory discovery. Include them when needed with:

arid . --hidden

This allows Arid to scan Python files under hidden directories such as .github/ while still honoring .gitignore and configured exclude patterns.

Exclude matching paths:

arid . --exclude 'generated/**'

--exclude may be repeated:

arid . \
    --exclude 'generated/**' \
    --exclude 'vendor/**'

Include the original source in each finding:

arid . --show-source

Choose an output format:

arid . --format text
arid . --format json
arid . --format markdown
arid . --format sarif

text is the default. The existing JSON shorthand remains supported:

arid . --json

Control text color explicitly when needed:

arid . --color auto
arid . --color always
arid . --color never

Create a baseline for existing duplicate debt:

arid . --write-baseline arid-baseline.json

Then enforce it explicitly:

arid . --baseline arid-baseline.json

or configure it in [tool.arid] so normal arid . scans enforce the baseline automatically.

Example diagnostic:

DUP001 4 duplicated lines
Context: declarative
Scope: class
Occurrences: 2 across 2 files (cross-file)

  src/models/user.py:12-15
  src/models/account.py:20-23

Found 1 duplicate group.
4 duplicate lines (2.31%).

Understanding Arid's output

Arid separates two questions:

Detection answers: "Is this code duplicated?"
Context helps answer: "What kind of code is duplicated?"

Arid deliberately does not assign a severity or decide whether duplication should be removed. Duplicate code can be intentional, harmless, framework-driven, or worth refactoring.

The structural metadata exists to help you make that decision.

Consider:

DUP001 4 duplicated lines
Context: declarative
Scope: class
Occurrences: 2 across 2 files (cross-file)

  src/models/user.py:12-15
  src/models/account.py:20-23

Found 1 duplicate group.
4 duplicate lines (2.31%).

DUP001 4 duplicated lines

DUP001 is Arid's duplicate-code diagnostic.

4 duplicated lines means the matching region contains four effective normalized lines that satisfy the configured duplicate threshold.

Arid compares source after its configured Python-aware normalization. Depending on configuration, this can remove constructs such as:

  • comments
  • docstrings
  • imports
  • function signatures

Blank lines do not count toward min-lines, and lines containing only non-substantive punctuation do not increase the effective-line count.

Because of that, a finding reported as four duplicated lines may span more than four physical source lines.

Context

Context describes the structural kind of Python code involved in the duplicate.

Possible values are:

Context Meaning
declarative The duplicate consists of declarations or definitions, such as direct module/class assignments or definitions.
executable The duplicate consists of executable statements, control flow, or function-body logic.
mixed The duplicate contains or occurs across more than one structural context.

For example:

Context: declarative

often appears for repeated class or module definitions.

Context: executable

often appears for repeated application logic inside functions.

[!NOTE] Context is descriptive, not a severity. declarative does not mean "safe to ignore," and executable does not mean "must refactor."

Arid describes Python structure without attempting to infer framework semantics or developer intent.

It therefore does not label findings as "ORM boilerplate," "configuration noise," "safe duplication," or similar framework-specific categories.

Scope

Scope describes where the duplicated code occurs structurally.

Possible values are:

Scope Meaning
module Module-level code.
class Code structurally associated with a class.
function Code structurally associated with a function or method.
mixed The duplicate spans or occurs across more than one scope.

For example:

Context: executable
Scope: function

indicates repeated executable logic within functions or methods.

By contrast:

Context: declarative
Scope: class

indicates repeated declarative code associated with classes.

Again, scope describes where the duplicate exists, not whether it is a problem.

Occurrences

The occurrence line tells you how widely the duplicate appears.

Occurrences: 2 across 2 files (cross-file)

contains three pieces of information:

  • the number of duplicate occurrences
  • the number of distinct files containing them
  • how those occurrences are distributed

Distribution values are:

Distribution Meaning
same-file All occurrences are contained in one file.
cross-file Occurrences are spread across multiple files, with one occurrence in each involved file.
mixed Multiple files are involved and at least one file contains multiple occurrences.

Examples:

Occurrences: 2 across 1 file (same-file)

means the same block appears twice in one file.

Occurrences: 3 across 3 files (cross-file)

means one occurrence appears in each of three files.

Occurrences: 4 across 3 files (mixed)

means the duplicate spans multiple files and at least one of those files contains more than one occurrence.

Source locations

Locations such as:

src/models/user.py:12-15

always refer to the original physical Python source, not Arid's internal normalized representation.

This remains true even when ignored comments, imports, signatures, docstrings, or blank lines appear within the physical range.

Use:

arid . --show-source

to include the original source text alongside each location.

Duplicate groups

When the same block appears more than twice, Arid reports it as one duplicate group rather than generating every possible pair.

For example, a block appearing in:

a.py
b.py
c.py

is one finding with three occurrences, not three separate pairwise findings.

Duplicate lines and duplication percentage

The final summary:

4 duplicate lines (2.31%).

measures redundant effective lines, not every line participating in a duplicate.

One occurrence of each duplicate group is treated as canonical. Only redundant copies beyond that canonical occurrence contribute duplicate lines.

For example:

10-line block × 2 occurrences

contributes:

10 duplicate lines

not 20.

A 10-line block appearing three times contributes:

20 duplicate lines

because two of the three copies are redundant.

Overlapping redundant regions are not counted repeatedly.

The duplication percentage is:

duplicate effective lines
───────────────────────── × 100
 analyzed effective lines

This makes the metric represent how much analyzed code is redundant rather than how much code merely participates in a duplicated region.

How to interpret findings

There is no universal rule for which duplicate should be refactored first, but Arid's metadata can help you triage a large report.

A practical review order is often:

  1. Look at larger duplicate regions before very short ones.
  2. Review executable / function findings for repeated application logic.
  3. Look at findings with many occurrences to identify patterns repeated broadly through the codebase.
  4. Use same-file, cross-file, and mixed to distinguish localized repetition from code repeated across modules.
  5. Review declarative findings in context. Repeated declarations may be intentional, generated by a common coding pattern, or candidates for consolidation depending on the project.

Arid intentionally stops short of saying:

high severity
low value
safe to ignore
must refactor

Those are project-specific judgments.

Its job is to provide accurate duplicate detection and enough objective structural information for the developer to make them.


Configuration

Arid uses [tool.arid] in pyproject.toml:

[tool.arid]
min-lines = 4
ignore-comments = true
ignore-docstrings = true
ignore-imports = true
ignore-signatures = true
same-file = true
hidden = false
exclude = [
    "generated/**",
    "vendor/**",
]

Current defaults are:

Option Default Meaning
min-lines 4 Minimum effective normalized lines required for a duplicate.
ignore-comments true Ignore Python comments during matching.
ignore-docstrings true Ignore structural Python docstrings.
ignore-imports true Ignore import statements.
ignore-signatures true Ignore function and method declaration signatures.
same-file true Detect non-overlapping duplicate regions within the same file.
hidden false Include hidden files and directories during directory discovery.
exclude [] Path patterns excluded from discovery.
baseline none Optional baseline file used to accept existing duplicate debt while reporting new debt.

Configuration precedence is:

CLI arguments
    ↓
pyproject.toml
    ↓
built-in defaults

For example:

[tool.arid]
min-lines = 6
ignore-docstrings = true
same-file = true
hidden = false

can be overridden for one scan with:

arid . \
    --min-lines 10 \
    --no-ignore-docstrings \
    --no-same-file \
    --hidden

Each configurable boolean has both an enabling and disabling CLI form. This matters when the project configuration differs from the built-in default. For example, if the project contains:

[tool.arid]
ignore-comments = false

then:

arid . --ignore-comments

explicitly enables comment filtering for that scan.

Likewise:

arid . --no-ignore-comments

explicitly disables it.

Supplying one or more --exclude options on the command line overrides the configured exclude list for that scan:

arid . \
    --exclude 'build/**' \
    --exclude 'generated/**'

To enforce an existing baseline on every normal scan:

[tool.arid]
baseline = "arid-baseline.json"

An explicit --baseline path overrides the configured baseline for that scan. --format, --color, --json, --write-baseline, and --show-source are CLI-only presentation or administrative options.


Pre-commit

Arid provides an official pre-commit hook that runs a whole-project arid . scan rather than limiting duplicate detection to staged Python files.

Arid must already be installed and available as arid on PATH, and the official hook requires pre-commit 4.4.0 or newer.

repos:
  - repo: https://github.com/sponge-b0b/arid
    rev: v1.1.0
    hooks:
      - id: arid

The hook honors normal [tool.arid] configuration, including baseline = "arid-baseline.json".

See Arid pre-commit integration for installation details and behavior.


Detection model

Arid is focused on exact duplicate source blocks after configurable Python-aware normalization.

For example, with comments and function signatures ignored:

def first():
    # explanation
    value = calculate_value()
    save_value(value)

and:

def second():
    # different explanation
    value = calculate_value()
    save_value(value)

can be considered duplicates.

Arid v1 does not normalize identifiers, so these are intentionally different:

value = calculate_value()
save_value(value)
result = calculate_value()
save_value(result)

Arid can attach structural context such as declarative, executable, class, or function to a duplicate that it has already detected.

That does not make Arid a structural clone detector. Two pieces of code that are merely structurally similar but do not become identical after normalization are not considered duplicates.

Semantic clone detection, identifier-renaming clone detection, and fuzzy AST similarity remain outside the v1 scope.


Suppressing intentional duplication

Arid supports source-level suppression regions:

# arid: disable

# intentionally duplicated code

# arid: enable

Code inside a disabled region does not participate in duplicate detection.

Suppression regions also create matching boundaries, so Arid does not construct a duplicate across disabled source.

Use suppression for duplication that is intentionally accepted by the project rather than expecting Arid to infer whether a particular framework pattern or coding convention should be ignored.


Architecture

The v1 architecture is intentionally small:

discover
   ↓
parse
   ↓
normalize
   ↓
intern lines
   ↓
suffix array
   ↓
LCP
   ↓
maximal repeats
   ↓
DUP001

Arid analyzes Python source entirely in Rust and never imports or executes the project being scanned.

Duplicate detection operates on Arid's normalized source representation. Structural context is derived from Python syntax and attached as reporting metadata; it does not alter whether two normalized regions match.


Installation

[!NOTE] See Project status for the current release stage and stability expectations.

uv

Install Arid as an isolated command-line tool:

uv tool install arid

pip

Install Arid from PyPI:

python -m pip install --pre arid

Verify the installation:

arid --version

Scan the current project:

arid .

Exit codes

Arid uses predictable exit codes for CLI and CI usage:

Exit code Meaning
0 Scan completed successfully and no duplicate findings failed the scan.
1 Duplicate-code findings were reported.
2 Invocation, configuration, parsing, or internal error.

A finding exit status is therefore distinct from an Arid execution failure.


License

Licensed under either of:

  • Apache License, Version 2.0
  • MIT License

at your option.


Development

Arid includes dedicated tooling and documentation for release qualification, performance benchmarking, and real-world validation:

  • Release qualification — automated acceptance of published release candidates and stable releases, including artifact validation, equivalence checks, benchmarks, and stable-promotion enforcement.
  • Benchmarks — reproducible performance comparisons against Pylint R0801 and jscpd, including corpus provisioning and benchmark execution.
  • Validation — real-world correctness and robustness validation across Black, Django, mypy, Rich, determinism checks, malformed-source handling, and filesystem edge cases.
  • Release roadmap — release stages, qualification gates, and release metadata preparation with ./release.sh.

Contributing

Contributions should preserve Arid's focused scope and existing product contract. Bug fixes, compatibility improvements, tests, documentation, and performance work are welcome; scope-expanding features should be discussed before implementation.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

arid-1.2.0rc1-py3-none-win_amd64.whl (2.1 MB view details)

Uploaded Python 3Windows x86-64

arid-1.2.0rc1-py3-none-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (2.4 MB view details)

Uploaded Python 3manylinux: glibc 2.17+ x86-64

arid-1.2.0rc1-py3-none-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (2.3 MB view details)

Uploaded Python 3manylinux: glibc 2.17+ ARM64

arid-1.2.0rc1-py3-none-macosx_11_0_arm64.whl (2.2 MB view details)

Uploaded Python 3macOS 11.0+ ARM64

arid-1.2.0rc1-py3-none-macosx_10_12_x86_64.whl (2.3 MB view details)

Uploaded Python 3macOS 10.12+ x86-64

File details

Details for the file arid-1.2.0rc1-py3-none-win_amd64.whl.

File metadata

  • Download URL: arid-1.2.0rc1-py3-none-win_amd64.whl
  • Upload date:
  • Size: 2.1 MB
  • Tags: Python 3, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for arid-1.2.0rc1-py3-none-win_amd64.whl
Algorithm Hash digest
SHA256 19292787979c82c3e8c64b55af8fe1e0c58d8b2196fb14a3a26c94f2787c7568
MD5 150d9b4f45fdaae5d4fea7d79acd4564
BLAKE2b-256 ee0676813a94273ee3b5007c79bddb83314217e198c8dccfc60289f5cf45d57b

See more details on using hashes here.

Provenance

The following attestation bundles were made for arid-1.2.0rc1-py3-none-win_amd64.whl:

Publisher: release.yml on sponge-b0b/arid

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file arid-1.2.0rc1-py3-none-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for arid-1.2.0rc1-py3-none-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 47c1965c341e8ffbfa84b78681e9b6c44b96e25471050f5d8e4555e9062107fe
MD5 dad79443ae3d14101057513bba15577c
BLAKE2b-256 38f98c3035c1d4eb02584710c0f1d4f6ae6c926cfc608a2a1c28aa0c7bc6c283

See more details on using hashes here.

Provenance

The following attestation bundles were made for arid-1.2.0rc1-py3-none-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: release.yml on sponge-b0b/arid

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file arid-1.2.0rc1-py3-none-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for arid-1.2.0rc1-py3-none-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 5af166163757679424e8df85812e461ce80f390dbcd58c49feed2fecfe203ac1
MD5 96b9aa03e7af49028c24b23c0fa8f924
BLAKE2b-256 10cf87f839a8298de5d32752a35d1e9f4597c19495f7b05584c31c61e8b0df86

See more details on using hashes here.

Provenance

The following attestation bundles were made for arid-1.2.0rc1-py3-none-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:

Publisher: release.yml on sponge-b0b/arid

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file arid-1.2.0rc1-py3-none-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for arid-1.2.0rc1-py3-none-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 e4e33fa23bbe5b5ed4f0d25871acf9e7a9726c6d7a8b6df4af8cd3fbb396fe64
MD5 f5aaea9fe882edbbf051e239f95acb51
BLAKE2b-256 4721b591dde49f01eab63e26be093417623dbebf4ac1a500f4d63b97cde0a9d3

See more details on using hashes here.

Provenance

The following attestation bundles were made for arid-1.2.0rc1-py3-none-macosx_11_0_arm64.whl:

Publisher: release.yml on sponge-b0b/arid

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file arid-1.2.0rc1-py3-none-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for arid-1.2.0rc1-py3-none-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 d0cdabda7afa39c79c6e9064ee9369c2281d4d189ac4e13fd9658ddc775b6683
MD5 8fa969d4d2c10108a813fa270a00126e
BLAKE2b-256 720683e42bb352181e8d9f0a790f5584113f6beec3ab493c475899836c9abe76

See more details on using hashes here.

Provenance

The following attestation bundles were made for arid-1.2.0rc1-py3-none-macosx_10_12_x86_64.whl:

Publisher: release.yml on sponge-b0b/arid

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page