Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

Arid banner

Status: Alpha Scope: Python-only Purpose: Duplicate-code detection Config: tool.arid License: MIT OR Apache-2.0

Arid icon Arid

Fast Python duplicate-code checker written in Rust. A focused replacement for Pylint R0801 that complements Ruff.

What is Arid? · Project status · Goals · Usage · Output · Schemas · Configuration · Pre-commit · Architecture · License


Project status

[!IMPORTANT] Arid is currently in alpha. The intended release functionality is code complete but remains under stabilization. Interfaces, behavior, defaults, and packaging details may still change before stable release.

Arid is a small, focused CLI for one job:

Detect duplicated Python source code quickly and accurately.


What is Arid?

Arid is a Python-specific CLI for duplicate-code detection.

It is designed to replace the duplicate-code functionality of Pylint R0801 / symilar without turning into another general-purpose linter. Arid is intentionally narrow in scope and is meant to run alongside Ruff, not compete with it.

Ruff
├── linting
├── formatting
├── imports
├── modernization
└── general code quality

Arid
└── duplicate-code detection

Why Arid? Because duplicated code isn't DRY.


Why not just use Pylint?

Pylint's R0801 checker provides useful Python-aware duplicate-code detection, but duplicate analysis can become very slow on larger codebases.

Arid aims to preserve the useful behavior of R0801 while using a Rust-native architecture designed specifically for duplicate detection.

The goal is not bug-for-bug compatibility. Where Pylint relies on textual heuristics, Arid prefers correct Python syntax interpretation.


Why not just use jscpd?

jscpd is a capable multi-language copy/paste detector, and its current implementation is also written in Rust.

Arid occupies a narrower niche:

  • Python only
  • focused on Pylint-style duplicate-code semantics
  • Python-aware filtering for comments, docstrings, imports, and signatures
  • designed to fit naturally into modern Python workflows
  • intentionally minimal in scope

Arid is not intended to replace jscpd for multi-language repositories.


Goals

Arid v1 is designed to:

  • detect duplicated Python source blocks across files
  • detect duplicated blocks within the same file
  • ignore comments, docstrings, imports, and function signatures when configured
  • preserve accurate original source locations
  • report concise DUP001 diagnostics
  • describe duplicate findings using Python structural context
  • provide deterministic duplication metrics
  • support pyproject.toml configuration via [tool.arid]
  • provide deterministic text, JSON, Markdown, and SARIF output
  • support baseline-based incremental adoption
  • support opt-in parallel file preparation while remaining serial by default
  • publish JSON Schemas for the report and baseline machine contracts
  • publish release artifacts for Linux x86_64 and ARM64, macOS x86_64 and ARM64, and Windows x86_64
  • integrate with pre-commit while preserving whole-project detection
  • run substantially faster than Pylint's duplicate-code checker
  • require no Python runtime to analyze Python source

Non-goals

Arid is intentionally not a general-purpose linter.

It does not aim to provide:

  • formatting
  • import sorting
  • type checking
  • dead-code detection
  • complexity analysis
  • security scanning
  • semantic clone detection
  • structural clone matching
  • fuzzy AST similarity
  • multi-language duplicate detection

If a feature belongs naturally in Ruff, it does not belong in Arid.


Usage

Arid is intended to fit naturally into a Python quality workflow:

ruff check .
arid .

Scan specific files or directories:

arid src tests

Require a larger duplicate before reporting it:

arid . --min-lines 8

Override normalization behavior for a single scan:

arid . --no-ignore-docstrings

Configurable boolean options support both positive and negative forms:

--ignore-comments       --no-ignore-comments
--ignore-docstrings     --no-ignore-docstrings
--ignore-imports        --no-ignore-imports
--ignore-signatures     --no-ignore-signatures
--same-file             --no-same-file
--hidden                --no-hidden

This allows command-line arguments to explicitly override either value from pyproject.toml.

Hidden files and directories are skipped by default during directory discovery. Include them when needed with:

arid . --hidden

This allows Arid to scan Python files under hidden directories such as .github/ while still honoring .gitignore and configured exclude patterns.

Exclude matching paths:

arid . --exclude 'generated/**'

--exclude may be repeated:

arid . \
    --exclude 'generated/**' \
    --exclude 'vendor/**'

Parallelism

Arid runs serially by default. Running:

arid .

is equivalent to:

arid . --workers 1

For larger projects, file preparation can run in parallel with an explicit worker count:

arid . --workers 4

Arid 1.2 also supports bounded automatic worker selection:

arid . --workers auto

auto chooses a conservative worker count capped at 4 and further bounded by available parallelism and the number of discovered Python files.

Parallelism applies only to file preparation: reading, parsing, and normalization. It does not change duplicate-detection semantics. Serial, numeric-worker, and auto execution produce the same findings, metrics, report ordering, and exit status.

Worker selection is intentionally a CLI-only execution setting and cannot be configured in [tool.arid].

Include the original source in each finding:

arid . --show-source

Choose an output format:

arid . --format text
arid . --format json
arid . --format markdown
arid . --format sarif

text is the default. The existing JSON shorthand remains supported:

arid . --json

Control text color explicitly when needed:

arid . --color auto
arid . --color always
arid . --color never

Create a baseline for existing duplicate debt:

arid . --write-baseline arid-baseline.json

Then enforce it explicitly:

arid . --baseline arid-baseline.json

or configure it in [tool.arid] so normal arid . scans enforce the baseline automatically.

Example diagnostic:

DUP001 4 duplicated lines
Context: declarative
Scope: class
Occurrences: 2 across 2 files (cross-file)

  src/models/user.py:12-15
  src/models/account.py:20-23

Found 1 duplicate group.
4 duplicate lines (2.31%).

Understanding Arid's output

Arid separates two questions:

Detection answers: "Is this code duplicated?"
Context helps answer: "What kind of code is duplicated?"

Arid deliberately does not assign a severity or decide whether duplication should be removed. Duplicate code can be intentional, harmless, framework-driven, or worth refactoring.

The structural metadata exists to help you make that decision.

Consider:

DUP001 4 duplicated lines
Context: declarative
Scope: class
Occurrences: 2 across 2 files (cross-file)

  src/models/user.py:12-15
  src/models/account.py:20-23

Found 1 duplicate group.
4 duplicate lines (2.31%).

DUP001 4 duplicated lines

DUP001 is Arid's duplicate-code diagnostic.

4 duplicated lines means the matching region contains four effective normalized lines that satisfy the configured duplicate threshold.

Arid compares source after its configured Python-aware normalization. Depending on configuration, this can remove constructs such as:

  • comments
  • docstrings
  • imports
  • function signatures

Blank lines do not count toward min-lines, and lines containing only non-substantive punctuation do not increase the effective-line count.

Because of that, a finding reported as four duplicated lines may span more than four physical source lines.

Context

Context describes the structural kind of Python code involved in the duplicate.

Possible values are:

Context Meaning
declarative The duplicate consists of declarations or definitions, such as direct module/class assignments or definitions.
executable The duplicate consists of executable statements, control flow, or function-body logic.
mixed The duplicate contains or occurs across more than one structural context.

For example:

Context: declarative

often appears for repeated class or module definitions.

Context: executable

often appears for repeated application logic inside functions.

[!NOTE] Context is descriptive, not a severity. declarative does not mean "safe to ignore," and executable does not mean "must refactor."

Arid describes Python structure without attempting to infer framework semantics or developer intent.

It therefore does not label findings as "ORM boilerplate," "configuration noise," "safe duplication," or similar framework-specific categories.

Scope

Scope describes where the duplicated code occurs structurally.

Possible values are:

Scope Meaning
module Module-level code.
class Code structurally associated with a class.
function Code structurally associated with a function or method.
mixed The duplicate spans or occurs across more than one scope.

For example:

Context: executable
Scope: function

indicates repeated executable logic within functions or methods.

By contrast:

Context: declarative
Scope: class

indicates repeated declarative code associated with classes.

Again, scope describes where the duplicate exists, not whether it is a problem.

Occurrences

The occurrence line tells you how widely the duplicate appears.

Occurrences: 2 across 2 files (cross-file)

contains three pieces of information:

  • the number of duplicate occurrences
  • the number of distinct files containing them
  • how those occurrences are distributed

Distribution values are:

Distribution Meaning
same-file All occurrences are contained in one file.
cross-file Occurrences are spread across multiple files, with one occurrence in each involved file.
mixed Multiple files are involved and at least one file contains multiple occurrences.

Examples:

Occurrences: 2 across 1 file (same-file)

means the same block appears twice in one file.

Occurrences: 3 across 3 files (cross-file)

means one occurrence appears in each of three files.

Occurrences: 4 across 3 files (mixed)

means the duplicate spans multiple files and at least one of those files contains more than one occurrence.

Source locations

Locations such as:

src/models/user.py:12-15

always refer to the original physical Python source, not Arid's internal normalized representation.

This remains true even when ignored comments, imports, signatures, docstrings, or blank lines appear within the physical range.

Use:

arid . --show-source

to include the original source text alongside each location.

Duplicate groups

When the same block appears more than twice, Arid reports it as one duplicate group rather than generating every possible pair.

For example, a block appearing in:

a.py
b.py
c.py

is one finding with three occurrences, not three separate pairwise findings.

Duplicate lines and duplication percentage

The final summary:

4 duplicate lines (2.31%).

measures redundant effective lines, not every line participating in a duplicate.

One occurrence of each duplicate group is treated as canonical. Only redundant copies beyond that canonical occurrence contribute duplicate lines.

For example:

10-line block × 2 occurrences

contributes:

10 duplicate lines

not 20.

A 10-line block appearing three times contributes:

20 duplicate lines

because two of the three copies are redundant.

Overlapping redundant regions are not counted repeatedly.

The duplication percentage is:

duplicate effective lines
───────────────────────── × 100
 analyzed effective lines

This makes the metric represent how much analyzed code is redundant rather than how much code merely participates in a duplicated region.

How to interpret findings

There is no universal rule for which duplicate should be refactored first, but Arid's metadata can help you triage a large report.

A practical review order is often:

  1. Look at larger duplicate regions before very short ones.
  2. Review executable / function findings for repeated application logic.
  3. Look at findings with many occurrences to identify patterns repeated broadly through the codebase.
  4. Use same-file, cross-file, and mixed to distinguish localized repetition from code repeated across modules.
  5. Review declarative findings in context. Repeated declarations may be intentional, generated by a common coding pattern, or candidates for consolidation depending on the project.

Arid intentionally stops short of saying:

high severity
low value
safe to ignore
must refactor

Those are project-specific judgments.

Its job is to provide accurate duplicate detection and enough objective structural information for the developer to make them.


Machine-readable schemas

Arid publishes JSON Schema documents for its own machine-readable contracts:

Published schema files describe versioned compatibility contracts. A future incompatible report or baseline format receives a new schema version rather than silently changing an existing schema file.

SARIF output remains SARIF 2.1.0 and uses the official SARIF schema rather than an Arid-owned schema.


Configuration

Arid uses [tool.arid] in pyproject.toml:

[tool.arid]
min-lines = 4
ignore-comments = true
ignore-docstrings = true
ignore-imports = true
ignore-signatures = true
same-file = true
hidden = false
exclude = [
    "generated/**",
    "vendor/**",
]

Current defaults are:

Option Default Meaning
min-lines 4 Minimum effective normalized lines required for a duplicate.
ignore-comments true Ignore Python comments during matching.
ignore-docstrings true Ignore structural Python docstrings.
ignore-imports true Ignore import statements.
ignore-signatures true Ignore function and method declaration signatures.
same-file true Detect non-overlapping duplicate regions within the same file.
hidden false Include hidden files and directories during directory discovery.
exclude [] Path patterns excluded from discovery.
baseline none Optional baseline file used to accept existing duplicate debt while reporting new debt.

Configuration precedence is:

CLI arguments
    ↓
pyproject.toml
    ↓
built-in defaults

For example:

[tool.arid]
min-lines = 6
ignore-docstrings = true
same-file = true
hidden = false

can be overridden for one scan with:

arid . \
    --min-lines 10 \
    --no-ignore-docstrings \
    --no-same-file \
    --hidden

Each configurable boolean has both an enabling and disabling CLI form. This matters when the project configuration differs from the built-in default. For example, if the project contains:

[tool.arid]
ignore-comments = false

then:

arid . --ignore-comments

explicitly enables comment filtering for that scan.

Likewise:

arid . --no-ignore-comments

explicitly disables it.

Supplying one or more --exclude options on the command line overrides the configured exclude list for that scan:

arid . \
    --exclude 'build/**' \
    --exclude 'generated/**'

To enforce an existing baseline on every normal scan:

[tool.arid]
baseline = "arid-baseline.json"

An explicit --baseline path overrides the configured baseline for that scan. --workers, --format, --color, --json, --write-baseline, and --show-source are CLI-only execution, presentation, or administrative options.


Pre-commit

Arid provides an official pre-commit hook that runs a whole-project arid . scan rather than limiting duplicate detection to staged Python files.

Arid must already be installed and available as arid on PATH, and the official hook requires pre-commit 4.4.0 or newer.

repos:
  - repo: https://github.com/sponge-b0b/arid
    rev: v1.2.0
    hooks:
      - id: arid

The hook honors normal [tool.arid] configuration, including baseline = "arid-baseline.json".

See Arid pre-commit integration for installation details and behavior.


Detection model

Arid is focused on exact duplicate source blocks after configurable Python-aware normalization.

For example, with comments and function signatures ignored:

def first():
    # explanation
    value = calculate_value()
    save_value(value)

and:

def second():
    # different explanation
    value = calculate_value()
    save_value(value)

can be considered duplicates.

Arid v1 does not normalize identifiers, so these are intentionally different:

value = calculate_value()
save_value(value)
result = calculate_value()
save_value(result)

Arid can attach structural context such as declarative, executable, class, or function to a duplicate that it has already detected.

That does not make Arid a structural clone detector. Two pieces of code that are merely structurally similar but do not become identical after normalization are not considered duplicates.

Semantic clone detection, identifier-renaming clone detection, and fuzzy AST similarity remain outside the v1 scope.


Suppressing intentional duplication

Arid supports source-level suppression regions:

# arid: disable

# intentionally duplicated code

# arid: enable

Code inside a disabled region does not participate in duplicate detection.

Suppression regions also create matching boundaries, so Arid does not construct a duplicate across disabled source.

Use suppression for duplication that is intentionally accepted by the project rather than expecting Arid to infer whether a particular framework pattern or coding convention should be ignored.


Architecture

The v1 architecture is intentionally small:

discover
   ↓
parse
   ↓
normalize
   ↓
intern lines
   ↓
suffix array
   ↓
LCP
   ↓
maximal repeats
   ↓
DUP001

Arid analyzes Python source entirely in Rust and never imports or executes the project being scanned.

Duplicate detection operates on Arid's normalized source representation. Structural context is derived from Python syntax and attached as reporting metadata; it does not alter whether two normalized regions match.


Installation

[!NOTE] See Project status for the current release stage and stability expectations.

uv

Install Arid as an isolated command-line tool:

uv tool install arid

pip

Install Arid from PyPI:

python -m pip install arid

Published platforms

Arid publishes release wheels and standalone archives for:

  • Linux x86_64
  • Linux ARM64 (aarch64)
  • macOS x86_64
  • macOS ARM64 (Apple silicon)
  • Windows x86_64

Linux release artifacts target manylinux_2_17 / glibc 2.17 compatibility.

Verify the installation:

arid --version

Scan the current project:

arid .

Exit codes

Arid uses predictable exit codes for CLI and CI usage:

Exit code Meaning
0 Scan completed successfully and no duplicate findings failed the scan.
1 Duplicate-code findings were reported.
2 Invocation, configuration, parsing, or internal error.

A finding exit status is therefore distinct from an Arid execution failure.


License

Licensed under either of:

  • Apache License, Version 2.0
  • MIT License

at your option.


Development

Arid includes dedicated tooling and documentation for release qualification, performance benchmarking, and real-world validation:

  • Release qualification — automated acceptance of published release candidates and stable releases, including artifact validation, equivalence checks, benchmarks, and stable-promotion enforcement.
  • Benchmarks — reproducible performance comparisons against Pylint R0801 and jscpd, including corpus provisioning and benchmark execution.
  • Validation — real-world correctness and robustness validation across Black, Django, mypy, Rich, determinism checks, malformed-source handling, and filesystem edge cases.
  • Release roadmap — release stages, qualification gates, and release metadata preparation with ./release.sh.

Contributing

Contributions should preserve Arid's focused scope and existing product contract. Bug fixes, compatibility improvements, tests, documentation, and performance work are welcome; scope-expanding features should be discussed before implementation.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

arid-2.0.0a1-py3-none-win_amd64.whl (2.2 MB view details)

Uploaded Python 3Windows x86-64

arid-2.0.0a1-py3-none-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (2.5 MB view details)

Uploaded Python 3manylinux: glibc 2.17+ x86-64

arid-2.0.0a1-py3-none-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (2.4 MB view details)

Uploaded Python 3manylinux: glibc 2.17+ ARM64

arid-2.0.0a1-py3-none-macosx_11_0_arm64.whl (2.3 MB view details)

Uploaded Python 3macOS 11.0+ ARM64

arid-2.0.0a1-py3-none-macosx_10_12_x86_64.whl (2.3 MB view details)

Uploaded Python 3macOS 10.12+ x86-64

File details

Details for the file arid-2.0.0a1-py3-none-win_amd64.whl.

File metadata

  • Download URL: arid-2.0.0a1-py3-none-win_amd64.whl
  • Upload date:
  • Size: 2.2 MB
  • Tags: Python 3, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for arid-2.0.0a1-py3-none-win_amd64.whl
Algorithm Hash digest
SHA256 367c4b1a81e98ea0c3bcb96c7c78b96c247594a72cc653659177ce3272f9e793
MD5 81a0ecb98ea385659697b5b04c3dfa55
BLAKE2b-256 6cadd2059ab3b25414e18cdbec66481a0bb66f02d32c24bbba2508803dd4c687

See more details on using hashes here.

Provenance

The following attestation bundles were made for arid-2.0.0a1-py3-none-win_amd64.whl:

Publisher: release.yml on sponge-b0b/arid

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file arid-2.0.0a1-py3-none-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for arid-2.0.0a1-py3-none-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 313f1440d5f97f2c67a849d9f7ea04d38d021b8a93926af90587e4e861c6d6c0
MD5 5e4fdc977f31da498f1986846a5f7b7d
BLAKE2b-256 ec6a558cefc53fecb09997dee43c212a169366052988c89e2139db2efcd45739

See more details on using hashes here.

Provenance

The following attestation bundles were made for arid-2.0.0a1-py3-none-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: release.yml on sponge-b0b/arid

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file arid-2.0.0a1-py3-none-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for arid-2.0.0a1-py3-none-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 2d3cb0ece368bca7f55a3f0e1fb517987e64b6c61473e926e4336b162856c78b
MD5 a63f2069fbc9a3d953898c1bd64312aa
BLAKE2b-256 b4066b0cb6b05a03b99e122b8505954bdaa3408e41b88ebeb1ee057d64735c74

See more details on using hashes here.

Provenance

The following attestation bundles were made for arid-2.0.0a1-py3-none-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:

Publisher: release.yml on sponge-b0b/arid

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file arid-2.0.0a1-py3-none-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for arid-2.0.0a1-py3-none-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 051814c979285dc978828d03d99e8f974b826ea6f6f07e9c82f6087d75163dde
MD5 b4a4a12dad94cc34a59beb2c783aa107
BLAKE2b-256 347554601e0eb914182589bbab77b6b740430ad1bb4059b8773a35348c616be7

See more details on using hashes here.

Provenance

The following attestation bundles were made for arid-2.0.0a1-py3-none-macosx_11_0_arm64.whl:

Publisher: release.yml on sponge-b0b/arid

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file arid-2.0.0a1-py3-none-macosx_10_12_x86_64.whl.

File metadata

File hashes

Hashes for arid-2.0.0a1-py3-none-macosx_10_12_x86_64.whl
Algorithm Hash digest
SHA256 9b5135f27b19b82bcfff74b121466f9db948154c1a950f8ff8539092d42dc839
MD5 52717ff5c206fd505f2139c0766fade3
BLAKE2b-256 2044e6f6abe63d1ede249e9b337fa7f0d12e8408d5b2e472620d89090c946843

See more details on using hashes here.

Provenance

The following attestation bundles were made for arid-2.0.0a1-py3-none-macosx_10_12_x86_64.whl:

Publisher: release.yml on sponge-b0b/arid

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page