Skip to main content

duplicatecode

Static (LLM-free) detection of duplicate / similar code in Python and TypeScript (TSX), aimed at catching an LLM re-implementing something that already exists. Input can be a git diff checked against existing source.

  • crates/duplicatecode-engine – Rust engine (tree-sitter based)
  • crates/duplicatecode-cli – duplicatecode CLI
  • dataset/ – benchmark: coding tasks + diverging implementations (see dataset/README.md)

Install

uv tool install duplicatecode            # from PyPI (prebuilt binary, no Rust needed)
uvx duplicatecode scan .                 # or run without installing
curl -fsSL https://raw.githubusercontent.com/bmsuisse/duplicatecode/main/install.sh | sh   # standalone binary

Binaries for Linux/macOS/Windows (x86_64, plus aarch64 on Linux/macOS) are also attached to each GitHub Release. Releases are automatic: bump version in [workspace.package] in Cargo.toml, merge to main, and CI tags and publishes it.

Usage

cargo build --release
duplicatecode units src/                              # list extracted units
git diff origin/main | duplicatecode diff --repo .    # check added code against the repo
duplicatecode scan packages/ services/api           # find duplicate groups inside one or more folders (monorepo)
duplicatecode scan . --exclude '**/generated/**' --skip-tests --fail-on-found   # CI-style run
duplicatecode scan . --pairs --json                  # machine-readable pairs instead of groups
duplicatecode bench --dataset dataset [--file-level] [--mutations] [--negatives <other repo>]

How it works

Units (functions, methods, classes, arrow-function components) are extracted with tree-sitter and normalized (identifiers/literals abstracted; comments, docstrings and type annotations dropped). Each pair gets several signals: token 4-/2-gram overlap, node-kind histogram, literals, API/attribute names, name subwords, callees, and statement-level multiset/alignment scores. combined is a weighted sum (0.6 structure, 0.2 name, 0.2 callees). Constructors and dunder methods are ignored by default. An inverted index over k-grams keeps matching against large corpora fast.

Benchmark findings (Haiku vs Sonnet implementations of the same 50 tasks)

  • Ranking is good (right match is top-1 in ~70–95% of cases) but absolute detection at a 1% false-positive rate is modest for independently written code: ~35–50% (function level) of same-task pairs.
  • Comparing whole files (which removes helper-splitting noise) gives 76–100% recall, mostly thanks to names; structure-only signals reach ~35–55% on loose/hard tasks. Statement-level features are roughly on par with token k-grams; a fitted logistic model does not beat the hand weights in cross-validation.
  • Mechanical rewrites (rename, reorder, temp variable, logging, dead code) are handled well individually (86–100% of pairs ≥ 0.7); all combined drops the mean score to ~0.58, mostly because renames zero the name signal.

Normalization

Before comparing, debug output (print, logging.*, console.*) is dropped; Python comprehensions are read as the equivalent loop; x += 1 as x = x + 1; a temp variable returned right after its definition is inlined; operators are canonicalized (===/==/is, and/&&, None/null, const/let). Mutation benchmark (bench --mutations): each of rename, statement swap, temp variable, logging, dead code and loop-to-comprehension keeps >= 93% of pairs at score >= 0.7; all combined 40% (renames zero the name signal).

Real-repo validation (hand-judged)

Self-scans of three internal repositories (Python + TypeScript), 189 pairs judged by reading both units (TRUE_DUP / PARTIAL / BOILERPLATE / FALSE):

  • Raw scores are a weak signal: only ~23% of pairs above 0.55 were true duplicates, 40% incl. partial; below 0.65 almost none. Real duplicates were almost all exact copies with matching names.
  • Dominant false alarms: react-query/fetch wrappers differing by endpoint, per-entity CRUD endpoints, one-line repository delegators, DTO/model classes, per-file test fixtures, generated clients.
  • Filters added from this: field-only classes, constructors/dunders, generated files, tiny units (unless near-identical with the same name), test code (unless identical), minimum name similarity (0.3).
  • --profile copies (default) re-weights signals for copy-paste (name 0.36, statements 0.27, literals 0.18); leave-one-repo-out AUC improved on all three repos (0.71/0.69/0.95 -> 0.75/0.79/0.96). At threshold 0.6 the judged sample gives ~50% true duplicates and ~76% incl. partial while keeping ~70% of the true ones. These numbers are partly in-sample (weights and filters were chosen on the same pairs) — re-judge a fresh sample before trusting them. A structure-heavy --profile reimpl exists for renamed re-implementations (use with --min-name 0) but has no real-repo precision data yet.

Fresh held-out check (copies profile, threshold 0.6)

A third, untouched sample of 57 pairs (nothing was tuned on it): 30% true duplicates, 47% incl. partial (OneSales 0/24, MDMApp 11/24, CCMT2 6/9 true). The 50%/76% above was optimistic because it was measured on the pairs used to choose the weights. Test code is about half of the noise but also holds real copies (25% true either way), so it is kept by default; --skip-tests drops it.

Re-implementations of existing helpers (reimpl-eval)

42 real helper functions from the three repos were described neutrally (no names/code) and re-implemented from scratch by Haiku and Sonnet without seeing the repo. Fraction where the original is found (copies profile) and where the best other match also crosses the threshold:

threshold Sonnet found Haiku found both other-code alarm
0.3 67% 60% 63% 21%
0.4 62% 33% 48% 8%
0.5 45% 14% 30% 1%
0.6 24% 5% 14% 0%

So diff (checking new code) defaults to threshold 0.4 while scan (existing copies) defaults to 0.6. About half of independent re-implementations are caught; the rest are genuinely different code. The structure-heavy reimpl profile is not better on this test.

Head-to-head: LLM-only vs LLM + CLI (MDMApp, OneSales)

Four Sonnet agents (report only, ~80 tool calls, max 30 groups) hunted duplicates in two repos: two with plain read/grep, two with this CLI. All distinct groups (82) were then judged blind by independent agents that read the code.

groups true dup. precision (true / incl. partial) pooled recall (true) tool calls
MDMApp, LLM only 25 15 60% / 80% 60% ~27–33
MDMApp, with CLI 26 20 77% / 92% 80% 7
OneSales, LLM only 28 10 36% / 71% 67% ~33–42
OneSales, with CLI 21 10 48% / 86% 67% 9

Only 18 of 82 groups were found by both, so the approaches are complementary. The agents' reports are capped at 30 groups, so the raw tool is a better measure of recall: scan --threshold 0.6 (defaults) finds 38 of the 40 judged true duplicates (25/25 MDMApp, 13/15 OneSales) and 23 of the 25 that the LLM-only agents found independently. What it still misses: a differently-written picker function (same purpose, different code) and a formatFileSize variant with different constants. Fixed after this test: tiny same-name exact copies (min_tokens 20 -> 8, near-exact same-name rule) and multi-line module-level values such as export const queryClient = new QueryClient({...}).

Cheap model + CLI + skill (Haiku) vs Sonnet (MDMApp, OneSales)

Same task and judging as above (blind judges, pooled recall against all distinct confirmed true duplicates: 37 in MDMApp, 18 in OneSales). review + skills/duplicate-code-review/SKILL.md were built from what the first runs missed. "cost index" = tokens x relative price (Haiku assumed 1/3 of Sonnet per token; estimate).

repo approach groups precision true / incl. partial recall tokens cost index
MDMApp Sonnet, no CLI 25 60% / 80% 41% 86k 259
MDMApp Sonnet + CLI 29 79% / 93% 54% 59k 177
MDMApp Haiku + CLI + skill (r1 / r2) 31 / 40 90% / 90% · 70% / 82% 65% / 70% 86k / 86k 86
OneSales Sonnet, no CLI 28 36% / 71% 56% 131k 393
OneSales Sonnet + CLI 21 52% / 90% 61% 62k 187
OneSales Haiku + CLI + skill (r1 / r2) 11 / 30 82% / 91% · 33% / 67% 50% / 50% 88k / 74k 74–88

r1 = first skill (top-60 groups only); r2 = skill with --brief paging. Read: on MDMApp Haiku + CLI beats Sonnet alone on recall (70% vs 41%) and precision at about a third of the cost; on OneSales it is at par with Sonnet alone on recall (50% vs 56%, one group) with better-or-equal precision at about a fifth of the cost, but below Sonnet + CLI. Broadening (r2) buys recall at the price of precision on OneSales. Small samples; judges and finders are the same model family.

Round 3 (after LIKELY-NOISE tagging and the line-diff view; hard cap 30 groups): MDMApp 29 groups, precision 69% / 86%, recall 51%, 69k tokens; OneSales 12 groups, precision 75% / 83%, recall 42%, 89k tokens. Recall pooled against 41 (MDMApp) and 19 (OneSales) confirmed duplicates. Run-to-run variation of Haiku (59–63% / 51% on MDMApp, 42–47% on OneSales) is as large as the differences between skill/tool versions, so the tag and diff view are not shown to help measurably; averaged over three runs Haiku + CLI + skill lands at ~58% (MDMApp) and ~45% (OneSales) recall versus 37% / 53% for Sonnet alone and 49% / 63% for Sonnet + CLI, at roughly a fifth to a third of the cost of Sonnet alone. Note the 30-group cap is itself a ceiling: the confirmed pool holds 41 real duplicates in MDMApp.

Uncapped comparison (same 80-group cap, ~120-call budget for every arm) — final

Pooled confirmed true duplicates: 48 (MDMApp), 35 (OneSales); every group anyone reported was judged blind by reading the code.

repo arm groups precision true / incl. partial recall tokens cost index (Haiku = 1/3 price)
MDMApp Sonnet, no CLI 65 42% / 65% 65% 174k 522
MDMApp Sonnet + CLI + skill 69 62% / 86% 77% 72k 214
MDMApp Haiku + CLI + skill 59 73% / 86% 67% 102k 101
OneSales Sonnet, no CLI 64 30% / 56% 54% 153k 459
OneSales Sonnet + CLI + skill 80 32% / 71% 66% 90k 270
OneSales Haiku + CLI + skill 47 38% / 68% 40% 85k 84
both tool list only, IDENTICAL+NEAR-COPY, minus LIKELY-NOISE 173 / 278 (judged subset: 66% / 85% MDMApp, 38% / 64% OneSales) 88% / 83% 0 0
  • With the same budget Haiku + CLI matches Sonnet alone on MDMApp (67% vs 65% recall, better precision) at ~1/5 of the cost; on OneSales it is below (40% vs 54%) at ~1/5.5 of the cost. Sonnet + CLI is best on recall in both.
  • The unfiltered tool list already covers 83–88% of the pool: the agents' job is pruning, and their lower recall comes from what they choose to report (they judge from brief lines and read few members), not from what the tool misses. Precision of the raw list is unknown beyond the judged subset (which is biased towards groups an agent reported).
  • Caveats: judges and finders are the same model family; the pool only contains duplicates someone found; single runs per arm (Haiku varies by +-10 points between runs).

Release files for duplicatecode 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for duplicatecode 0.1.0
File Size Uploaded
duplicatecode-0.1.0.tar.gz 60.7 kB Details

Built distributions (wheels)

Table of built distributions (wheels) for duplicatecode 0.1.0
File
duplicatecode-0.1.0-py3-none-win_amd64.whl Python 3 none Windows x86-64 Details
duplicatecode-0.1.0-py3-none-manylinux_2_17_x86_64.manylinux2014_x86_64.whl Python 3 none Linux glibc 2.17+ x86-64 Details
duplicatecode-0.1.0-py3-none-manylinux_2_17_aarch64.manylinux2014_aarch64.whl Python 3 none Linux glibc 2.17+ ARM64 Details
duplicatecode-0.1.0-py3-none-macosx_11_0_arm64.whl Python 3 none macOS 11.0+ ARM64 Details
duplicatecode-0.1.0-py3-none-macosx_10_12_x86_64.whl Python 3 none macOS 10.12+ x86-64 Details

Total release size: 16.2 MB

Release files / duplicatecode-0.1.0.tar.gz

Download URL duplicatecode-0.1.0.tar.gz
Size 60.7 kB
Tags Source
SHA-256 checksum
How to use checksums
933572c67b6cb2fbedd60f9903ad17f5312ebf351705379144115750536b5cbe
BLAKE2b-256 checksum
How to use checksums
5221917074a14b4e4233123eedd7f082d597ba89f551f80da0d7fd41c73edbf9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / duplicatecode-0.1.0-py3-none-win_amd64.whl

Download URL duplicatecode-0.1.0-py3-none-win_amd64.whl
Size 2.9 MB
Tags Python 3 Windows x86-64
SHA-256 checksum
How to use checksums
875b8a24f2e7ad1cd49cc092253aaf85c0b0b6f738474ac3c77546aab269bd8d
BLAKE2b-256 checksum
How to use checksums
123e4b1143b4a6837246d32960155270214e61cedec55511ae4ce13f1c8cd93b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / duplicatecode-0.1.0-py3-none-manylinux_2_17_x86_64.manylinux2014_x86_64.whl

Download URL duplicatecode-0.1.0-py3-none-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Size 3.4 MB
Tags Linux glibc 2.17+ x86-64 Python 3
SHA-256 checksum
How to use checksums
0f4efb92b04d09dd431af40e8d719ded553da201619056265ab19b1eb88f4f42
BLAKE2b-256 checksum
How to use checksums
331a3f92388100059cb125c091b2adc8314811099fbdca6930643c3056f138d3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / duplicatecode-0.1.0-py3-none-manylinux_2_17_aarch64.manylinux2014_aarch64.whl

Download URL duplicatecode-0.1.0-py3-none-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Size 3.4 MB
Tags Linux glibc 2.17+ ARM64 Python 3
SHA-256 checksum
How to use checksums
6cbd781f8823422354e8b2e27eda89c20958bec4a3765a615dcc40ac6684529a
BLAKE2b-256 checksum
How to use checksums
0046f897baadc24419059ebe639368e20f4810d211f96fbb09f312ee515ba870
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / duplicatecode-0.1.0-py3-none-macosx_11_0_arm64.whl

Download URL duplicatecode-0.1.0-py3-none-macosx_11_0_arm64.whl
Size 3.1 MB
Tags Python 3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
15c02a0ee768a7f417553535c09825009adb44e7394587ce1018cd5dbf2a2d00
BLAKE2b-256 checksum
How to use checksums
af3b99bcfe7146d3935e3b98536533252df26cbe6feb166e95e7dd271640fe10
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / duplicatecode-0.1.0-py3-none-macosx_10_12_x86_64.whl

Download URL duplicatecode-0.1.0-py3-none-macosx_10_12_x86_64.whl
Size 3.2 MB
Tags Python 3 macOS 10.12+ x86-64
SHA-256 checksum
How to use checksums
3d4d987365f85da64a432ffa895d7a4b7be41c7446438f73583215e9da600768
BLAKE2b-256 checksum
How to use checksums
08767d48f510b8158d0dbb3e489fc614c00b53920372923748bcba02d9977973
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

6 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page