Slopo
A CLI tool for detecting non-exact code duplication using embedding models.
It focuses on the similar code that is hardest to detect and most harmful: snippets written similarly, sitting far apart in the codebase, often spread across different modules or separated within a large file. Exact copy-paste is easy to spot by other tools, and duplicates that are close together are easy to spot by humans or AI.
For high-level description of the problem and example LLM prompts, see slopo.dev.
For details, see Embedding models benchmark for code duplication detection. The author is also the developer of this tool, so both parts are compatible. The sample configuration in this documentation is the one that gives the best results based on this research.
Supported languages
Python, TypeScript, JavaScript, Java, Kotlin, C#, Go, Rust, PHP, Elixir
How it works
It takes a different approach than typical duplication detection. For every code unit, it calculates an embedding, then looks for pairs whose embeddings are close. Similar code is not necessarily a duplicate, so each pair is a potential duplicate to confirm.
The result is clusters of similar code units, ranked by similarity and by distance in the codebase. These clusters are meant as input for your AI coding agent, which can check whether a cluster is a real duplicate. Reviewed clusters can be marked as ignored or passed on for refactoring.
Example report generated from Slopo code (src directory, git tag v0.2.0).
Accessing embedding model
According to the benchmark, two providers are recommended:
- Jina AI - code-focused models available via API and for local use
- Voyage AI - general-purpose model only, not code-focused ones which are not suitable here
Jina AI offers API keys with free tokens for non-commercial use, without registration. Paid options are available for commercial use. Simply open the main page, and the token will be generated. You should see "You have 10,000,000 tokens left in the API key below." However, free tokens are not always granted - this is probably because of abuse prevention mechanisms based on IP address or other measures (it's a guess, not official information).
Local model
One tested solution is to use Ollama with jina-embeddings-v2-base-code. This is a small model that also runs on a CPU, but it may be (significantly) slower than the API, depending on the hardware you use. Pull the model from here.
Other models
Any model provider supported by LiteLLM can be configured. Additionally, any OpenAI-compatible server is supported.
Quick start
Once you figure out the model, the rest is quick.
Installation
uv tool install slopo
Update to the latest version (there are no automatic updates)
uv tool upgrade slopo
This command uses uv (installing uv), a Python package manager, to install Slopo from PyPI in an isolated virtual environment. No need to get Python separately.
Setup
Run slopo init to create a config file template containing further instructions. Only the directory with code for analysis and embedding model configuration is required.
Model configuration
Jina AI
embedding_model: jina_ai/jina-code-embeddings-0.5b
embedding_dimensions: 256
embedding_params:
task: code2code.query
If you are embedding a large project with the free API key and hitting rate limits, add the embedding_request_delay: 6 option to slow down.
Voyage AI
embedding_model: voyage/voyage-4-large
embedding_dimensions: 256
Local model
embedding_model: ollama/unclemusclez/jina-embeddings-v2-base-code
embedding_dimensions: 768
Analysis
Run slopo show-config to validate your config and show all configurable parameters, most are optional with sensible defaults.
Now you are ready to index code, calculate embeddings and generate a report:
slopo index
slopo embed
slopo analyze
Real workflow
This section demonstrates how Slopo can be used in a real development workflow.
It utilizes incremental re-indexing (update index with changed files only) and slopo.ignore.txt to discard already reviewed clusters.
- Create your first analysis and check results. You will notice
index.mdcontaining a list of all clusters and cluster details per file. - You may want to exclude some directories or file patterns, usually excluding tests is a good idea. You can also tune thresholds if the result is too big or too small.
- Once satisfied with analysis results, ask your AI coding agent to filter out clusters that are not real duplicates. This is a common case because not every similar code is a duplication to act on. Ask the AI agent to add discarded cluster hashes to
slopo.ignore.txt. - Re-run the analysis to generate a report without reviewed clusters. This is a basis for refactoring, which can be done by an AI agent.
ignorefile can be committed to your Git repository and reused cross-team. New and modified clusters will reappear in the report. A configuration file without an API key can also be committed. Don't commitslopo.db, this is your local data.
Configuration
Run slopo --help and slopo show-config to explore it by yourself anytime.
Most configuration is done with a configuration file with two exceptions:
- The location of the configuration file can be overridden with the
--configoption. - The API key can be set with the
SLOPO_EMBEDDING_API_KEYenvironment variable, also picked up from a.envfile in the current directory.
Be aware that some parameters can't be changed after first indexing. You need to remove slopo.db and index/embed from the beginning: source_dir, embedding_model, embedding_dimensions, body_node_count_threshold.
All configurable parameters
source_dir: Source directory with code to index, absolute or relative path.source_dir_exclude: .gitignore-style patterns to exclude from indexing.db_file: SQLite database file with tool data.report_dir: Output directory for analysis report.ignore_file: Text file with ignored clusters.embedding_model: Embedding model name in LiteLLM format.embedding_dimensions: Embedding dimensions compatible with the used model. This value is also used to verify received embeddings dimensions.embedding_api_key: API key for embedding provider, alternatively configured with an environment variable. Optional, no need to set for local models.embedding_params: Additional properties passed to the LiteLLM and embedding API. Anything supported by the API provider is valid. Example:
embedding_params:
api_base: http://example.com:123
task: code2code.query
embedding_batch_sizeandembedding_batch_chars: Requests to the embedding API are batched for performance. Defaults are fine for most cases.embedding_request_delay: Delay in seconds after every batched request, by default no delay. Increase if you reach rate limits.similarity_threshold: Controls minimal cosine similarity between embeddings.rerank_threshold: Controls minimal similarity after applying a boost reflecting distance in the codebase.body_node_count_threshold: Number of AST nodes inside the body (excluding signature and annotations). This value reflects the minimum code complexity of the included code unit, more precise than text length. Increase if you notice unwanted, too-small code units in the report.
Details
Ranking thresholds
Similar code units are filtered in two passes, each with its own configurable threshold. The pipeline is as follows:
similarity_thresholdfilters out code unit pairs whose embeddings are not similar enough. The calculated value is cosine similarity ranging from-1to1where1means the same.- Similar pairs are grouped in clusters.
- Units in clusters are reranked after applying a boost. Boost is calculated based on the number of directory hops required to reach the other file in the pair (max. 15%). If they are in the same file, the boost is calculated based on distance in number of lines (max. 10%).
rerank_thresholdfilters out clusters whose highest-scoring pair is not high enough.
Exact-copy duplicates
The main goal of this tool is to detect non-exact code duplication, but exact copies (identical code at multiple paths) are reported too, just handled a little differently from merely similar code:
- The report shows the code once, listing every path where it appears, instead of repeating identical snippets.
- The
analyzecommand reports the "similarity ratio" (the share of code units flagged as similar) in two variants: including and excluding exact copies.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file slopo-0.5.0.tar.gz.
File metadata
- Download URL: slopo-0.5.0.tar.gz
- Upload date:
- Size: 39.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/6.1.0 CPython/3.13.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c92ac84624f7d5b87fb3c8d3b457ed1a9bdc511708785e353cc37202a4f35edb
|
|
| MD5 |
1ff67df0bf7debb025a16b2e7275aebd
|
|
| BLAKE2b-256 |
3c12b72ac338c34ddc30c8bcc756600ee40595b09e80d55edb964ef55c789278
|
Provenance
The following attestation bundles were made for slopo-0.5.0.tar.gz:
Publisher:
publish.yml on rafal-qa/slopo
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
slopo-0.5.0.tar.gz -
Subject digest:
c92ac84624f7d5b87fb3c8d3b457ed1a9bdc511708785e353cc37202a4f35edb - Sigstore transparency entry: 2532109411
- Sigstore integration time:
-
Permalink:
rafal-qa/slopo@08e01d07e4b82cb20144d51322cd458f3247f3ac -
Branch / Tag:
refs/tags/v0.5.0 - Owner: https://github.com/rafal-qa
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@08e01d07e4b82cb20144d51322cd458f3247f3ac -
Trigger Event:
release
-
Statement type:
File details
Details for the file slopo-0.5.0-py3-none-any.whl.
File metadata
- Download URL: slopo-0.5.0-py3-none-any.whl
- Upload date:
- Size: 56.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/6.1.0 CPython/3.13.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e8082364c06fb63961949d3909636a858571b5454aa35b1ef98ef7b580ea9571
|
|
| MD5 |
1aa9bb2e1a7412c57abcf6740f0e0524
|
|
| BLAKE2b-256 |
1953d6499f05870c16c88aff3dfcfd5c780567542c09bd5d1e69d09d07811dc5
|
Provenance
The following attestation bundles were made for slopo-0.5.0-py3-none-any.whl:
Publisher:
publish.yml on rafal-qa/slopo
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
slopo-0.5.0-py3-none-any.whl -
Subject digest:
e8082364c06fb63961949d3909636a858571b5454aa35b1ef98ef7b580ea9571 - Sigstore transparency entry: 2532109836
- Sigstore integration time:
-
Permalink:
rafal-qa/slopo@08e01d07e4b82cb20144d51322cd458f3247f3ac -
Branch / Tag:
refs/tags/v0.5.0 - Owner: https://github.com/rafal-qa
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@08e01d07e4b82cb20144d51322cd458f3247f3ac -
Trigger Event:
release
-
Statement type: