Skip to main content

A short description for your project.

Project description

Library for detecting plagiarism in source code. Userful for online judges, teachers, developers and maybe lawyers.

Plagiarism uses the method described [here](http://…). The basic idea is to classify each submitted file according to different metrics and perform a series of k-means based clusterizations to determine which objects are most similar to each other. This approach has a N log N cost and scales fairly well to big samples.

The algorithm can be applied to natural text, source code and can even be adapted to run on arbitrary data structures (such as the parse tree of a computer program, ASM output, even binary executables). It requires some tuning for each application and accuracy may vary widely depending on application. You should expect better results grading Python and C source code. Performance on other programming languages or even in other domains may vary.

Project details

Release history Release notifications

This version
History Node


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Filename, size & hash SHA256 hash help File type Python version Upload date
plagiarism-0.1.0.tar.gz (21.1 kB) Copy SHA256 hash SHA256 Source None

Supported by

Elastic Elastic Search Pingdom Pingdom Monitoring Google Google BigQuery Sentry Sentry Error logging AWS AWS Cloud computing DataDog DataDog Monitoring Fastly Fastly CDN SignalFx SignalFx Supporter DigiCert DigiCert EV certificate StatusPage StatusPage Status page