Skip to main content

A short description for your project.

Project description

Library for detecting plagiarism in source code. Userful for online judges, teachers, developers and maybe lawyers.

Plagiarism uses the method described [here](http://…). The basic idea is to classify each submitted file according to different metrics and perform a series of k-means based clusterizations to determine which objects are most similar to each other. This approach has a N log N cost and scales fairly well to big samples.

The algorithm can be applied to natural text, source code and can even be adapted to run on arbitrary data structures (such as the parse tree of a computer program, ASM output, even binary executables). It requires some tuning for each application and accuracy may vary widely depending on application. You should expect better results grading Python and C source code. Performance on other programming languages or even in other domains may vary.

Project details

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Files for plagiarism, version 0.1.0
Filename, size File type Python version Upload date Hashes
Filename, size plagiarism-0.1.0.tar.gz (21.1 kB) File type Source Python version None Upload date Hashes View

Supported by

Pingdom Pingdom Monitoring Google Google Object Storage and Download Analytics Sentry Sentry Error logging AWS AWS Cloud computing DataDog DataDog Monitoring Fastly Fastly CDN DigiCert DigiCert EV certificate StatusPage StatusPage Status page