C-Rank

Extracts Keyphrases from Documents

These details have not been verified by PyPI

Project links

Project description

# C-Rank

C-Rank is an unsupervised keyphrase extraction algorithm that uses Concept Linking in order to improve its results.

It does not need external data to be inputted by the user other than the document to have its keyphrases extracted.

It is necessary to create an account to Babelfy (http://babelfy.org/login) as C-Rank uses its services. Then, your Babelfy key must be inserted in orde to C-Rank work properly.

## Installation
The following packages must be installed to use C-Rank:

networkx (https://networkx.github.io/):
```
pip install networkx
```

nltk:(https://www.nltk.org/index.html)
```
sudo pip install -U nltk
```

pybabelfy: (https://github.com/aghie/pybabelfy)
```
sudo pip install pybabelfy
```

## Getting started
```
from crank import CRank as cr

crank = cr.CRank(BABELFY_KEY, LIST_OF_INPUT_DOCUMENTS, OUTPUT_DIRECTORY)
#Exemple
#crank = cr.CRank("3ejklasd-a456-41ae-647f-0a1234546dd3", ['./document1.txt', './document2.txt'], './')
crank.keyphrasesExtraction()

printKeyphrases()
```
## Functionalities
```
# all printing options
printKeyphrases(self, nKeyphrases = 10, documentIndex=-1, showRanking = True, stem = False)

# save options to persist keyphrases in a single file (as in SemEval)
saveKeyphrasesSingleFile(self, fileName, nKeyphrases = 10, documentIndex=-1, showRanking = True, stem = False)

# save options to persist keyphrases in diferent files
saveKeyphrasesDiferentFiles(self, nKeyphrases = 10, documentIndex=-1, showRanking = True, stem = False)

# variables used in above functionalities
##nKeyphrases = number of kyphrases to print | nKeyphrases = 0 for all keyphrases
##documentIndex = index of document to print | documentIndex = -1 for all documents
##showRanking = show or not weight of keyphrases
##stem = stem or not keyphrases
##fileName = name of the file
```
### Intermediate results and available variables
```
self.key = BabelfyKey
self.inputFiles = inputFiles
self.outputDirectory = outputDirectory
self.lang = language
self.distance = dist
self.graphName = []
self.splitted_text = []
self.dictionary = []
self.dictionaryCode = []
self.weight = []
self.paragraphs_annotations = []
self.paragraphs_text = []
self.paragraphs_code = []
self.graphs = []
self.graphs2 = []
self.keyPhrases = []
```
## Citation
Available soon

Project details

These details have not been verified by PyPI

Project links

Release history Release notifications | RSS feed

Apr 18, 2019

This version

Apr 18, 2019

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

C-Rank-2.tar.gz (9.8 kB view details)

Uploaded Apr 18, 2019 Source

File details

Details for the file C-Rank-2.tar.gz.

File metadata

Download URL: C-Rank-2.tar.gz
Upload date: Apr 18, 2019
Size: 9.8 kB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/1.13.0 pkginfo/1.5.0.1 requests/2.21.0 setuptools/41.0.0 requests-toolbelt/0.9.1 tqdm/4.31.1 CPython/3.6.7

File hashes

Hashes for C-Rank-2.tar.gz
Algorithm	Hash digest
SHA256	`02fcfe8f00fbf2f4c8926d5955ddbae5d5034ad71a82ac9c533ab7d3d09ec0a9`
MD5	`004b1abbff4f15fa8a80eecc41ded727`
BLAKE2b-256	`4c3bb92f0dd77d51133ad78fa6e29e2f0fa9b30e316f52ec51372c979456d3d1`

See more details on using hashes here.

C-Rank 2

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

File details

File metadata

File hashes