Skip to main content

GitHub Stars PyPi Version PyPi Downloads

🔑 Keyword spaCy

keyword spacy

Keyword spaCy is a spaCy pipeline component for extracting keywords from text using cosine similarity. The basis for this comes from KeyBERT: A Minimal Method for Keyphrase Extraction using BERT, a transformer-based approach to keyword extraction. The methods employed by Keyword spaCy follow this methodology closely. It allows users to specify the range of n-grams to consider and can operate in a strict mode, which limits results to the specified n-gram range.

Installation

Before using Keyword spaCy, make sure you have spaCy installed:

pip install keyword-spacy

Then, download the en_core_web_md model:

python -m spacy download en_core_web_md

Usage

To use the Keyword Extractor, first, create a spaCy nlp object:

import spacy
nlp = spacy.load("en_core_web_md")

Then, add the KeywordExtractor to the pipeline:

nlp.add_pipe("keyword_extractor", last=True, config={"top_n": 10, "min_ngram": 3, "max_ngram": 3, "strict": True})

Now you can process text and extract keywords:

text = "Natural language processing is a fascinating domain of artificial intelligence. It allows computers to understand and generate human language."
doc = nlp(text)
print("Top Keywords:", doc._.keywords)

Output:

Top Keywords: ['generate human language', 'Natural language processing']

Each token that is not a punctuation also receives a special attribute ._.keyword_value, this is the value of a given word's similarity to the doc.vector. This may be helpful for other downstream tasks.

Configuration

The KeywordExtractor can be configured using the following parameters:

  • top_n: The number of top keywords to extract.
  • min_ngram: The minimum size for n-grams.
  • max_ngram: The maximum size for n-grams.
  • strict: If set to True, only n-grams within the min_ngram to max_ngram range are considered. If False, individual tokens and the specified range of n-grams are considered.

Methodology

The methodology employed by Keyword spaCy is inspired by KeyBERT. It utilizes cosine similarity between tokens (and n-grams) and the entire document to determine the relevance of terms. The most similar terms are then considered as keywords.

References

Metadata

Release files for keyword-spacy 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for keyword-spacy 0.1.2
File Size Uploaded
keyword_spacy-0.1.2.tar.gz 3.6 kB Details

Release files / keyword_spacy-0.1.2.tar.gz

Download URL keyword_spacy-0.1.2.tar.gz
Size 3.6 kB
Tags Source
SHA-256 checksum
How to use checksums
f1235f8e5fbff1429f70cd07953e3993d7e71df7925b45fa46d6915a14f16bbf
BLAKE2b-256 checksum
How to use checksums
b56a6ac144946514b8564d9854a8c0e1743b0d6c01f16004d222f1ef46843954
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/4.0.2 CPython/3.7.15

Release history Release notifications | RSS feed

This release

0.1.2 This release

1 release file

0.1.1

2 release files

0.1

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page