No project description provided
Project description
NLP Primitives
nlp_primitives is a Python library with Natural Language Processing Primitives, intended for use with Featuretools.
nlp_primitives allows you to make use of text data in your machine learning pipeline in the same pipeline as the rest of your data.
Install
There are two options for installing nlp_primitives. Both of the options will also install Featuretools if it is not already installed.
The first option is to install a version of nlp_primitives that does not include Tensorflow. With this option, primitives that depend on Tensorflow cannot be used. Currently, the only primitive that can not be used with this install option is UniversalSentenceEncoder
.
nlp_primitives without Tensorflow can be installed with pip:
pip install nlp_primitives
or from the conda-forge channel on conda:
conda install -c conda-forge nlp-primitives
The second option is to install the complete version of nlp_primitives, which will also install Tensorflow and allow use of all primitives.
To install the complete version of nlp_primitives with pip:
pip install "nlp_primitives[complete]"
or from the conda-forge channel on conda:
conda install -c conda-forge nlp-primitives-complete
Demos
Calculating Features
With nlp_primitives primtives in featuretools
, this is how to calculate the same feature.
from featuretools.nlp_primitives import PolarityScore
data = ["hello, this is a new featuretools library",
"this will add new natural language primitives",
"we hope you like it!"]
pol = PolarityScore()
pol(data)
0 0.365
1 0.385
2 1.000
dtype: float64
Combining Primitives
In featuretools
, this is how to combine nlp_primitives primitives with built-in or other installed primitives.
import featuretools as ft
from featuretools.nlp_primitives import TitleWordCount
from featuretools.primitives import Mean
entityset = ft.demo.load_retail()
feature_matrix, features = ft.dfs(entityset=entityset, target_entity='products', agg_primitives=[Mean], trans_primitives=[TitleWordCount])
feature_matrix.head(5)
MEAN(order_products.quantity) MEAN(order_products.unit_price) MEAN(order_products.total) TITLE_WORD_COUNT(description)
product_id
10002 16.795918 1.402500 23.556276 3.0
10080 13.857143 0.679643 8.989357 3.0
10120 6.620690 0.346500 2.294069 2.0
10123C 1.666667 1.072500 1.787500 3.0
10124A 3.2000 0.6930 2.2176 5.0
Development
To install from source, clone this repo and run
make installdeps-test
This will install all pip dependencies.
Feature Labs
NLP Primitives is an open source project created by Feature Labs. To see the other open source projects we're working on visit Feature Labs Open Source. If building impactful data science pipelines is important to you or your business, please get in touch.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
File details
Details for the file nlp_primitives-1.1.0.tar.gz
.
File metadata
- Download URL: nlp_primitives-1.1.0.tar.gz
- Upload date:
- Size: 17.8 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/3.2.0 pkginfo/1.6.1 requests/2.24.0 setuptools/47.1.0 requests-toolbelt/0.9.1 tqdm/4.51.0 CPython/3.7.9
File hashes
Algorithm | Hash digest | |
---|---|---|
SHA256 | e78e63fdec2750c3aeca2108307e8fb29d6729007d11c0b224640ea2ea7450cc |
|
MD5 | 28e51c53d7bb4228ec017bc65546a3b2 |
|
BLAKE2b-256 | 059ce78d8bc47e8addc52bcf9eb64378936d814a7744e2d3344f4e8c27971290 |
File details
Details for the file nlp_primitives-1.1.0-py3-none-any.whl
.
File metadata
- Download URL: nlp_primitives-1.1.0-py3-none-any.whl
- Upload date:
- Size: 18.0 MB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/3.2.0 pkginfo/1.6.1 requests/2.24.0 setuptools/47.1.0 requests-toolbelt/0.9.1 tqdm/4.51.0 CPython/3.7.9
File hashes
Algorithm | Hash digest | |
---|---|---|
SHA256 | e3c2eb6d0fb70115f17fae4b101fe5927734914380c0a6c225a28d4afe0e0cb5 |
|
MD5 | d105976cf62d4a98ebf98f146dec96ae |
|
BLAKE2b-256 | a0b8c5465a6ed6d5fce76703e62d60fde588cee1af65a8a3d7340b26e1a1c290 |