regex4dummies·PyPI

A NLP library that simplifies pattern finding in strings

Project description

https://travis-ci.org/DarkmatterVale/regex4dummies.svg?branch=master

Simple pattern finding in strings and natural language processing. Checkout regex4dummies’ website at https://darkmattervale.github.io/regex4dummies/

Features

Automatic pattern detection ( semantic and literal )
Multiple parsers ( implementations of nltk, pattern, and nlpnet )
Keyword searching to find specific phrases within text
Topic analysis & Important information extraction
Tokenizer and sentence dependency identifier
Phrase extraction ( noun, verb, prepositional )
String comparison

Roadmap

Some features I plan to implement in the future:

Machine Learning. This will allow the parsers to learn multiple grammatical “styles” and be able to successfully parse a much wider selection of strings
Additional parsers

If you have feature requests, feel free to add an issue to the Github issue tracker. All contributions and requests are appreciated!

Usage

regex4dummies is very easy to use. Simply import the library, get some strings, and compare them!

from regex4dummies import regex4dummies
from regex4dummies import Toolkit

# Creating strings
strings = [ "This is the first test string.", "This is the second test string." ]

regex = regex4dummies()

# Identifying literal patterns in strings
print regex.compare_strings( parser='default', pattern_detection="literal", text=strings )

# Identifying semantic patterns in strings using the nltk parser
print regex.compare_strings( parser='nltk', pattern_detection="semantic", text=strings )

Above is regex4dummies in its simplest form. It allows for additional features as well, including:

# Display the version of regex4dummies you are using
print regex.__version__

# To use the other parsers, replace the above line of code with either of the following:
# print regex.compare_strings( parser='pattern', pattern_detection="semantic", text=strings )
# print regex.compare_strings( parser='nlpnet', pattern_detection="semantic", text=strings )

# To call all of the parsers, replace the above line of code with the following:
# print regex.compare_strings( parser='', pattern_detection="semantic", text=strings )

# To get the topics of the strings, call the get_topics function
print regex.get_topics( text=strings )

# Printing pattern information
pattern_information = regex.get_sentence_information()
  for objects in pattern_information:
      print "[ Pattern ]             : " + objects.pattern
      print "[ Subject ]             : " + objects.subject
      print "[ Verb ]                : " + objects.verb
      print "[ Object ]              : " + objects.object[0]
      print "[ Prep Phrases ]        : " + str( objects.prepositional_phrases )
      print "[ Reliability Score ]   : " + str( objects.reliability_score )
      print "[ Applicability Score ] : " + str( objects.applicability_score )
      print ""

A newly released set of features include a tokenizer and dependency finder function. To use them, simply give a string as the first parameter, and the name of the parser you would like to use as the second for both functions.

# Testing the toolkit functions
tool_tester = Toolkit()

# Testing the tokenizer functions
print tool_tester.tokenize( text="This is a test string.", parser="" )

# Testing the dependency functions
print tool_tester.find_dependencies( text="This is a test string.", parser="pattern" )

Other features included are demonstrated below.

# Testing the information extraction functions
regex.extract_important_information( text=[ "This is a test string." ] )

# Testing the ability to extract phrases
print "Noun Phrases: " + str( tool_tester.extract_noun_phrases( text="This is a test string." ) )
print "Verb Phrases(Pattern): " + str( tool_tester.extract_verb_phrases( text="This is a test string.", parser="pattern" ) )
print "Verb Phrases(Nlpnet): " + str( tool_tester.extract_verb_phrases( text="This is a test string.", parser="nlpnet" ) )
print "Prepositional Phrases: " + str( tool_tester.extract_prepositional_phrases( text="This is a test string in the house." ) )

print "String comparison: " + str( tool_tester.compare_strings( String1="This is a test string.", String2="This is a test string." ) )

Installation

To install this library, use pip.

$ pip install regex4dummies

In addition to the library, wget is a required command-line command to use the nlpnet parser. If you do not have wget or cannot get it, follow the below directions to still get the functionality of the nlpnet parser.

Instructions to install the required dependency for nlpnet:

Download the nlpnet_dependency file on the most recent release found in Github ( please not, when uncompressed, this file is over 350 MB large ).
Place this directory into the same directory that nltk-data is located ( if you don’t have that installed, just run the library and go through the GUI downloader )

That’s it! The nlpnet parser should now be able to be used.

Patch Notes

v1.4.6: Code refactoring, Download system rework

Brought code up to PEP8 standards
Redid download system. No more non-functional GUI; it uses automatic installation. In addition, the over size of the dependencies has decreased to ~700 MB from ~1.5 GB

Contributing

Contributors are welcome and much needed! regex4dummies is still under heavy development, and needs all of the help it can get. If you have any feature ideas, feel free to create an issue on the github repository ( https://github.com/darkmattervale/regex4dummies/issues ) or fork the repository and create your addition.

Any help you can give is much appreciated. The more help we get, the better regex4dummies will perform. Thanks for contributing!

License

Please see LICENSE.txt for information about the MIT license

Citations

nlpnet:

Fonseca, E. R. and Rosa, J.L.G. Mac-Morpho Revisited: Towards Robust Part-of-Speech Tagging. Proceedings of the 9th Brazilian Symposium in Information and Human Language Technology, 2013. p. 98-107 [PDF]

Project details

Release history Release notifications | RSS feed

This version

1.4.6

Dec 25, 2015

1.4.5

Nov 18, 2015

1.4.4

Nov 6, 2015

1.4.3

Sep 24, 2015

1.4.2

Aug 27, 2015

1.4.0

Aug 1, 2015

1.3.7

Jul 28, 2015

1.3.6

Jul 25, 2015

1.3.5

Jul 24, 2015

1.3.4

Jul 20, 2015

1.3.3

Jul 13, 2015

1.3.2

Jul 8, 2015

1.3.1

Jul 7, 2015

1.3.0

Jul 6, 2015

1.2.1

Jul 1, 2015

1.1.3

Jun 28, 2015

1.1.2

Jun 28, 2015

1.1.1

Jun 28, 2015

1.1.0

Jun 27, 2015

1.0.4

Jun 26, 2015

1.0.3

Jun 25, 2015

1.0.2

Jun 25, 2015

1.0.1

Jun 25, 2015

1.0.0

Jun 23, 2015

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

regex4dummies-1.4.6.tar.gz (24.5 kB view details)

Uploaded Dec 25, 2015 Source

Built Distribution

regex4dummies-1.4.6-py2.py3-none-any.whl (52.0 kB view details)

Uploaded Dec 25, 2015 Python 2Python 3

File details

Details for the file regex4dummies-1.4.6.tar.gz.

File metadata

Download URL: regex4dummies-1.4.6.tar.gz
Upload date: Dec 25, 2015
Size: 24.5 kB
Tags: Source
Uploaded using Trusted Publishing? No

File hashes

Hashes for regex4dummies-1.4.6.tar.gz
Algorithm	Hash digest
SHA256	`f473eb2c7f161952d78f97fc5a9ac50331d6f9afd7a7ba8633a822bdc791a598`
MD5	`1fd09dc2a06d0b039993ab9ec8680ceb`
BLAKE2b-256	`2a2a70b9cfb025208b0b01f278ab55103a3406d88a080be5be10042f91f99f96`

See more details on using hashes here.

File details

Details for the file regex4dummies-1.4.6-py2.py3-none-any.whl.

File metadata

Download URL: regex4dummies-1.4.6-py2.py3-none-any.whl
Upload date: Dec 25, 2015
Size: 52.0 kB
Tags: Python 2, Python 3
Uploaded using Trusted Publishing? No

File hashes

Hashes for regex4dummies-1.4.6-py2.py3-none-any.whl
Algorithm	Hash digest
SHA256	`034162cb76069672236abfb7dec0d7ee3bc2e592e9e2e0291e2c96b00e43c16c`
MD5	`373a47b4090c523be05e85962e970fd8`
BLAKE2b-256	`272def3fe8fefa3f2390170e1bce3aabae0532b5f0bd5516b483fb07e3ee2311`

See more details on using hashes here.

regex4dummies 1.4.6

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

Features

Roadmap

Usage

Installation

Patch Notes

Contributing

License

Citations

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes