xml-cleaner·PyPI

Word and sentence tokenization.

These details have not been verified by PyPI

Project links

Intended Audience
- Science/Research
Operating System
- OS Independent
Programming Language
- Python :: 2.7
- Python :: 3.3
Topic
- Text Processing :: Linguistic

Project description

XML cleaner

Word and sentence tokenization in Python.

[![PyPI version](https://badge.fury.io/py/xml-cleaner.svg)](https://badge.fury.io/py/xml-cleaner) [![Build Status](https://travis-ci.org/JonathanRaiman/xml_cleaner.svg?branch=master)](https://travis-ci.org/JonathanRaiman/xml_cleaner) ![Jonathan Raiman, author](https://img.shields.io/badge/Author-Jonathan%20Raiman%20-blue.svg)

[![License](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE.md)

Usage

Use this package to split up strings according to sentence and word boundaries. For instance, to simply break up strings into tokens:

` tokenize("Joey was a great sailor.") #=> ["Joey ", "was ", "a ", "great ", "sailor ", "."] `

To also detect sentence boundaries:

` sent_tokenize("Cat sat mat. Cat's named Cool.", keep_whitespace=True) #=> [["Cat ", "sat ", "mat", ". "], ["Cat ", "'s ", "named ", "Cool", "."]] `

sent_tokenize can keep the whitespace as-is with the flags keep_whitespace=True and normalize_ascii=False.

Installation

` pip3 install xml_cleaner `

Testing

Run nose2.

Project details

These details have not been verified by PyPI

Project links

Intended Audience
- Science/Research
Operating System
- OS Independent
Programming Language
- Python :: 2.7
- Python :: 3.3
Topic
- Text Processing :: Linguistic

Release history Release notifications | RSS feed

This version

2.0.4

Dec 29, 2016

2.0.3

Dec 16, 2016

2.0.2

Dec 10, 2016

2.0.1

Dec 7, 2016

2.0.0

Dec 5, 2016

1.0.21

Sep 26, 2015

1.0.20

Sep 22, 2015

1.0.19

Sep 22, 2015

1.0.18

Jun 26, 2015

1.0.17

Mar 9, 2015

1.0.16

Mar 9, 2015

1.0.15

Feb 20, 2015

1.0.14

Feb 20, 2015

1.0.13

Feb 20, 2015

1.0.12

Oct 31, 2014

1.0.11

Oct 31, 2014

1.0.10

Oct 31, 2014

1.0.9

Oct 27, 2014

1.0.8

Oct 27, 2014

1.0.7

Oct 27, 2014

1.0.6

Oct 27, 2014

1.0.5

Oct 27, 2014

1.0.4

Oct 27, 2014

1.0.3

Oct 27, 2014

1.0.2

Oct 27, 2014

1.0.1

Oct 27, 2014

1.0.0

Oct 27, 2014

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

xml-cleaner-2.0.4.tar.gz (10.8 kB view details)

Uploaded Dec 29, 2016 Source

File details

Details for the file xml-cleaner-2.0.4.tar.gz.

File metadata

Download URL: xml-cleaner-2.0.4.tar.gz
Upload date: Dec 29, 2016
Size: 10.8 kB
Tags: Source
Uploaded using Trusted Publishing? No

File hashes

Hashes for xml-cleaner-2.0.4.tar.gz
Algorithm	Hash digest
SHA256	`7b79fbab3d9b9a57dbd5b3b6d7457f576722e2e20f42653b4939de6065fbb90d`
MD5	`b23b49660513ae444b91f74cd4a5ae39`
BLAKE2b-256	`c16b859b2f7730a32e049db35b3b6e9640c6b55c1cd4195c19fe5e8a063b4066`

See more details on using hashes here.

xml-cleaner 2.0.4

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

XML cleaner

Usage

Installation

Testing

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

File details

File metadata

File hashes