Constractor (derived from “Content Extractor’) allows one to use machine learning for web pages content extraction. Library provide following functionality:
Extendable features api.
Gui tools for simple train-set creation.
Simple training and testing process.
Simple usage of trained model.
Models dumping.
etc
Installation
NOTE: Project was developed and tested under Ubuntu 12.04. Other operation systems may require enhancements of library.
Ubuntu >=12.04 instructions:
Install apt dependecies: sudo apt-get install pip gcc g++ python-dev python-qt4
Install constractor: sudo pip install constractor
Usage
Following code will run gui train helper:
#!/usr/bin/env python
from constractor.train import GuiTrainer
if __name__ == '__main__':
GuiTrainer()
And following code will print html of predicted element in DOM:
#!/usr/bin/env python
from constractor.parser import Parser
if __name__ == '__main__':
predicted = Parser('https://pypi.python.org', model_file='model.txt').predicted
for element in predicted:
print unicode(element.toInnerXml()).encode('utf-8')
Contribution
Project is completely open for contribution. See more on bitbucket repo: https://bitbucket.org/dkuryakin/constractor
Metadata
Release files for constractor 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| constractor-0.1.0.tar.gz | 10.1 kB | Details |
Release files / constractor-0.1.0.tar.gz
| Download URL | constractor-0.1.0.tar.gz |
|---|---|
| Size | 10.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
fc5ed87c49f9f334ffbe19f773881c919f8843dde9ce1e49a5d2d41b6ccec345
|
|
BLAKE2b-256 checksum How to use checksums |
d6769ca8b9a73de04f0ec53cccff7e648df8186198e2f197bba04692e6052738
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |