Skip to main content

extraction toolkit

Project description

ETK: Information Extraction Toolkit

ETK is a Python library for high precision information extraction from many document formats. It proivdes a flexible framework of composable extractors that enables you to combine a host of predefined extractors provided in ETK with custom extractors that you may need to develop for your application. It supports extraction from HTML pages, text documents, CSV and Excel files and JSON documents. ETK is open-source software, released under the MIT license.

MIT License travis ci


Read the documentation here


  • Extraction from HTML, text, CSV, Excel, JSON
  • High-precision predefined extractors for common entities (dates, phones, email, cities, ...)
  • Extraction of microdata, and RDFa markup
  • Integration with spaCy for text processing
  • Automatic identification and extraction of HTML tables containing data
  • Automatic identification and extraction of time series
  • Semi-automatic generation of Web wrappers
  • Scalable execution and management of extraction pipelines
  • Automatic provenance recording



Operating system:macOS / OS X, Linux, Windows
Python version:Python 3.6+

Project details

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Files for etk, version 2.2.6
Filename, size File type Python version Upload date Hashes
Filename, size etk-2.2.6-py3-none-any.whl (203.1 kB) File type Wheel Python version py3 Upload date Hashes View
Filename, size etk-2.2.6.tar.gz (150.9 kB) File type Source Python version None Upload date Hashes View

Supported by

Pingdom Pingdom Monitoring Google Google Object Storage and Download Analytics Sentry Sentry Error logging AWS AWS Cloud computing DataDog DataDog Monitoring Fastly Fastly CDN DigiCert DigiCert EV certificate StatusPage StatusPage Status page