Automate information extraction for multimodal LLMs.

These details have not been verified by PyPI

Project links

Homepage

Project description

The Pipe

python-gh-action

Feed PDFs, word docs, slides, web pages and more into Vision-LLMs with one line of code ⚡

The Pipe is a multimodal-first tool for feeding files and web pages into vision-language models such as GPT-4V. It is best for LLM and RAG applications that require a deep understanding of tricky data sources. The Pipe is available as a hosted API at thepi.pe, or it can be set up locally.

Demo

Features 🌟

Extracts text and visuals from files or web pages 📚
Outputs chunks optimized for multimodal LLMs 🖼️
Interpret complex PDFs, web pages, slides, CSVs, and more 🧠
Auto-compress prompts exceeding your chosen token limit 📦
Works even with missing file extensions, in-memory data streams 💾
Works with codebases, git repos, and custom integrations 🌐
Multi-threaded ⚡️

Getting Started 🚀

The Pipe handles a wide array of complex filetypes, and thus has many dependencies that must be installed separately. It also requires a strong machine for good response times. For this reason, we host it as an API that works out-of-the-box.

First, install The Pipe.

pip install thepipe_api

The Pipe is available as a hosted API, or it can be set up locally. An API key is recommended for out-of-the-box functionality (alternatively, see the local installation section). Ensure the THEPIPE_API_KEY environment variable is set. Don't have a key yet? Get one here.

Now you can extract comprehensive text and visuals from any file:

from thepipe_api import thepipe
messages = thepipe.extract("example.pdf")

Or any website:

messages = thepipe.extract("https://example.com")

Then feed it into GPT-4-Vision:

response = client.chat.completions.create(
    model="gpt-4-vision-preview",
    messages = messages,
)

Just call OpenAI

You can also use The Pipe from the command line. Here's how to recursively extract from a directory, matching only a specific file type:

thepipe path/to/folder --match *jsx

Supported File Types 📚

Source Type	Input types	Token Compression 🗜️	Image Extraction 👁️	Notes 📌
Directory	Any `/path/to/directory`	✔️	✔️	Extracts from all files in directory, supports match and ignore patterns
Code	`.py`, `.tsx`, `.js`, `.html`, `.css`, `.cpp`, etc	✔️ (varies)	❌	Combines all code files. `.c`, `.cpp`, `.py` are compressible with ctags, others are not
Plaintext	`.txt`, `.md`, `.rtf`, etc	✔️	❌	Regular text files
PDF	`.pdf`	✔️	✔️	Extracts text and images of each page; can use AI for extraction of table data and images within pages
Image	`.jpg`, `.jpeg`, `.png`	❌	✔️	Extracts images, uses OCR if text_only
Data Table	`.csv`, `.xls`, `.xlsx`	✔️	❌	Extracts data from spreadsheets; converts to text representation. For very large datasets, will only extract column names and types
Jupyter Notebook	`.ipynb`	❌	✔️	Extracts code, markdown, and images from Jupyter notebooks
Microsoft Word Document	`.docx`	✔️	✔️	Extracts text and images from Word documents
Microsoft PowerPoint Presentation	`.pptx`	✔️	✔️	Extracts text and images from PowerPoint presentations
Website	URLs (inputs containing `http`, `https`, `ftp`)	✔️	✔️	Extracts text from web page along with image (or images if scrollable); text-only extraction available
GitHub Repository	GitHub repo URLs	✔️	✔️	Extracts from GitHub repositories; supports branch specification
ZIP File	`.zip`	✔️	✔️	Extracts contents of ZIP files; supports nested directory extraction

How it works 🛠️

The input source is either a file path, a URL, or a directory. The pipe will extract information from the source and process it for downstream use with language models, vision transformers, or vision-language models. The output from the pipe is a sensible list of multimodal messages representing chunks of the extracted information, carefully crafted to fit within context windows for any models from gemma-7b to GPT-4. The messages returned should look like this:

[
  {
    "type": "text",
    "content": "Extracted text here..."
  },
  {
    "type": "image_url",
    "image_url": {
      "url": "data:image/jpeg;base64,..."}
  },
]

The text and images from these messages may also be prepared for a vector database with thepipe.core.create_chunks_from_messages or for downstream use with RAG frameworks. LiteLLM can be used to easily integrate The Pipe with any LLM provider.

It uses a variety of heuristics for optimal performance with vision-language models, including AI filetype detection with filetype detection, opt-in AI table, equation, and figure extraction, efficient token compression, automatic image encoding, reranking for lost-in-the-middle effects, and more, all pre-built to work out-of-the-box.

Local Installation 🛠️

The Pipe handles a wide array of complex filetypes, and thus requires installation of many different packages to function. It also requires a very capable machine for good response times. For this reason, we host it as an API that works out-of-the-box. To use The Pipe locally for free instead, you will need playwright, ctags, pytesseract, and the local python requirements, which differ from the more lightweight API requirements:

git clone https://github.com/emcf/thepipe
pip install -r requirements_local.txt

Tip for windows users: Install the python-libmagic binaries with pip install python-magic-bin. Ensure the tesseract-ocr binaries and the ctags binaries are in your PATH.

Now you can use The Pipe with Python:

from thepipe_api import thepipe
chunks = thepipe.extract("example.pdf", local=True)

or from the command line:

thepipe path/to/folder --local

Arguments are:

source (required): can be a file path, a URL, or a directory path.
local (optional): Use the local version of The Pipe instead of the hosted API.
match (optional): Regex pattern to match files in the directory.
ignore (optional): Regex pattern to ignore files in the directory.
limit (optional): The token limit for the output prompt, defaults to 100K. Prompts exceeding the limit will be compressed. This may not work as expected with the API, as it is in active development.
ai_extraction (optional): Extract tables, figures, and math from PDFs using our extractor. Incurs extra costs.
text_only (optional): Do not extract images from documents or websites. Additionally, image files will be represented with OCR instead of as images.

Sponsors

Thank you to Cal.com for sponsoring this project. Contact emmett@thepi.pe for sponsorship information.

Project details

These details have not been verified by PyPI

Project links

Homepage

Release history Release notifications | RSS feed

1.3.9

Sep 30, 2024

1.3.8

Sep 13, 2024

1.3.7

Sep 9, 2024

1.3.5

Sep 9, 2024

1.3.3

Sep 6, 2024

1.2.9

Sep 5, 2024

1.2.8

Sep 5, 2024

1.2.6

Sep 4, 2024

1.2.5

Sep 4, 2024

1.2.4

Sep 2, 2024

1.2.3

Sep 1, 2024

1.2.2

Sep 1, 2024

1.2.1

Jul 26, 2024

1.2.0

Jul 26, 2024

1.1.9

Jul 24, 2024

1.1.8

Jul 21, 2024

1.1.6

Jul 21, 2024

1.1.5

Jul 21, 2024

1.1.4

Jul 21, 2024

1.1.3

Jul 21, 2024

1.1.1

Jul 21, 2024

1.1.0

Jul 21, 2024

0.3.9

Jun 19, 2024

0.3.6

May 17, 2024

0.3.5

May 17, 2024

0.3.4

Apr 30, 2024

0.3.3

Apr 28, 2024

0.3.2

Apr 27, 2024

0.3.1

Apr 27, 2024

This version

0.3.0

Apr 20, 2024

0.2.9

Apr 20, 2024

0.2.8

Apr 20, 2024

0.2.7

Apr 19, 2024

0.2.6

Apr 17, 2024

0.2.5

Apr 17, 2024

0.2.4

Apr 17, 2024

0.2.3

Apr 16, 2024

0.2.2

Apr 16, 2024

0.2.0

Apr 16, 2024

0.1.9

Apr 15, 2024

0.1.8

Apr 15, 2024

0.1.7

Apr 15, 2024

0.1.6

Apr 14, 2024

0.1.5

Apr 14, 2024

0.1.4

Apr 13, 2024

0.1.3

Apr 13, 2024

0.1.2

Apr 13, 2024

0.1.0

Apr 13, 2024

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

thepipe_api-0.3.0.tar.gz (19.9 kB view details)

Uploaded Apr 20, 2024 Source

Built Distribution

thepipe_api-0.3.0-py3-none-any.whl (18.4 kB view details)

Uploaded Apr 20, 2024 Python 3

File details

Details for the file thepipe_api-0.3.0.tar.gz.

File metadata

Download URL: thepipe_api-0.3.0.tar.gz
Upload date: Apr 20, 2024
Size: 19.9 kB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: twine/3.8.0 pkginfo/1.9.6 readme-renderer/37.3 requests/2.31.0 requests-toolbelt/0.10.1 urllib3/1.26.18 tqdm/4.66.2 importlib-metadata/6.11.0 keyring/24.3.0 rfc3986/1.5.0 colorama/0.4.6 CPython/3.10.8

File hashes

Hashes for thepipe_api-0.3.0.tar.gz
Algorithm	Hash digest
SHA256	`16c83a1131b34b9148ab1b7405c327efe5942966eea7f777eef4ec5845cfdeb8`
MD5	`e6e7d8a6a87aca1ed7aebee1a31d3f3b`
BLAKE2b-256	`f5de0f1177da06a84fd82283c4f8b26b43dd6841b6f3de46c4c9bec028e15d2e`

See more details on using hashes here.

Provenance

File details

Details for the file thepipe_api-0.3.0-py3-none-any.whl.

File metadata

Download URL: thepipe_api-0.3.0-py3-none-any.whl
Upload date: Apr 20, 2024
Size: 18.4 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: twine/3.8.0 pkginfo/1.9.6 readme-renderer/37.3 requests/2.31.0 requests-toolbelt/0.10.1 urllib3/1.26.18 tqdm/4.66.2 importlib-metadata/6.11.0 keyring/24.3.0 rfc3986/1.5.0 colorama/0.4.6 CPython/3.10.8

File hashes

Hashes for thepipe_api-0.3.0-py3-none-any.whl
Algorithm	Hash digest
SHA256	`3f36bacdf30ba62211fa654160856cbd88c77f9bc473b9f822eb1bfb318230d5`
MD5	`1dee6ebd29b8d48d13350f7d23a59c8f`
BLAKE2b-256	`be516203e133e3f9c063a0ba49bcfdb3ab28139499e8d39a3b556063310c0356`