Xberg document extraction tools for CrewAI agents — text, metadata, and batch extraction across 97 file formats
Project description
crewai-xberg
CrewAI tools backed by Xberg. Give an agent document intelligence: extract text, metadata, keywords, entities, and summaries from 97 file formats — PDF, DOCX, XLSX, HTML, images with OCR, and more. Extraction is async at the core; the batch tool routes many files through Xberg's extract_batch in a single native call.
Install
pip install crewai-xberg
Requires Python 3.10+.
Tools
| Tool | Input | Returns |
|---|---|---|
XbergExtractTool |
file_path |
Extracted text, plus any requested rich results. |
XbergExtractBatchTool |
file_paths |
One section per document via extract_batch, then errors. |
XbergExtractMetadataTool |
file_path |
Title, authors, dates, page/table/image counts, format info. |
Use
from crewai import Agent
from crewai_xberg import XbergExtractTool, XbergExtractBatchTool, XbergExtractMetadataTool
agent = Agent(
role="Document Analyst",
goal="Extract and analyze document content",
backstory="You process documents of any format.",
tools=[XbergExtractTool(), XbergExtractBatchTool(), XbergExtractMetadataTool()],
)
Call a tool directly to see its output:
tool = XbergExtractTool()
# Plain extraction
text = tool.run(file_path="report.pdf")
# Force OCR and surface keywords, entities, and a summary
enriched = tool.run(
file_path="scan.pdf",
output_format="markdown",
force_ocr=True,
extract_keywords=True,
extract_entities=True,
summarize=True,
)
# Many files in one batched extraction
combined = XbergExtractBatchTool().run(file_paths=["report.pdf", "notes.docx", "sheet.xlsx"])
Options
Both extraction tools accept the same options. Each toggles an Xberg ExtractionConfig capability:
| Option | Default | Effect |
|---|---|---|
output_format |
"markdown" |
plain, markdown, or html. |
force_ocr |
False |
Run OCR on every page, even with a text layer. |
chunk |
False |
Split into semantic chunks; report the chunk count. |
extract_keywords |
False |
Append a keyword list. |
extract_entities |
False |
Append named entities (people, orgs, locations). |
summarize |
False |
Append a short summary. |
Detected languages and tables are surfaced automatically when present.
Async
The tools are async at the core. Inside an event loop, await tool.arun(...); the synchronous run bridges to it and must not be called from a running loop.
Errors
A missing file raises from Xberg directly. In batch mode, per-file failures land in ExtractionResult.errors and are reported in a trailing Errors section instead of aborting the batch.
For the full API, see the Xberg documentation.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file crewai_xberg-1.0.0rc37.tar.gz.
File metadata
- Download URL: crewai_xberg-1.0.0rc37.tar.gz
- Upload date:
- Size: 9.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: uv/0.11.32 {"installer":{"name":"uv","version":"0.11.32","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d060136c22600b76e09f52f61d036ba871e5a173c1c7e59f792dce95ecdeac5c
|
|
| MD5 |
7614a536dee476db8334295ae00f4104
|
|
| BLAKE2b-256 |
0cfc8616718c52f018a4256efa0ba66cb160c4615c5d8add7065048b879adea6
|
File details
Details for the file crewai_xberg-1.0.0rc37-py3-none-any.whl.
File metadata
- Download URL: crewai_xberg-1.0.0rc37-py3-none-any.whl
- Upload date:
- Size: 8.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: uv/0.11.32 {"installer":{"name":"uv","version":"0.11.32","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2ea8f2d1fe61cd971babbea4379dc4019e818240a7649bbb780318335ca77609
|
|
| MD5 |
05fe066abda23b3216b44a8d2b2c4db8
|
|
| BLAKE2b-256 |
da28abd9ab7d30201fcf42afc4f70d2eea90efdd681a38210e71e0b666e7d736
|