Skip to main content

PyPI - Python Version GitHub Contributors GitHub Last Commit GitHub Repo Size GitHub Issues GitHub Pull Requests Github License

Aurelio SDK

The Aurelio Platform SDK. API references

Installation

To install the Aurelio SDK, use pip or poetry:

pip install aurelio-sdk

Authentication

The SDK requires an API key for authentication. Get key from Aurelio Platform. Set your API key as an environment variable:

export AURELIO_API_KEY=your_api_key_here

Usage

See examples for more details.

Initializing the Client

from aurelio_sdk import AurelioClient
import os

client = AurelioClient(api_key=os.environ["AURELIO_API_KEY"])

or use asynchronous client:

from aurelio_sdk import AsyncAurelioClient

client = AsyncAurelioClient(api_key="your_api_key_here")

Chunk

from aurelio_sdk import ChunkingOptions, ChunkResponse

# All options are optional with default values
chunking_options = ChunkingOptions(
    chunker_type="semantic", max_chunk_length=400, window_size=5
)

response: ChunkResponse = client.chunk(
    content="Your text here to be chunked", processing_options=chunking_options
)

Extracting Text from Files

PDF Files

from aurelio_sdk import ExtractResponse

# From a local file
file_path = "path/to/your/file.pdf"

response_pdf_file: ExtractResponse = client.extract_file(
    file_path=file_path, quality="low", chunk=True, wait=-1
)

Video Files

from aurelio_sdk import ExtractResponse

# From a local file
file_path = "path/to/your/file.mp4"


response_video_file: ExtractResponse = client.extract_file(
    file_path=file_path, quality="low", chunk=True, wait=-1
)

Extracting Text from URLs

PDF URLs

from aurelio_sdk import ExtractResponse

# From URL
url = "https://arxiv.org/pdf/2408.15291"
response_pdf_url: ExtractResponse = client.extract_url(
    url=url, quality="low", chunk=True, wait=-1
)

Video URLs

from aurelio_sdk import ExtractResponse

# From URL
url = "https://storage.googleapis.com/gtv-videos-bucket/sample/ForBiggerMeltdowns.mp4"
response_video_url: ExtractResponse = client.extract_url(
    url=url, quality="low", chunk=True, wait=-1
)

Waiting for completion and checking document status

# Set wait time for large files with `high` quality
# Wait time is set to 10 seconds
response_pdf_url: ExtractResponse = client.extract_url(
    url="https://arxiv.org/pdf/2408.15291", quality="high", chunk=True, wait=10
)

# Get document status and response
document_response: ExtractResponse = client.get_document(
    document_id=response_pdf_file.document.id
)
print("Status:", document_response.status)

# Use a pre-built function, which helps to avoid long hanging requests (Recommended)
document_response = client.wait_for(
    document_id=response_pdf_file.document.id, wait=300
)

Embeddings

from aurelio_sdk import EmbeddingResponse

response: EmbeddingResponse = client.embedding(
    input="Your text here to be embedded",
    model="bm25")

# Or with a list of texts
response: EmbeddingResponse = client.embedding(
    input=["Your text here to be embedded", "Your text here to be embedded"]
)

Response Structure

The ExtractResponse object contains the following key information:

  • status: The current status of the extraction task
  • usage: Information about token usage, pages processed, and processing time
  • message: Any relevant messages about the extraction process
  • document: The extracted document information, including its ID
  • chunks: The extracted text, divided into chunks if chunking was enabled

The EmbeddingResponse object contains the following key information:

  • message: Any relevant messages about the embedding process
  • model: The model name used for embedding
  • usage: Information about token usage, pages processed, and processing time
  • data: The embedded documents

Best Practices

  1. Use appropriate wait times based on your use case and file sizes.
  2. Use async client for better performance.
  3. For large files or when processing might take longer, enable polling for long-hanging requests.
  4. Always handle potential exceptions and check the status of the response.
  5. Adjust the quality parameter based on your needs. "low" is faster but less accurate, while "high" is slower but more accurate.

Metadata

Release files for aurelio-sdk 0.0.19

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for aurelio-sdk 0.0.19
File Size Uploaded
aurelio_sdk-0.0.19.tar.gz 15.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for aurelio-sdk 0.0.19
File Interpreter ABI Platform
aurelio_sdk-0.0.19-py3-none-any.whl Python 3 none any Details

Total release size: 32.6 kB

Release files / aurelio_sdk-0.0.19.tar.gz

Download URL aurelio_sdk-0.0.19.tar.gz
Size 15.3 kB
Tags Source
SHA-256 checksum
How to use checksums
14107e7440ff2efd0b4a08c52fb595e7680bd4bc973a0ddfb3b64157c6666b91
BLAKE2b-256 checksum
How to use checksums
270ec2e369ad173fb3d76448e46d10beb3dcc53388318933ddf8169a3f21a810
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.5.1 CPython/3.12.9 Linux/6.8.0-1021-azure

Release files / aurelio_sdk-0.0.19-py3-none-any.whl

Download URL aurelio_sdk-0.0.19-py3-none-any.whl
Size 17.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
390c0212b59ce99116df8722d3badced88c5ef0bb742a6222d479ceed0ed3948
BLAKE2b-256 checksum
How to use checksums
a21faa74b23b6eea4cf9b79ace914df59123c4c8e7e4bd32dd22d09c126422d9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.5.1 CPython/3.12.9 Linux/6.8.0-1021-azure

Release history Release notifications | RSS feed

This release

0.0.19 This release

2 release files

0.0.18

2 release files

0.0.17

2 release files

0.0.16

2 release files

0.0.15

2 release files

0.0.14

2 release files

0.0.13

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page