Skip to main content

Maxun Python SDK

The Maxun Python SDK turns websites and documents into structured data from your Python code. You create data scraping robots and run them whenever you need fresh data.

import asyncio
from maxun import Maxun

async def main():
    async with Maxun(api_key="your-api-key") as maxun:
        robot = await maxun.scrape("Maxun", "https://maxun.dev", formats=["markdown"])
        result = await robot.run()
        print(result.markdown)


asyncio.run(main())

Installation

pip install maxun

Requirements

  • Python 3.8+
  • A Maxun Cloud account or a self-hosted Maxun instance
  • An API key from the Maxun Dashboard

Configuration

Pass your API key directly:

from maxun import Maxun

maxun = Maxun(api_key="your-api-key")

Or set it in the environment and create Maxun() with no arguments:

MAXUN_API_KEY=your-api-key
MAXUN_TEAM_ID=your-team-uuid                      # optional, Maxun Cloud teams
MAXUN_BASE_URL=http://localhost:8080/api/sdk/     # only for self-hosted Maxun
from dotenv import load_dotenv
from maxun import Maxun

load_dotenv()          # only needed if your variables are in a .env file
maxun = Maxun()

The SDK connects to Maxun Cloud by default. For a self-hosted instance, set MAXUN_BASE_URL or pass base_url:

maxun = Maxun(api_key="your-api-key", base_url="http://localhost:8080/api/sdk/")

Maxun keeps one connection open. Use it with async with (as above), or call await maxun.close() when you are done.

Everything starts from maxun

Each call takes the robot name first, then what to work on (a URL, a search query or a file), then any settings as keyword arguments. It returns a Robot saved on your account.

Call What the robot does Read the result from
maxun.scrape(name, url) Turns a page into Markdown, HTML, text, links, a summary or screenshots result.markdown, result.html, ...
maxun.extract(name, url, prompt=...) Extracts structured data, described in plain English or with selectors result.list_data, result.text_data
maxun.crawl(name, url) Visits many pages of a website result.crawl_data
maxun.search(name, query) Searches the web and optionally scrapes the results result.search_data
maxun.documents.extract(name, file, prompt) Extracts data from a PDF, DOCX, XLSX, CSV, JPG or PNG result.document_data
maxun.documents.parse(name, file) Converts a document to Markdown, HTML, links or a summary result.markdown, ...
maxun.robots Lists, finds and deletes robots of any type

The name is required. It is how the robot appears in the Maxun dashboard.

Without async

Prefer plain function calls? MaxunSync has exactly the same methods, without await. It also works in Jupyter notebooks.

from maxun import MaxunSync

with MaxunSync(api_key="your-api-key") as maxun:
    robot = maxun.scrape("Example", "https://example.com")
    result = robot.run()
    print(result.markdown)

:::note The examples in these docs use await, so they need to run inside an async function like the one at the top of this page. With MaxunSync, drop the await. :::

Errors

Every API error is a MaxunError with .status_code and .details. More specific errors:

Error When
AuthenticationError The API key is missing or invalid
NotFoundError The robot or run does not exist
ConflictError A robot with that name already exists with different settings
ValidationError Maxun rejected the input
RunFailedError A run failed or was aborted
from maxun import ConflictError, MaxunError

try:
    robot = await maxun.scrape("Pricing page", "https://example.com/pricing")
except ConflictError:
    robot = await maxun.robots.find("Pricing page")
except MaxunError as error:
    print(error.status_code, error)

What's next

  • Scrape: turn pages into clean content
  • Extract: pull structured data out of pages
  • Crawl: collect content from a whole website
  • Search: search the web
  • Document: extract data from and convert documents
  • Monitoring: get notified when a page changes
  • Robot Management: run, schedule and manage robots

Metadata

Release files for maxun 0.0.13

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for maxun 0.0.13
File Size Uploaded
maxun-0.0.13.tar.gz 45.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for maxun 0.0.13
File Interpreter ABI Platform
maxun-0.0.13-py3-none-any.whl Python 3 none any Details

Total release size: 90.1 kB

Release files / maxun-0.0.13.tar.gz

Download URL maxun-0.0.13.tar.gz
Size 45.3 kB
Tags Source
SHA-256 checksum
How to use checksums
6af232f80d5fa116715cb64e92102d09fa96862b9de1bc908637817ef4aa787f
BLAKE2b-256 checksum
How to use checksums
20e7d94a400ffc2421c85daae375cb6285ffe813ba1618a4422bb9017a90f6f6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.5

Release files / maxun-0.0.13-py3-none-any.whl

Download URL maxun-0.0.13-py3-none-any.whl
Size 44.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b7f3c939aada431f4022362d07bd1c2c035f0c7c863b8128156f61ef32fc59ac
BLAKE2b-256 checksum
How to use checksums
cfddbb7fc9340cd8c9cd21b48c827a20fbe171e16eccc86eaff2f3f9dc66e68b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.10.5

Release history Release notifications | RSS feed

This release

0.0.13 This release

2 release files

0.0.12

2 release files

0.0.11

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.2

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page