pareta
Python client for Pareta. One model id — "auto" — and
Pareta plans each request, routes it to benchmark-proven open specialists,
verifies the result, and falls back to a frontier model when that's the right
call. One request, one bill; you never pay for Pareta's orchestration or
cold starts.
pip install pareta # or: uv add pareta / poetry add pareta
from pareta import Pareta
pa = Pareta.from_env() # reads PARETA_API_KEY
# or: Pareta(api_key="pareta_sk_…", base_url="https://api.pareta.ai")
resp = pa.chat.completions.create(
model="auto", # the routing brain — the product
messages=[{"role": "user", "content": "Extract the total from this invoice: …"}],
)
print(resp.choices[0].message.content)
# Streaming (progress while Pareta plans + executes, then tokens)
for chunk in pa.chat.completions.create(model="auto", messages=[...], stream=True):
print(chunk.choices[0].delta.content or "", end="")
Async mirrors the sync client:
from pareta import AsyncPareta
async with AsyncPareta.from_env() as pa:
resp = await pa.chat.completions.create(model="auto", messages=[...])
Is it actually good? Measure it on YOUR data
Don't take the routing brain on faith — benchmark it. pa.evals runs "auto"
head-to-head against frontier models on your own ground truth and prices every
contender honestly:
run = pa.evals.runs.create(eval_set=es.id, models=["auto"],
frontier=["claude-opus-4-7"], wait=True)
And watch what your live traffic is doing — spend, success rate, and the projected savings vs calling a frontier directly:
pa.auto.metrics() # requests, success rate, spend, savings vs frontier
pa.auto.compare_frontier( # one prompt, metered, side-by-side with auto
model="gpt-5.5",
messages=[{"role": "user", "content": "…"}],
)
Inference is OpenAI-compatible
You don't even need this SDK to call Pareta — point the openai client at
base_url + your key and set model="auto":
from openai import OpenAI
client = OpenAI(api_key="pareta_sk_…", base_url="https://api.pareta.ai/v1")
resp = client.chat.completions.create(model="auto", messages=[...])
This SDK's unique value is everything AROUND that call — evals on your data, auto metrics, and the benchmark catalog — as Python methods, a CLI, and an MCP server.
Discovery
pa.tasks.match resolves a plain-language intent to the benchmarked task (or
capability lane) Pareta covers it with — feed the matched task into pa.evals
to prove "auto" on your own data:
m = pa.tasks.match("extract the key fields from these contracts")
m.type, m.chosen.task_id # "task", "contract-key-fields"
for model in pa.models.list(): # everything your org can call
print(model.id)
Auth
Mint a pareta_sk_ key in the dashboard (key management is browser-only) and
pass it as api_key= or via PARETA_API_KEY. The SDK only ever consumes a
key; it never creates, lists, or revokes them.
CLI
pip install "pareta[cli]" adds the pareta command (or pipx install "pareta[cli]" for an isolated, always-on-PATH install):
export PARETA_API_KEY=pareta_sk_…
pareta chat "Summarize this contract: …" # model:"auto" by default
pareta auto metrics # your auto traffic, rolled up
pareta auto compare "…prompt…" --frontier gpt-5.5 # auto vs a frontier, metered
pareta tasks match "extract fields from invoices" # intent → task
pareta evals run --task invoice-extraction --file rows.jsonl \
--models auto --frontier --wait # prove auto on your data
Add --json to any command for machine-readable output; pareta --help (or
pareta <group> --help) documents the full tree — chat, auto, tasks,
models, evals, audio.
MCP server
pareta-mcp is a Model Context Protocol
server (stdio) that exposes Pareta to an AI agent (Claude Desktop, Cursor, …) as
tools — chat (defaults to model="auto"), auto_metrics, compare_frontier,
run_eval / get_eval_run, discovery (match_task, list_tasks, get_task,
list_models), and audio (transcribe, speak).
Run it in its own isolated environment — like any MCP server it has its own
dependency tree, so don't pip install it into an app/project venv. The simplest
is uvx (no install, runs on demand). Register it
(Claude Desktop → Settings → Developer → Edit Config):
{
"mcpServers": {
"pareta": {
"command": "uvx",
"args": ["--from", "pareta[mcp]", "pareta-mcp"],
"env": { "PARETA_API_KEY": "pareta_sk_…" }
}
}
}
Prefer a persistent install? pipx install "pareta[mcp]" puts pareta-mcp on
your PATH in a dedicated venv — then use "command": "pareta-mcp". (Avoid a plain
pip install "pareta[mcp]" into a shared environment: its mcp/starlette
dependencies can clash with an app's FastAPI, and the console script may not land
on your PATH.) Inference, eval, and audio tools spend money; your MCP client's
per-tool-call approval is the guardrail.
Errors
All errors subclass pareta.ParetaError:
| Exception | When |
|---|---|
AuthenticationError (401) |
bad/missing key |
InsufficientCreditsError (402) |
org out of credit — top up in the dashboard |
NotFoundError (404) |
unknown resource (task, eval set, run, …) |
EndpointNotReadyError (503) |
serving capacity cold / provider down (retryable) |
RateLimitError (429) |
throttled (auto-retried) |
BadRequestError (400/422) |
malformed request |
APIConnectionError / APITimeoutError |
transport failure (auto-retried) |
Idempotent GETs and 429/5xx/timeouts are retried with exponential backoff
(max_retries, default 2).
Status
Live: model="auto" inference (buffered + streaming) with pa.auto metrics and
frontier comparison, plus models, tasks (browse + match), evals
(bring-your-own-data, with "auto" as a first-class contender), and audio —
and two interfaces over it: the pareta CLI (pip install "pareta[cli]")
and the pareta-mcp MCP server (pip install "pareta[mcp]"). Sync + async
clients.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file pareta-1.1.0.tar.gz.
File metadata
- Download URL: pareta-1.1.0.tar.gz
- Upload date:
- Size: 305.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ccc30e4ca852d3015a783f90864a3863bdfc96ebc837c2f3fca39df428c81112
|
|
| MD5 |
17db74f646560c98bb04388ca2f5877a
|
|
| BLAKE2b-256 |
0a666109e4dcc4586bf7eb8665212e8fb6fcd2cf3c9e7c8a65cb9e4705cd5d56
|
Provenance
The following attestation bundles were made for pareta-1.1.0.tar.gz:
Publisher:
publish.yml on Pareta-AI/pareta
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pareta-1.1.0.tar.gz -
Subject digest:
ccc30e4ca852d3015a783f90864a3863bdfc96ebc837c2f3fca39df428c81112 - Sigstore transparency entry: 2138399776
- Sigstore integration time:
-
Permalink:
Pareta-AI/pareta@53b7c60f29c5d49a9999c3b02880dcb113eb3f43 -
Branch / Tag:
refs/tags/v1.1.0 - Owner: https://github.com/Pareta-AI
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@53b7c60f29c5d49a9999c3b02880dcb113eb3f43 -
Trigger Event:
release
-
Statement type:
File details
Details for the file pareta-1.1.0-py3-none-any.whl.
File metadata
- Download URL: pareta-1.1.0-py3-none-any.whl
- Upload date:
- Size: 41.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d434f8f2282beef928a340db6c9d8b33d085d9fa938e538c55925a19f6bb43b8
|
|
| MD5 |
f1a89e60a043608d124983ec4bae51cf
|
|
| BLAKE2b-256 |
384659d0a145db36aa5bb7d1fef0c5335413a5601404fae5aa018c98e395960c
|
Provenance
The following attestation bundles were made for pareta-1.1.0-py3-none-any.whl:
Publisher:
publish.yml on Pareta-AI/pareta
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pareta-1.1.0-py3-none-any.whl -
Subject digest:
d434f8f2282beef928a340db6c9d8b33d085d9fa938e538c55925a19f6bb43b8 - Sigstore transparency entry: 2138399796
- Sigstore integration time:
-
Permalink:
Pareta-AI/pareta@53b7c60f29c5d49a9999c3b02880dcb113eb3f43 -
Branch / Tag:
refs/tags/v1.1.0 - Owner: https://github.com/Pareta-AI
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@53b7c60f29c5d49a9999c3b02880dcb113eb3f43 -
Trigger Event:
release
-
Statement type: