Skip to main content

JustAI

Package to make working with Large Language Models in Python super easy. Supports OpenAI, Anthropic Claude, Google Gemini, X Grok, DeepSeek, Perplexity, Reve, OpenRouter, Kimi (Moonshot), MiniMax and local GGUF models.

Author: Hans-Peter Harmsen (hp@harmsen.nl)
Current version: 5.6.10

Installation

  1. Install the package:
pip install justai
  1. Create an API key for the provider(s) you intend to use:

  2. Create a .env file with the relevant keys:

OPENAI_API_KEY=your-openai-api-key
ANTHROPIC_API_KEY=your-anthropic-api-key
GOOGLE_API_KEY=your-google-api-key
X_API_KEY=your-x-ai-api-key
DEEPSEEK_API_KEY=your-deepseek-api-key
PERPLEXITY_API_KEY=your-perplexity-api-key
MOONSHOT_API_KEY=your-moonshot-api-key
MINIMAX_API_KEY=your-minimax-api-key

Basic usage

from justai import Model

model = Model('gpt-5-mini')
model.system = """You are a movie critic. I feed you with movie
                  titles and you give me a review in 50 words."""

response = model.chat("Forrest Gump", cached=True)
print(response)

The cached=True parameter tells justai to cache the prompt and response locally.

Models

The provider is chosen automatically based on the model name prefix:

Prefix Provider
gpt*, o1*, o3* OpenAI
claude* Anthropic
gemini* Google
grok* X AI
deepseek* DeepSeek
sonar* Perplexity
reve* Reve
openrouter/* OpenRouter
kimi*, moonshot* Moonshot
minimax* (case-insensitive) MiniMax
*.gguf Local GGUF

Features

JSON and structured output

model = Model('gemini-2.5-flash')
prompt = 'Give me the main characters from Seinfeld. Return json with keys name, profession and weirdness'
data = model.chat(prompt, return_json=True)

For typed structured output, pass a Pydantic model or Python type as response_format:

from pydantic import BaseModel as PydanticModel

class Character(PydanticModel):
    name: str
    profession: str
    weirdness: str

result = model.chat(prompt, response_format=list[Character])

Images

Pass images as URLs, raw bytes or PIL images:

model = Model('gpt-5-nano')
url = 'https://upload.wikimedia.org/wikipedia/commons/9/94/Common_dolphin.jpg'
message = model.chat("What is in this image", images=url)

Image generation

model = Model('gpt-5')
pil_image = model.generate_image("A dolphin reading a book")

Input images can be passed for editing or style transfer:

model = Model('gemini-2.5-flash-image-preview')
pil_image = model.generate_image("Convert to Van Gogh style", images=source_image)

Async streaming

import asyncio

async def stream(model_name, prompt):
    model = Model(model_name)
    async for word in model.chat_async(prompt):
        print(word, end='')

asyncio.run(stream('sonar-pro', 'Give me 5 names for a juice bar'))

Prompt caching (Anthropic)

On by default. JustAI marks two cache breakpoints on every Anthropic request: one after the system prompt, which also covers the tool definitions, and one after the last turn of the conversation. Multi-turn chats and agent runs reuse their prefix instead of paying full price for it on every call.

The second breakpoint is only set from the second message on. A cache write costs 1.25x, so on a genuine one-shot call there would be nothing to earn it back.

cached_prompt still has a job: it moves a large fixed text into the cached prefix, ahead of the varying question.

model = Model('claude-sonnet-4-6')
model.system_message = 'You are an experienced book analyzer'
model.cached_prompt = SOME_LONG_TEXT
response = model.chat('Who is the main character?', cached=False)

Two settings, both optional:

model = Model('claude-sonnet-4-6', cache_ttl='1h')       # default '5m'
model = Model('claude-sonnet-4-6', prompt_cache=False)   # no breakpoints at all

A 1-hour write costs 2x rather than 1.25x and needs three requests to break even, so it is worth it only when your traffic has gaps longer than five minutes.

Below the minimum, caching silently does nothing. The shortest cacheable prefix is 1024 tokens on Sonnet 5, Sonnet 4.6 and Opus 4.8, 512 on Opus 5, 2048 on Opus 4.7, and 4096 on Opus 4.6 and Haiku 4.5. Under that the API caches nothing, without an error and without charging you for it. Zero counters on a short prompt are expected, not a bug.

Cache counters

model.cache_read_input_tokens      # served from cache, ~0.1x price
model.cache_creation_input_tokens  # written to cache, ~1.25x price

Available on every provider. Anthropic reports both; OpenAI and Gemini cache server-side on their own and report reads only. Providers that report nothing stay at zero, which means "not measured" rather than "no cache".

The two families count differently. On OpenAI and Gemini the cached tokens are a subset of the input tokens. On Anthropic they sit beside them, so the full prompt size is input_token_count + cache_read_input_tokens + cache_creation_input_tokens. If an agent ran for an hour and input_token_count reads 4000, the rest came from cache: check the sum, not the single field.

Effort

Control reasoning depth with a single portable setting. The library translates it to each provider's native parameter.

model = Model('claude-fable-5', effort='low')
# or
model = Model('gpt-5.6-terra')
model.effort = 'xhigh'

Valid values: 'low', 'medium', 'high', 'xhigh', 'max', or None (default — send nothing, provider default applies).

None vs 'none' — two different things:

Model('gpt-5.6-sol', effort=None)    # default, no reasoning field sent
Model('gpt-5.6-sol', effort='none')  # explicitly turn reasoning off (GPT-5.6 only)

'none' as a string is a pass-through only accepted by GPT-5.6 models. On every other provider it raises ValueError.

Provider Native support
Anthropic (Fable 5, Mythos 5, Opus 4.7/4.8, Sonnet 5) full set (low/medium/high/xhigh/max)
Anthropic (Opus 4.6, Sonnet 4.6) xhigh maps up to max with warning
Anthropic (Opus 4.5) xhigh and max map down to high with warning
OpenAI (gpt-5.6-*) full set; max maps to xhigh (SDK cap) with warning; also accepts 'none'
Google (Gemini 3.x) low/medium/high; xhigh/max map to HIGH with warning
xAI (grok-4.5, grok-4.3, grok-4.20-multi-agent) low/medium/high; xhigh/max map to high with warning
OpenRouter passed through raw; OpenRouter maps server-side
Older Anthropic (Sonnet 4.5, Haiku 4.5), older OpenAI (o1, o3), Gemini 2.x, DeepSeek, Perplexity, Reve, GGUF ignored with a warning (dedup'd per Model instance)

Downmap warnings use a dedicated EffortDownmapWarning category so you can filter them:

import warnings
from justai import EffortDownmapWarning

warnings.filterwarnings('ignore', category=EffortDownmapWarning)
# or promote to an error for strict pipelines:
warnings.filterwarnings('error', category=EffortDownmapWarning)

Effort is captured at request initiation. Mutating model.effort during an in-flight chat_async does not affect that request.

Related knobs to consider when raising effort:

  • Anthropic reasoning tokens count against max_tokens. On effort >= 'high' justai auto-raises max_tokens to 4096 if you did not set it explicitly (emits a UserWarning). Set max_tokens yourself to disable.
  • For effort='xhigh' or 'max' on any provider, the default timeout=120 seconds is often insufficient. Set timeout=300 (or higher) explicitly.

Level names are not calibrated across providers'high' on Anthropic burns different tokens than 'high' on OpenAI. Re-test cost/latency when switching models.

A failed call still reports its tokens. When the provider returns a response that justai then rejects — no text block because the whole budget went to thinking, JSON that will not parse, output truncated at the ceiling — the tokens were billed. last_token_count() reports them after the exception, so a caller adding up what a run cost does not lose that spend:

try:
    model.prompt('...')
except BadRequestException:
    spent = model.last_token_count()   # (input, output, total) of the failed call

When nothing came back at all (connection error, rate limit), the counters read (0, 0, 0) rather than the previous call's numbers.

Agent

JustAI includes an Agent class for autonomous, tool-using agent execution. The agent runs in a loop: it reads a task file, calls tools as needed, and returns a final answer.

Basic agent usage

import asyncio
from justai import Agent, FileSystemTool

agent = Agent(
    model='claude-sonnet-4-6',
    role='Code reviewer',
    goal='Review Python files and report issues',
    tools=[FileSystemTool(read=['/path/to/src'])],
    max_iterations=10,
)

async def main():
    async for event in agent.run('tasks.md'):
        if event.type == 'response':
            print(event.content, end='')
        elif event.type == 'done':
            print(f'\nAnswer: {event.result.answer}')

asyncio.run(main())

Built-in tools

FileSystemTool — read/write files with path traversal protection:

FileSystemTool(read=['/allowed/read/dir'], write=['/allowed/write/dir'])

ShellTool — run shell commands with allowlist-based security:

ShellTool(allowlist=['echo', 'ls', 'python'])

WebFetchTool — fetch URLs with SSRF protection:

WebFetchTool()

Custom tools

@agent.tool
def search_database(ctx, query: str) -> str:
    """Search the database for matching records."""
    return db.search(query)

Dynamic instructions

@agent.instructions
def inject_context(ctx) -> str:
    return f'Current user: {ctx.deps["username"]}'

Skills

Load .md skill files to extend the agent's system prompt:

agent = Agent(
    model='claude-sonnet-4-6',
    role='Assistant',
    goal='Help with tasks',
    skills_dir='./skills',
)

Agent events

The agent.run() async generator yields AgentEvent objects with these types:

  • status — status messages
  • response — streamed text from the model
  • tool_call — tool invocation (with name, arguments, tool_result)
  • error — error messages
  • done — final result with AgentResult (answer, audit trail, token usage, iterations)

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

justai-5.6.10.tar.gz (359.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

justai-5.6.10-py3-none-any.whl (69.5 kB view details)

Uploaded Python 3

File details

Details for the file justai-5.6.10.tar.gz.

File metadata

  • Download URL: justai-5.6.10.tar.gz
  • Upload date:
  • Size: 359.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for justai-5.6.10.tar.gz
Algorithm Hash digest
SHA256 195c1d0e9a18ee56c44cc277defe9036bd9089554f10559e459a4a9e92dbb369
MD5 f32c84f6ad7d057a4290eb6e44eebb0d
BLAKE2b-256 3188ce4dd0624b0fe5877a7abae0a028c4f4452675b8a7020e6ca6a5c9ff5fd3

See more details on using hashes here.

File details

Details for the file justai-5.6.10-py3-none-any.whl.

File metadata

  • Download URL: justai-5.6.10-py3-none-any.whl
  • Upload date:
  • Size: 69.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for justai-5.6.10-py3-none-any.whl
Algorithm Hash digest
SHA256 d5fc10f04bf1c78d1dee588dc2c995608fea860732d135124815ce8bd7f41e0d
MD5 744737cf32b8ecb3d43aff547a7cec50
BLAKE2b-256 2e06102f5e466c540c519bd7d9aa6cca916449a12ebb1b2c4229f4f2b137fe80

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

5.6.10 This release

2 files

5.6.9

2 files

5.6.8

2 files

5.6.7

2 files

5.6.6

2 files

5.6.5

2 files

5.6.4

2 files

5.6.3

2 files

5.6.2

2 files

5.6.1

2 files

5.5.7

2 files

5.5.6

2 files

5.5.5

2 files

5.5.4

2 files

5.5.3

2 files

5.5.2

2 files

5.5.1

2 files

5.5.0

2 files

5.4.16

2 files

5.4.15

2 files

5.4.14

2 files

5.4.13

2 files

5.4.12

2 files

5.4.11

2 files

5.4.9

2 files

5.4.7

2 files

5.4.6

2 files

5.4.5

2 files

5.4.3

2 files

5.4.1

2 files

5.4.0

2 files

5.3.0

2 files

5.2.1

2 files

5.2.0

2 files

5.1.1

2 files

5.1.0

2 files

5.0.0

2 files

4.2.5

2 files

4.2.4

2 files

4.2.3

2 files

4.2.2

2 files

4.2.1

2 files

4.2.0

2 files

4.1.4

2 files

4.1.3

2 files

4.1.1

2 files

4.1.0

2 files

4.0.13

2 files

4.0.12

2 files

4.0.11

2 files

4.0.10

2 files

4.0.9

2 files

4.0.8

2 files

4.0.7

2 files

4.0.6

2 files

4.0.4

2 files

4.0.3

2 files

4.0.2

2 files

4.0.1

2 files

4.0.0

2 files

3.11.7

2 files

3.11.6

2 files

3.11.5

2 files

3.11.4

2 files

3.11.3

2 files

3.11.2

2 files

3.11.1

2 files

3.11.0

2 files

3.10.2

2 files

3.10.1

2 files

3.10.0

2 files

3.9.0

2 files

3.8.1

2 files

3.8.0

2 files

3.7.1

2 files

3.7.0

2 files

3.6.0

2 files

3.5.2

2 files

3.5.1

2 files

3.5.0

2 files

3.4.7

2 files

3.4.6

2 files

3.4.5

2 files

3.4.3

2 files

3.4.2

2 files

3.4.1

2 files

3.4.0

2 files

3.3.2

2 files

3.3.1

2 files

3.2.12

2 files

3.2.11

2 files

3.2.10

2 files

3.2.9

2 files

3.2.8

2 files

3.2.7

2 files

3.2.6

2 files

3.2.5

2 files

3.2.4

2 files

3.2.3

2 files

3.2.2

2 files

3.2.1

2 files

3.2.0

2 files

3.1.10

2 files

3.1.9

2 files

3.1.8

2 files

3.1.7

2 files

3.1.6

2 files

3.1.5

2 files

3.1.4

2 files

3.1.3

2 files

3.1.2

2 files

3.1.1

2 files

3.0.1

2 files

3.0.0

2 files

2.1.6

2 files

2.1.5

2 files

2.1.4

2 files

2.1.3

2 files

2.1.2

2 files

2.1.1

2 files

2.1.0

2 files

2.0.20

2 files

2.0.19

2 files

2.0.18

2 files

2.0.17

2 files

2.0.16

2 files

2.0.4

2 files

2.0.3

2 files

2.0.2

2 files

2.0.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page