Skip to main content

Ollama Python Library

The Ollama Python library provides the easiest way to integrate Python 3.8+ projects with Ollama.

Prerequisites

  • Ollama should be installed and running
  • Pull a model to use with the library: ollama pull <model> e.g. ollama pull gemma4
    • See Ollama.com for more information on the models available.

Install

pip install ollama

Usage

from ollama import chat
from ollama import ChatResponse

response: ChatResponse = chat(
  model='gemma4',
  messages=[
    {
      'role': 'user',
      'content': 'Why is the sky blue?',
    },
  ],
)
print(response['message']['content'])
# or access fields directly from the response object
print(response.message.content)

See _types.py for more information on the response types.

Streaming responses

Response streaming can be enabled by setting stream=True.

from ollama import chat

stream = chat(
  model='gemma4',
  messages=[{'role': 'user', 'content': 'Why is the sky blue?'}],
  stream=True,
)

for chunk in stream:
  print(chunk['message']['content'], end='', flush=True)

Cloud Models

Run larger models by offloading to Ollama’s cloud while keeping your local workflow.

  • Supported models: deepseek-v3.1:671b-cloud, gpt-oss:20b-cloud, gpt-oss:120b-cloud, kimi-k2:1t-cloud, qwen3-coder:480b-cloud, kimi-k2-thinking See Ollama Models - Cloud for more information

Run via local Ollama

  1. Sign in (one-time):
ollama signin
  1. Pull a cloud model:
ollama pull gpt-oss:120b-cloud
  1. Make a request:
from ollama import Client

client = Client()

messages = [
  {
    'role': 'user',
    'content': 'Why is the sky blue?',
  },
]

for part in client.chat('gpt-oss:120b-cloud', messages=messages, stream=True):
  print(part.message.content, end='', flush=True)

Cloud API (ollama.com)

Access cloud models directly by pointing the client at https://ollama.com.

  1. Create an API key from ollama.com , then set:
export OLLAMA_API_KEY=your_api_key
  1. (Optional) List models available via the API:
curl https://ollama.com/api/tags
  1. Generate a response via the cloud API:
import os
from ollama import Client

client = Client(host='https://ollama.com', headers={'Authorization': 'Bearer ' + os.environ.get('OLLAMA_API_KEY')})

messages = [
  {
    'role': 'user',
    'content': 'Why is the sky blue?',
  },
]

for part in client.chat('gpt-oss:120b', messages=messages, stream=True):
  print(part.message.content, end='', flush=True)

Custom client

A custom client can be created by instantiating Client or AsyncClient from ollama.

All extra keyword arguments are passed into the httpx.Client.

from ollama import Client

client = Client(host='http://localhost:11434', headers={'x-some-header': 'some-value'})
response = client.chat(
  model='gemma4',
  messages=[
    {
      'role': 'user',
      'content': 'Why is the sky blue?',
    },
  ],
)

Async client

The AsyncClient class is used to make asynchronous requests. It can be configured with the same fields as the Client class.

import asyncio
from ollama import AsyncClient


async def chat():
  message = {'role': 'user', 'content': 'Why is the sky blue?'}
  response = await AsyncClient().chat(model='gemma4', messages=[message])


asyncio.run(chat())

Setting stream=True modifies functions to return a Python asynchronous generator:

import asyncio
from ollama import AsyncClient


async def chat():
  message = {'role': 'user', 'content': 'Why is the sky blue?'}
  async for part in await AsyncClient().chat(model='gemma4', messages=[message], stream=True):
    print(part['message']['content'], end='', flush=True)


asyncio.run(chat())

API

The Ollama Python library's API is designed around the Ollama REST API

Chat

ollama.chat(model='gemma4', messages=[{'role': 'user', 'content': 'Why is the sky blue?'}])

Generate

ollama.generate(model='gemma4', prompt='Why is the sky blue?')

List

ollama.list()

Show

ollama.show('gemma4')

Create

ollama.create(model='example', from_='gemma4', system='You are Mario from Super Mario Bros.')

Copy

ollama.copy('gemma4', 'user/gemma4')

Delete

ollama.delete('gemma4')

Pull

ollama.pull('gemma4')

Push

ollama.push('user/gemma4')

Embed

ollama.embed(model='gemma4', input='The sky is blue because of rayleigh scattering')

Embed (batch)

ollama.embed(model='gemma4', input=['The sky is blue because of rayleigh scattering', 'Grass is green because of chlorophyll'])

Ps

ollama.ps()

System One

import ollama

response = ollama.systemone(
  model='nimble',
  state='Our checkout has returned 500 errors since 9am.',
  questions={
    'team': {
      'type': 'choice',
      'instructions': 'Which team should handle this ticket?',
      'criteria': {'billing': 'Payments and refunds', 'technical': 'Software errors'},
    },
  },
)
print(response.answers['team'])

System One uses POST /v1/systemone and requires Ollama v0.35.0 or later with a compatible local model such as nimble. It returns one JSON response; streaming and cloud models are not supported.

state and question instructions accept text, JSON objects, or arrays. Questions are evaluated in their supplied order:

  • choice: 2–26 option keys mapped to descriptions; None uses the key as its description.
  • noul: probability of true, with optional {"false": "No", "true": "Yes"} descriptions.
  • score: 2–26 descriptions ordered lowest to highest; returns a potentially fractional, zero-based score.

Responses contain model, answers, and usage.input_tokens / usage.output_tokens. Confidence measures probability concentration, not calibrated correctness. Token usage comes from the server; output tokens are not necessarily zero. Optional keep_alive accepts seconds or a duration string. Requests use the client's existing host, headers, and HTTP error handling. The server validates its body and model context limits without truncating input.

Available as ollama.systemone, Client.systemone, and await AsyncClient.systemone. Question and answer types, including SystemOneResponse, are exported from ollama. See the combined question example.

Errors

Errors are raised if requests return an error status or if an error is detected while streaming.

model = 'does-not-yet-exist'

try:
  ollama.chat(model)
except ollama.ResponseError as e:
  print('Error:', e.error)
  if e.status_code == 404:
    ollama.pull(model)

Release files for ollama 0.6.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ollama 0.6.3
File Size Uploaded
ollama-0.6.3.tar.gz 56.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ollama 0.6.3
File Interpreter ABI Platform
ollama-0.6.3-py3-none-any.whl Python 3 none any Details

Total release size: 73.5 kB

Release files / ollama-0.6.3.tar.gz

Download URL ollama-0.6.3.tar.gz
Size 56.9 kB
Tags Source
SHA-256 checksum
How to use checksums
41fc49a8095c4a75939c4c1f8582e4d0671692fb6eac2a5a7ede8c9872b67096
BLAKE2b-256 checksum
How to use checksums
b897eeafe65594e4f4b25e443e068ef7d83aa3105b023e12e1c408c38669fc07
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.

Transparency log

Release files / ollama-0.6.3-py3-none-any.whl

Download URL ollama-0.6.3-py3-none-any.whl
Size 16.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6a20bc42c1a5f889295d7ec490d35e5132fc31f339561530f43a8abd4dbfe508
BLAKE2b-256 checksum
How to use checksums
4d6487505d9e006461233c21c8e66dc1ecee49c996090584b216abd0dd4a8322
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.6.3 This release

2 release files

0.6.2

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.4

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.9

2 release files

0.4.8

2 release files

0.4.7

2 release files

0.4.6

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.9

2 release files

0.1.8

2 release files

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.0

2 release files

0.0.1

2 release files

0.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page