Ollama Python Library
The Ollama Python library provides the easiest way to integrate Python 3.8+ projects with Ollama.
Prerequisites
- Ollama should be installed and running
- Pull a model to use with the library:
ollama pull <model>e.g.ollama pull gemma3- See Ollama.com for more information on the models available.
Install
pip install ollama
Usage
from ollama import chat
from ollama import ChatResponse
response: ChatResponse = chat(model='gemma3', messages=[
{
'role': 'user',
'content': 'Why is the sky blue?',
},
])
print(response['message']['content'])
# or access fields directly from the response object
print(response.message.content)
See _types.py for more information on the response types.
Streaming responses
Response streaming can be enabled by setting stream=True.
from ollama import chat
stream = chat(
model='gemma3',
messages=[{'role': 'user', 'content': 'Why is the sky blue?'}],
stream=True,
)
for chunk in stream:
print(chunk['message']['content'], end='', flush=True)
Cloud Models
Run larger models by offloading to Ollama’s cloud while keeping your local workflow.
- Supported models:
deepseek-v3.1:671b-cloud,gpt-oss:20b-cloud,gpt-oss:120b-cloud,kimi-k2:1t-cloud,qwen3-coder:480b-cloud,kimi-k2-thinkingSee Ollama Models - Cloud for more information
Run via local Ollama
- Sign in (one-time):
ollama signin
- Pull a cloud model:
ollama pull gpt-oss:120b-cloud
- Make a request:
from ollama import Client
client = Client()
messages = [
{
'role': 'user',
'content': 'Why is the sky blue?',
},
]
for part in client.chat('gpt-oss:120b-cloud', messages=messages, stream=True):
print(part.message.content, end='', flush=True)
Cloud API (ollama.com)
Access cloud models directly by pointing the client at https://ollama.com.
- Create an API key from ollama.com , then set:
export OLLAMA_API_KEY=your_api_key
- (Optional) List models available via the API:
curl https://ollama.com/api/tags
- Generate a response via the cloud API:
import os
from ollama import Client
client = Client(
host='https://ollama.com',
headers={'Authorization': 'Bearer ' + os.environ.get('OLLAMA_API_KEY')}
)
messages = [
{
'role': 'user',
'content': 'Why is the sky blue?',
},
]
for part in client.chat('gpt-oss:120b', messages=messages, stream=True):
print(part.message.content, end='', flush=True)
Custom client
A custom client can be created by instantiating Client or AsyncClient from ollama.
All extra keyword arguments are passed into the httpx.Client.
from ollama import Client
client = Client(
host='http://localhost:11434',
headers={'x-some-header': 'some-value'}
)
response = client.chat(model='gemma3', messages=[
{
'role': 'user',
'content': 'Why is the sky blue?',
},
])
Async client
The AsyncClient class is used to make asynchronous requests. It can be configured with the same fields as the Client class.
import asyncio
from ollama import AsyncClient
async def chat():
message = {'role': 'user', 'content': 'Why is the sky blue?'}
response = await AsyncClient().chat(model='gemma3', messages=[message])
asyncio.run(chat())
Setting stream=True modifies functions to return a Python asynchronous generator:
import asyncio
from ollama import AsyncClient
async def chat():
message = {'role': 'user', 'content': 'Why is the sky blue?'}
async for part in await AsyncClient().chat(model='gemma3', messages=[message], stream=True):
print(part['message']['content'], end='', flush=True)
asyncio.run(chat())
API
The Ollama Python library's API is designed around the Ollama REST API
Chat
ollama.chat(model='gemma3', messages=[{'role': 'user', 'content': 'Why is the sky blue?'}])
Generate
ollama.generate(model='gemma3', prompt='Why is the sky blue?')
List
ollama.list()
Show
ollama.show('gemma3')
Create
ollama.create(model='example', from_='gemma3', system="You are Mario from Super Mario Bros.")
Copy
ollama.copy('gemma3', 'user/gemma3')
Delete
ollama.delete('gemma3')
Pull
ollama.pull('gemma3')
Push
ollama.push('user/gemma3')
Embed
ollama.embed(model='gemma3', input='The sky is blue because of rayleigh scattering')
Embed (batch)
ollama.embed(model='gemma3', input=['The sky is blue because of rayleigh scattering', 'Grass is green because of chlorophyll'])
Ps
ollama.ps()
Errors
Errors are raised if requests return an error status or if an error is detected while streaming.
model = 'does-not-yet-exist'
try:
ollama.chat(model)
except ollama.ResponseError as e:
print('Error:', e.error)
if e.status_code == 404:
ollama.pull(model)
Release files for ollama 0.6.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ollama-0.6.2.tar.gz | 53.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ollama-0.6.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 68.3 kB
Release files / ollama-0.6.2.tar.gz
| Download URL | ollama-0.6.2.tar.gz |
|---|---|
| Size | 53.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
936d55daa684f474364c098611c933626f8d6c7d67065c5b7ae0c477b508b07f
|
|
BLAKE2b-256 checksum How to use checksums |
fc725f12423b6b39ca8430fbe56f77fcf4ef60f63067c7c4a2e30e200ed9ec16
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Apr 29, 2026.
Transparency logRelease files / ollama-0.6.2-py3-none-any.whl
| Download URL | ollama-0.6.2-py3-none-any.whl |
|---|---|
| Size | 15.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3ad7daab28e5a973445c36a73882a3ef698c2ebb00e21e308652741577509f7d
|
|
BLAKE2b-256 checksum How to use checksums |
c4abd6722beeb2d10f7a3b9ff49375708904fde18f82b5609a0bc4aeb5996a4d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Apr 29, 2026.
Transparency log