Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

Django AI SDK

A Django SDK for building AI-powered applications with support for multiple LLM providers, RAG (Retrieval-Augmented Generation), and streaming responses.

Project Status: Read This First

This is an early preview. We're actively iterating on the API and learning from real usage. Here's what that means for you:

  • Expect breaking changes: APIs will shift as we find better patterns.
  • Migrations might be reset: Don't rely on database schema stability between versions.
  • Not for production: Use this for experimentation, prototypes, and side projects. Keep critical workloads elsewhere.
  • Watch the repo: Things change quickly. Star & watch to stay in the loop.
  • Your feedback shapes the SDK: Break things, open issues, tell us what hurts.

We'd love to have you along for the ride, just keep your seatbelt on.

Install

pip install django-ai-sdk

Or with uv:

uv add django-ai-sdk

Quick Start

1. Add to INSTALLED_APPS

# settings.py
INSTALLED_APPS = [
    ...
    "django_ai_sdk",
]

Then run python manage.py migrate.

2. Define your agent

# agents.py
from django_ai_sdk import Agent
from django_ai_sdk.adapters.base import Stream
from django_ai_sdk.generators import openai_responses_chat
from haystack import Pipeline


class HelpDeskAgent(Agent):
    name = "Help Desk"
    model = "gpt-5-mini"
    instructions = "You are a helpful support agent."
    llm = openai_responses_chat

    async def get_pipeline_adapter(self, thread_id=None, user=None):
        storage_adapter = await self.get_storage_adapter(thread_id)
        generator = self.get_llm()
        return Stream(
            pipeline=Pipeline(),
            generator=generator,
            storage_adapter=storage_adapter,
        )

For non-streaming tasks (title generation, structured output), use Run instead:

from django_ai_sdk.adapters.base import Run

    async def get_run_adapter(self, thread_id=None, user=None):
        return Run(generator=self.get_llm())

3. Return a streaming response

# views.py
from .agents import HelpDeskAgent

agent = HelpDeskAgent()


@router.post("/chat")
async def chat(request, payload: ChatRequest):
    return await agent.as_view(
        payload.messages,
        thread_id=payload.thread_id,
    )

Integrations

An integration gives an agent tools from a third party: an MCP server, or an API you wrap yourself. Each one is a small Django app that registers itself on ready().

Enable a shipped integration by installing its app and naming it on an agent:

# settings.py
INSTALLED_APPS = [
    "django_ai_sdk",
    "django_ai_sdk.integrations.mcp",  # required by any MCP-backed integration
    "django_ai_sdk.integrations.github",  # also: .linear, .notion, .weather
]

# INSTALLED_APPS decides which integrations exist; this configures them. The same
# shape as DATABASES or CACHES, keyed by integration name.
AI_SDK_INTEGRATIONS = {
    "github": {"TOKEN": env("GITHUB_MCP_TOKEN")},
}
class HelpDeskAgent(Agent):
    integrations = ["github"]

A missing credential never breaks boot: the integration reports that it needs setup and contributes no tools. Alongside secrets, each entry accepts URL, TOOLS, LABEL, SCOPE and AUTH, so a self-hosted server, a narrower tool allow-list, or per-user OAuth instead of a shared token is a settings change rather than a subclass.

It's an ordinary dict, so pulling values from a vault or an ini file needs no hook. Just call whatever you like inside it.

Adding your own

Three files, no models and no migrations:

# myapp/integrations/zendesk/apps.py
from django_ai_sdk.integrations import IntegrationAppConfig


class ZendeskConfig(IntegrationAppConfig):
    default = True  # Django needs this to pick your AppConfig over the base
    name = "myapp.integrations.zendesk"
    integration = "myapp.integrations.zendesk.integration.ZendeskIntegration"
# myapp/integrations/zendesk/integration.py
from django_ai_sdk.integrations import MCPIntegration


class ZendeskIntegration(MCPIntegration):
    name = "zendesk"  # registry key, and the AI_SDK_INTEGRATIONS key
    label = "Zendesk"
    url = "https://mcp.zendesk.example/mcp"
    auth = "token"  # "static" | "token" | "oauth"
    default_tools = []  # [] discovers every tool the server offers

Add "myapp.integrations.zendesk" to INSTALLED_APPS, add "zendesk": {"TOKEN": env("ZENDESK_API_TOKEN")} to AI_SDK_INTEGRATIONS, and list "zendesk" on an agent. For an API you wrap by hand, subclass APIIntegration and set tools to @haystack.tools.tool-decorated functions. django_ai_sdk/integrations/weather/ is a complete, credential-free example, and github/, linear/ and notion/ cover token and OAuth MCP servers.

Tools are namespaced per integration (zendesk_search_tickets), so two servers exposing the same tool name don't collide. Tool lists are cached stale-while-revalidate behind a per-integration circuit breaker, so a slow or dead server costs one bounded wait and then reports itself as degraded.

HTTP endpoints

The SDK ships no integrations router; it doesn't pick your web framework. Build list/connect/disconnect/reconnect over IntegrationService (demo/apps/integrations/views/ninja.py is a working reference), and include the OAuth callback, which must sit at a fixed URL:

(path("api/integrations/", include("django_ai_sdk.integrations.mcp.urls")),)

Tracing

Every Haystack span of a run — the agent, each step, each LLM call, each tool call — persisted as a Trace row and linked to the Thread and Message it was produced for. Opt in the same way Haystack itself treats tracing: installing the SDK changes nothing until you add the app.

INSTALLED_APPS = [
    "django_ai_sdk",
    "django_ai_sdk.tracing",
]

That's the whole switch. The app ships the model plus migrations, and its AppConfig.ready() enables the tracer at startup. If you'd rather call it yourself, tracing.enable_tracing(DefaultTracer()) swaps Haystack's single global tracer — last call wins.

Spans record timing, hierarchy, model name and token counts. They do not record prompts or replies unless you set HAYSTACK_CONTENT_TRACING_ENABLED=true, which stores the content of queries, documents and answers in the tags JSON column — treat that data as sensitive, and note that deleting a Thread or Message cascades to its traces. Token counts are captured either way: Haystack reports usage through the same content tags it gates, so the numbers are harvested before the flag decides whether the payload is kept. That flag is read once, when haystack.tracing is first imported, so it must be a real environment variable set before the process starts — assigning it in settings.py does nothing.

Some tags are static configuration rather than observability data, and they are large: haystack.agent.tools serializes every tool definition in full. Drop them by key:

AI_SDK_TRACING_EXCLUDED_TAGS = [
    "haystack.agent.tools",
    "haystack.agent.state_schema",
]

Reading traces

TraceService is the permission-checked entry point. A trace belongs to its thread, so every method resolves the thread through ThreadService — enforcing VIEW_THREAD and raising PermissionDenied — and raises ValueError when the thread or message doesn't exist. There's no separate trace permission domain.

from django_ai_sdk.tracing.services import TraceService

await TraceService.thread_traces(
    thread_id, user=user, message_id=None, operation_name=None, limit=100, offset=0
)
await TraceService.message_traces(message_id, user=user)
await TraceService.thread_token_usage(thread_id, user=user)
await TraceService.message_token_usage(message_id, user=user)

Message-scoped calls resolve the owning thread themselves, so a caller holding only a message id doesn't need to know its thread. All four return pydantic schemas from django_ai_sdk.tracing.schemas. For sync contexts (DRF class-based views), import the module-level thread_traces / thread_token_usage aliases instead.

For direct ORM work, where you're doing your own access control, the manager carries the same query helpers:

Trace.objects.for_thread(thread.id).token_usage()
# {'prompt_tokens': 1841, 'completion_tokens': 317, 'total_tokens': 2158}

Trace.objects.for_message(message.id).roots()  # the top span of the run
Trace.objects.llm_calls()  # one row per LLM call

Never Sum the token columns over raw rows. Usage is recorded on every span that wraps an LLM call and aggregated again onto the agent's own haystack.agent.run span, so a plain sum counts each call twice — three times for an Agent inside a Pipeline. .llm_calls() selects leaf spans, which are exactly the LLM calls in every path and never a rollup. Both the service and token_usage() already count that way.

Behaviour worth knowing

Correlation is automatic for anything streamed through the SDK: each run stamps its spans with the assistant Message it mints and the Thread it belongs to. For pipelines you run yourself, wrap them in bind():

from django_ai_sdk.tracing import bind

with bind(thread_id=thread.id, message_id=message.id):
    await pipeline.run_async({"messages": messages})

Writes are buffered: a whole trace tree lands in one bulk_create when its root span exits, scheduled on the event loop so the ORM never blocks it. Spans appear only once the run completes — a hung pipeline has no rows yet, so this is not a live view. await aflush() is the barrier when you need one (tests, shutdown).

One provider caveat: the OpenAI Responses API reports usage on a streamed response without being asked, but Chat Completions omits it unless stream_options.include_usage is requested. The adapters never reconfigure a generator, so declare that on the agent (llm_kwargs) when streaming through openai_chat; a generator that reports no usage leaves the token columns null.

Features

  • RAG Pipelines: BM25, ChromaDB, and Qdrant hybrid search with query expansion.
  • Streaming Responses: Built-in SSE streaming. Works with Vercel AI SDK protocol.
  • Conversation Storage: Automatic message persistence. Thread-based history out of the box.
  • Tool Calling: MCP, memory, and custom tools, all managed by your Agent.
  • Artifacts: 16 structured UI types (tables, plans, approval cards, code blocks, and more) submitted by the LLM via tool calls.
  • File Processing: Document upload with pipeline-based processing (text, CSV, JSON, DOCX, PPTX, XLSX. Extraction transforms for metadata embedding.
  • Integrations: Third-party tools as self-registering Django apps, with caching, circuit breaking and OAuth built in. See Integrations.
  • Tracing: Opt-in Haystack tracing persisted to the ORM, with per-thread and per-message token accounting. See Tracing.
  • Reindexing: Hot-reload documents. Cached embeddings with simple refresh API.

Documentation

Full documentation and examples: github.com/django-ai-sdk/django-ai-sdk

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

django_ai_sdk-0.1.2a6.tar.gz (187.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

django_ai_sdk-0.1.2a6-py3-none-any.whl (234.9 kB view details)

Uploaded Python 3

File details

Details for the file django_ai_sdk-0.1.2a6.tar.gz.

File metadata

  • Download URL: django_ai_sdk-0.1.2a6.tar.gz
  • Upload date:
  • Size: 187.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.2 {"installer":{"name":"uv","version":"0.11.2","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for django_ai_sdk-0.1.2a6.tar.gz
Algorithm Hash digest
SHA256 46495a50f1741f194f102f18221e64e87cd38d78413a3b7e00479faf6c9f8dd3
MD5 4ec7813a313f69527b5644d6f5693459
BLAKE2b-256 32c451e189ae61559a10a0419cc2a20845da15f4f4e8af589a2095772fbb9eaf

See more details on using hashes here.

File details

Details for the file django_ai_sdk-0.1.2a6-py3-none-any.whl.

File metadata

  • Download URL: django_ai_sdk-0.1.2a6-py3-none-any.whl
  • Upload date:
  • Size: 234.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.2 {"installer":{"name":"uv","version":"0.11.2","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for django_ai_sdk-0.1.2a6-py3-none-any.whl
Algorithm Hash digest
SHA256 cc9137a81278c7e5bf401e1393343c2f662fdd914c4ce1f60aecff9e5845659a
MD5 b2d91ce5046c0b2454d9a58f109abf47
BLAKE2b-256 61988a5f5977f87bc77582624a53b8d185ac8c2d164e77123ab3346ffc2c2ce3

See more details on using hashes here.

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page