castia
Idiomatic, FastAPI-style Python SDK for Microsoft Foundry hosted agents.
castia lets a hosted agent speak Foundry's three wire protocols — Activity
(Teams/Bot Framework), OpenAI responses, and invocations — through
protocol-named decorators, dependency injection (Depends), and typed builders
for messages, Adaptive Cards, entities, and invoke envelopes. You decorate a
handler, return a value, and the framework does the rest — the auth chains,
hosting, Activity routing, and telemetry stay out of your file.
Install
pip install castia
# or, with uv:
uv add castia
Requires Python 3.11+.
Quickstart
from castia import Agent, Depends, Model, Teams
app = Agent(name="my-agent")
def gpt4o() -> Model:
return Model("gpt-4o")
@app.activity(Teams.direct, Teams.group, Teams.channel_mention)
async def reply(text: str, model: Model = Depends(gpt4o)) -> str:
return await model.respond(text)
if __name__ == "__main__":
app.run()
Decorate a handler with the surfaces it answers on, return a str, and the
framework sends it as the Teams reply. The model is built once for the
process and injected via Depends — the model choice stays visible in your file
instead of being buried in the framework. (Model() with no argument falls back
to AZURE_AI_MODEL_DEPLOYMENT_NAME; get_model / use_model("gpt-4o") are
zero-config conveniences.)
Composing protocols with routers
Like FastAPI's include_router, an Agent composes Routers so each protocol
can live in its own module:
from castia import Agent
from handlers import activity, responses, invocations
app = Agent(name="my-agent")
app.include(activity.router, responses.router, invocations.router)
Richer replies
Handlers can take a Message and reach for typed builders — Adaptive Cards,
suggested actions, citations, mentions, sensitivity labels, live-typing
streamers, and reactions:
from castia import Depends, Message, Model, Reaction, Router, Teams, get_model
router = Router()
@router.activity(Teams.direct)
async def reply(text: str, msg: Message, model: Model = Depends(get_model)) -> None:
await msg.react(Reaction.eyes)
answer = await model.respond(text)
await msg.say(answer)
Consume a toolbox
A Foundry toolbox is a curated set of tools the platform exposes behind one
MCP-compatible endpoint, with centralized auth, governance, and versioning.
castia turns it into a single Responses-API mcp tool spec the model service
resolves server-side — no local impl, no function loop. Build the spec and
hand it to the model as an extra_specs entry:
from castia import Depends, Model, Teams, get_model, toolbox_mcp_tool, toolbox_token
@app.activity(Teams.direct)
async def reply(text: str, model: Model = Depends(get_model)) -> str:
tool = toolbox_mcp_tool(token=await toolbox_token()) # reads TOOLBOX_* env
return await model.respond_with_tools(
text, tools=[], activity=None, extra_specs=[tool] if tool else []
)
toolbox_mcp_tool() (called with no endpoint) reads the environment via
resolve_toolbox_endpoint, whose precedence is:
- an explicit full URL in
TOOLBOX_ENDPOINT/TOOLBOX_MCP_ENDPOINT; - the platform-native
TOOLBOX_<NAME>_MCP_ENDPOINTthat theazd ai toolboxextension writes, keyed offTOOLBOX_NAME(seeplatform_endpoint_env); - composed from
FOUNDRY_PROJECT_ENDPOINT+TOOLBOX_NAME(+ optionalTOOLBOX_VERSION). The unversioned URL resolves the promoted default version, so a version bump needs no redeploy.
It returns None when no toolbox is configured, so "no toolbox" just attaches no
tool. Auth is either a bearer token (minted from the container's managed
identity by toolbox_token()) or a stored-connection project_connection_id.
A Foundry IQ knowledge base is the same shape via knowledge_base_mcp_tool.
For optimizer-ready toolbox guidance, pass an override map for the selected
tools and declare the same pure spec provider with app.tools(...) so
python -m castia optimize can emit the federated tools to tools.json:
def contract_kb():
tool = toolbox_mcp_tool(
allowed_tools=["contracts-kb-mcp___knowledge_base_retrieve"],
descriptions={
"knowledge_base_retrieve": "Search governing contracts and billing policies."
},
param_guidance={
"knowledge_base_retrieve": {
"query": "A natural-language contract or billing-policy question."
}
},
)
return [tool] if tool else []
app.tools(contract_kb)
Override keys may be the selected toolbox tool name
(contracts-kb-mcp___knowledge_base_retrieve) or the bare tool name
(knowledge_base_retrieve). The Responses API reports server_label
separately from the tool name, so castia emits the selected toolbox tool name
to tools.json; the older toolbox___... spelling is accepted as an alias.
Unknown or ambiguous names raise immediately. Runtime still uses the validated
server-side mcp call path; castia folds the overrides into
server_description for model-visible guidance and into a private optimizer
sidecar that is stripped before the Responses API call.
Deploying: the azd ai toolbox extension writes TOOLBOX_<NAME>_MCP_ENDPOINT
into the azd environment, but does not auto-inject it into a hosted
container — declare that env passthrough on your container yourself (there is no
azure.yaml/manifest step in castia for it).
Gotcha: do not copy
rai_config.rai_policy_name: Microsoft.Defaultfrom theazd ai toolbox create --helpexample — it is invalid on the project and 500s at tool enumeration (tools/list). Omit thepoliciesblock.
Validation status. Validated live (Foundry Responses path, App Insights-traced): the end-to-end pipe (env → compose URL → attach one
mcptool →tools/list+tools/call), a rawhttps://ai.azure.combearer minted in-container (noproject_connection_idrequired), both env forms, and unversioned→default-version resolution. Doc-derived / not yet live:knowledge_base_mcp_tool(Foundry IQ), connection-backed tools (Azure AI Search / remote-MCP / A2A), the Activity path with a toolbox, and approval-gated tools (require_approvalother than"never").
Protocols
castia publishes handlers for the protocols in PUBLISHABLE_PROTOCOLS:
- Activity — Teams / Bot Framework message and invoke turns.
responses— the OpenAIresponseswire shape.invocations— Foundry invoke envelopes (tool execution, agent-to-agent).
Observability & evaluation
castia configures Foundry/Agent 365 telemetry for you when the agent starts.
By default it emits GenAI spans (the chat {model} spans the Foundry Traces UI
keys off) but does not record the prompt/response content onto them.
Recording content is what makes an agent's traces evaluable — trace-based
evaluators read the input/output text from the GenAI spans, which is only present
when content recording is enabled. Turn it on deliberately via
configure_observability:
from castia.observability import configure_observability
# Records prompt/response text onto GenAI spans so traces can be evaluated.
configure_observability(enable_content_recording=True)
Resolution order for each flag is explicit argument > environment variable > default:
| Flag | Argument | Environment variable | Default |
|---|---|---|---|
| Content recording | enable_content_recording |
AZURE_TRACING_GEN_AI_CONTENT_RECORDING_ENABLED |
off |
| GenAI tracing | enable_genai_tracing |
AZURE_EXPERIMENTAL_ENABLE_GENAI_TRACING |
on |
Passing nothing preserves the default behavior. Telemetry setup is best-effort: a failure is logged, never raised, so it can't break startup or a turn.
Security caveat: enabling content recording writes prompt and response text to Application Insights. Only enable it where storing that content is acceptable for your data-handling and privacy requirements.
Building a scored eval suite
Once your traces are evaluable, python -m castia eval wraps the
azd ai agent eval extension to synthesize and run a scored eval suite — a
generated JSONL dataset plus an auto-generated, weighted rubric (a custom
multi-dimension evaluator):
# Offline gate — validate eval.yaml + rubric files, no Azure, free in CI:
python -m castia eval check
# Synthesize a rubric + dataset from the agent instruction (billable):
python -m castia eval generate --agent my-agent --max-samples 25
# Re-upload locally edited rubric/dataset files as a new version:
python -m castia eval update --evaluator-only
# Submit a scored run against the deployed agent (billable):
python -m castia eval run
check is a pure, offline referential-integrity gate: it resolves every
evaluator/dataset local_uri and validates each rubric dimensions file.
generate and run submit billable Foundry jobs, so both accept
--dry-run to print the resolved azd command line without submitting
anything. The azd wrappers need the build-time extra: pip install 'castia[deploy]'.
The rubric dimensions file is a bare JSON list where each entry is keyed by
id (a stable slug like correct_outcome), with an optional
always_applicable: true on the catch-all dimension. That cross-SDK shape is
pinned in the monorepo at spec/conformance/rubric/.
Optimizer-readiness
The Foundry Agent Optimizer searches for a better system prompt (and, when you declare tools, better tool descriptions) by running candidates against your eval suite. Making a castia agent optimizer-ready is three things: install the runtime resolver, ship a baseline config, and source your model + instructions from that config instead of hardcoding them — so the optimizer can swap in a candidate with zero handler changes.
Bind your model dependency with configured_model() and thread its resolved
instructions through:
from castia import Depends, Model, Router, configured_model
router = Router()
gpt = configured_model() # resolves baseline (or the injected candidate) once
@router.responses()
async def reply(text: str, model: Model = Depends(gpt)) -> str:
return await model.respond(text) # instructions flow into responses.create
configured_model() calls load_agent_config(), which is best-effort: if
the optimizer package isn't installed, resolution fails, or no config is found,
it degrades to environment defaults (AZURE_AI_MODEL_DEPLOYMENT_NAME, no
instructions) — the agent runs identically with or without the optimizer.
Ship a baseline under .agent_configs/baseline/:
.agent_configs/baseline/
metadata.yaml # model, instruction_file, (optional) tool_file pointers
instructions.md # the system prompt the optimizer tunes
tools.json # optional: tool specs the optimizer may reword
If your agent declares tools with app.tools(...), keep the baseline
tools.json in sync with the code using the build-time reconciler:
python -m castia optimize # write/refresh .agent_configs/baseline/tools.json
python -m castia optimize --check # CI drift gate (exits non-zero, writes nothing)
Then submit optimizer jobs directly through castia. This path reads the same
eval.yaml and baseline files, inlines local JSONL datasets for the request,
and preserves tools.json so toolbox/federated tool descriptions are visible to
the optimizer:
python -m castia optimize run --dry-run # FREE: print the exact payload
python -m castia optimize run # billable: submit to Foundry
python -m castia optimize status --watch # poll the latest castia-submitted job
python -m castia optimize apply # write the best candidate locally
python -m castia optimize cancel # cancel the latest job
optimize run uses FOUNDRY_PROJECT_ENDPOINT (or --project-endpoint) and
the model/evaluator/dataset declarations in eval.yaml. optimize apply
materializes the candidate under .agent_configs/<candidate-id>/ with
metadata.yaml, instructions.md, tools.json, and skills/ as returned by
Foundry. Deployment is still an azd handoff: set
OPTIMIZATION_LOCAL_DIR=.agent_configs and
OPTIMIZATION_CANDIDATE_ID=<candidate-id>, then deploy the hosted agent with
your existing azd workflow.
For preview-service drift detection, the repo includes
.github/workflows/foundry-optimizer-live.yml. It runs daily (and on manual
dispatch), submits one billable optimizer candidate against a pre-deployed
Foundry smoke agent, waits for completion, and applies the best candidate into a
throwaway runner directory. Configure the foundry-live GitHub environment with
OIDC Azure login secrets (AZURE_CLIENT_ID, AZURE_TENANT_ID,
AZURE_SUBSCRIPTION_ID) and these variables:
| Variable | Purpose |
|---|---|
FOUNDRY_PROJECT_ENDPOINT |
Target Foundry project endpoint. |
FOUNDRY_OPTIMIZER_AGENT_NAME |
Pre-deployed smoke agent name. |
FOUNDRY_OPTIMIZER_AGENT_VERSION |
Optional pinned hosted-agent version. |
FOUNDRY_EVAL_MODEL |
Optional evaluator model, defaults to gpt-4o. |
FOUNDRY_OPTIMIZE_MODEL |
Optional optimizer model, defaults to gpt-5. |
Config resolution order is first-wins: OPTIMIZATION_CONFIG (inline JSON) →
resolver API (OPTIMIZATION_CANDIDATE_ID + OPTIMIZATION_RESOLVE_ENDPOINT) →
local .agent_configs/ → environment defaults. An explicit config_dir
argument (or OPTIMIZATION_LOCAL_DIR) affects only the local source — pass
it anchored to your app root so the baseline resolves the same under
python app.py and python -m castia. That contract is pinned for every SDK in
spec/conformance/optimization/.
Switching to a reasoning (or RFT-tuned) model
The optimizer's model search can land on a reasoning model — an o-series or
GPT-5 deployment, or one you mint yourself with reinforcement fine-tuning (RFT).
Those models take a reasoning.effort control that plain chat models don't.
Model exposes it as reasoning_effort (minimal|low|medium|high):
o4 = use_model("o4-mini-rft-2025", reasoning_effort="high")
An unset effort omits the field entirely, so chat models are called exactly as
before; a bad level raises at construction rather than as a 400 mid-turn. An
operator can also switch a deployed agent onto a reasoning model with zero
code by setting MODEL_REASONING_EFFORT — an explicit argument still wins.
Because configured_model() builds its Model through the same path, that env
override flows through to the resolved candidate automatically.
Responses-only constraint: the optimizer accepts only single-protocol
responsesagents — submitting a multi-protocol agent (one that also speaks activity/invocations) is rejected with a400at submission. Project a responses-only sibling from the same handler code withapp.responses_only(), deploy that as its own service, optimize it, then apply the winning.agent_configscandidate back to your live agent.
The runtime resolver and the reconciler need the optimizer extra: pip install 'castia[optimize]'.
Reinforcement fine-tuning (RFT)
RFT is the fourth lifecycle step — build → evaluate → optimize → switch
models. It trains a reasoning model against a grader (a reward function)
instead of labeled answers, minting a new fine-tuned deployment that becomes a
candidate in the optimizer's model search. castia ships the build-time tooling
to prepare, validate, and (behind one guarded seam) submit an RFT job — the same
three-seam shape as the eval suite: pure builders, offline validators, and one
billable submit seam.
# Offline gate — validate an RFT dataset (+ grader), no Azure, free in CI:
python -m castia finetune check --dataset train.jsonl --validation val.jsonl --grader grader.json
# Bridge an eval rubric into a score_model grader (offline):
python -m castia finetune grader --rubric rubric.json --model gpt-4o --out grader.json
# Submit a billable RFT job (use --dry-run to print the payload and submit nothing):
python -m castia finetune submit --model o4-mini --dataset train.jsonl \
--validation val.jsonl --grader grader.json --reasoning-effort high --dry-run
A grader is one of string_check, text_similarity, score_model, python,
multi, or endpoint (preview); templates reference two namespaces only —
{{ sample.output_text }} and {{ item.<field> }}. Datasets are JSONL chat
messages[] rows whose final message role must be user, with extra
top-level keys as the item.* ground truth; both train and validation
splits are required. rubric_to_score_model() bridges an eval rubric straight
into a score_model grader — the natural tie between the evaluate and
switch-models steps.
⚠️ Provisional / doc-derived. The grader JSON schema, the RFT hyperparameter names, and that
fine_tuning.jobs.createaccepts this payload are derived from the Foundry RFT how-to and have not been confirmed against a live RFT job. The builders and validators are fully offline-tested; treat the submitted wire shape as provisional until a real submission validates it. The language-neutral contract is pinned inspec/conformance/graders/.
Design
castia is deliberately import-cheap: import castia never pulls in the
instrumented Azure/OpenAI/httpx stacks, so telemetry can be configured before
those libraries load. The heavy imports are deferred into the methods that need
them.
License
MIT © 2026 Seth Juarez
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file castia-0.5.0.tar.gz.
File metadata
- Download URL: castia-0.5.0.tar.gz
- Upload date:
- Size: 124.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
fb91cb4bd2d67a41a91f4d076147d0efb0fb27ba8dc7600cd61247baed426d0f
|
|
| MD5 |
617d41db93a63d6b57858351d5c30da3
|
|
| BLAKE2b-256 |
c5e0ee818e43e5b8f88e28d0cf700ac491f71d899fffe31c0e22f989939b33a8
|
Provenance
The following attestation bundles were made for castia-0.5.0.tar.gz:
Publisher:
release-please.yml on sethjuarez/castia
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
castia-0.5.0.tar.gz -
Subject digest:
fb91cb4bd2d67a41a91f4d076147d0efb0fb27ba8dc7600cd61247baed426d0f - Sigstore transparency entry: 2787271394
- Sigstore integration time:
-
Permalink:
sethjuarez/castia@85fcaafc13d752f71a14bf87b1a472cb54a05d6a -
Branch / Tag:
refs/heads/main - Owner: https://github.com/sethjuarez
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release-please.yml@85fcaafc13d752f71a14bf87b1a472cb54a05d6a -
Trigger Event:
push
-
Statement type:
File details
Details for the file castia-0.5.0-py3-none-any.whl.
File metadata
- Download URL: castia-0.5.0-py3-none-any.whl
- Upload date:
- Size: 115.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
072da5e9cb8a7efbd5418ddfaa23ac0cd5d88e306e7dcf99923e0a8ec9a2fdee
|
|
| MD5 |
f2dcb11ccc30dd86885477b6b834d24f
|
|
| BLAKE2b-256 |
104cbb43959069c4d38758575665ab764c3827f2b0bcd0476b05d9d6d0b1577f
|
Provenance
The following attestation bundles were made for castia-0.5.0-py3-none-any.whl:
Publisher:
release-please.yml on sethjuarez/castia
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
castia-0.5.0-py3-none-any.whl -
Subject digest:
072da5e9cb8a7efbd5418ddfaa23ac0cd5d88e306e7dcf99923e0a8ec9a2fdee - Sigstore transparency entry: 2787271424
- Sigstore integration time:
-
Permalink:
sethjuarez/castia@85fcaafc13d752f71a14bf87b1a472cb54a05d6a -
Branch / Tag:
refs/heads/main - Owner: https://github.com/sethjuarez
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release-please.yml@85fcaafc13d752f71a14bf87b1a472cb54a05d6a -
Trigger Event:
push
-
Statement type: