Skip to main content

agent-learning

Native reinforcement learning SDK for AI agents. An in-process Learner optimizes a small, interpretable TaskPolicy over discrete agent choices (understand intent and complete task by choosing the right outcome).

TaskPolicies model reusable decisions among executable alternatives such as models, skills, tools, workflows, or workloads. Factual questions, ordinary chat, reporting, and learning automation are not policy tasks.

How it works

The SDK improves agents without LLM weight fine-tuning. There are no GPU fine-tune jobs and no opaque update cycles — just three pieces that run in your existing Python process:

  1. TaskPolicy is a softmax distribution over N discrete actions (e.g., "take action A", "take action B", "take action C"). It lives in Python and updates in milliseconds.

  2. Score evaluates each episode on-device with three stdlib scorers for intent resolution, task adherence, and task completion. Their scores are combined into a single scalar reward with no scoring endpoint or environment variables required. Azure AI evaluators remain available as an opt-in.

  3. Learner applies REINFORCE-with-baseline to update TaskPolicy logits directly from logged episodes. Updates are tiny gradient steps that run on local compute and persist through a pluggable store — in-memory or local files by default, with Azure Cosmos DB optional.

task-policy-decide closes the loop at execution time by returning the selected action plus historical correctness, reward, result summaries, and per-metric quality feedback for the agent to use on its next delegated decision.

It also returns a complexity-proportional autonomy assessment. A persisted profile covers intent ambiguity, context variability, outcome observability, decision impact, reversibility, and mandatory approval; action-space size is derived. The resulting low, standard, high, or critical tier scales required outcomes, Wilson confidence, reward, probability, margin, stable snapshots, and drift-audit rate. Autonomous executions continue learning from observable outcomes, while tier-scaled samples request user feedback to detect drift.

Every episode, reward, run, and deployment is captured by the configured store — in-memory or local files by default, or Azure Cosmos DB — giving you a complete lineage and audit trail of how the policy evolved over time.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

agent_learning-0.6.1.tar.gz (108.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

agent_learning-0.6.1-py3-none-any.whl (109.2 kB view details)

Uploaded Python 3

File details

Details for the file agent_learning-0.6.1.tar.gz.

File metadata

  • Download URL: agent_learning-0.6.1.tar.gz
  • Upload date:
  • Size: 108.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for agent_learning-0.6.1.tar.gz
Algorithm Hash digest
SHA256 a057bad17ece6e67d6e4422c47bdc498c18eb932917c7d6947f9d66e964645e8
MD5 b017be856ba5ac7192ce4346af370e8d
BLAKE2b-256 43170fbae217b7c87bd72196c251f16cd78566b7cb6a432b3e73c7269b8aa723

See more details on using hashes here.

Provenance

The following attestation bundles were made for agent_learning-0.6.1.tar.gz:

Publisher: publish.yaml on microsoft/agent-learning

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file agent_learning-0.6.1-py3-none-any.whl.

File metadata

  • Download URL: agent_learning-0.6.1-py3-none-any.whl
  • Upload date:
  • Size: 109.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for agent_learning-0.6.1-py3-none-any.whl
Algorithm Hash digest
SHA256 1b8bc578c5a2e19c2eb881d52d97bd2cbf562db17eadb8826a66a8130d7a7818
MD5 7d7b109ff95fe69eb2a2b9ed30f2344b
BLAKE2b-256 c0ea77997543bcbcace4e6405715efdf80c6bdfffeb0d4671233179f5b8fc00f

See more details on using hashes here.

Provenance

The following attestation bundles were made for agent_learning-0.6.1-py3-none-any.whl:

Publisher: publish.yaml on microsoft/agent-learning

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page