Skip to main content

audr-adapter-nemo-relay

PyPI Python versions

The AUDR plugin for NVIDIA NeMo Relay turns every completed Relay LLM call and tool execution into one AUDR record for an audr.Client your application owns. It reads usage, identifiers and timings only: never prompts, responses, tool arguments or tool results.

Status: alpha. The record model tracks AUDR v1.0.0; until 1.0.0, a minor release may change the public API.

Setup

pip install "audr-adapter-nemo-relay[runtime]"

Requires Python 3.11 or later and nemo-relay 0.8 within 0.8.x, which the runtime extra installs. Relay is imported at activation, so importing this package alone leaves it out of the process.

Usage

Register the plugin once, then run Relay-managed work inside the plugin context. Every LLM call made with a response codec, and every tool execution, is then metered:

import asyncio

import nemo_relay
from audr import Attribution, Client, FileSink
from nemo_relay import plugin as relay_plugin
from audr_adapter_nemo_relay import PLUGIN_KIND, NeMoRelayConfig, NeMoRelayPlugin


async def call_model(request: nemo_relay.LLMRequest) -> nemo_relay.JsonValue:
    # Call your provider here; this returns a canned OpenAI Chat Completions response.
    return {
        "model": "gpt-4o-mini",
        "choices": [{"message": {"role": "assistant", "content": "It shipped."}}],
        "usage": {"prompt_tokens": 12, "completion_tokens": 7},
    }


async def main() -> None:
    async with Client(FileSink("audr.jsonl")) as client:
        usage_plugin = NeMoRelayPlugin(client=client)  # on the loop that owns the client
        config = NeMoRelayConfig(attribution_defaults=Attribution(environment="production"))
        relay_config = relay_plugin.PluginConfig(
            components=[relay_plugin.ComponentSpec(kind=PLUGIN_KIND, config=config.to_dict())]
        )
        relay_plugin.register(PLUGIN_KIND, usage_plugin)
        try:
            async with relay_plugin.plugin(relay_config):
                with nemo_relay.scope.scope(
                    "support-agent",
                    nemo_relay.ScopeType.Agent,
                    metadata={"audr": {"account_id": "acct_42", "subscription_id": "sub_7"}},
                ):
                    request = nemo_relay.LLMRequest(
                        {},
                        {
                            "model": "gpt-4o-mini",
                            "messages": [{"role": "user", "content": "Where is order 42?"}],
                        },
                    )
                    await nemo_relay.llm.execute(
                        "openai",
                        request,
                        call_model,
                        model_name="gpt-4o-mini",
                        response_codec=nemo_relay.codecs.OpenAIChatCodec(),
                    )
            await usage_plugin.drain(timeout=5)
        finally:
            relay_plugin.deregister(PLUGIN_KIND)


asyncio.run(main())

Leaving the async with block drains the client's queue and closes its sink. A runnable version with a canned model response, without network access or provider keys, is examples/agent_scope.py. examples/nemo_relay_chat.py is an interactive terminal chat against an OpenAI-compatible endpoint; install the example extra to run it.

Attribution

Attribution is resolved when a root scope starts, from its metadata["audr"] merged field by field over attribution_defaults. The scope wins, and labels merge by key:

import nemo_relay

with nemo_relay.scope.scope(
    "support-agent",
    nemo_relay.ScopeType.Agent,
    metadata={
        "audr": {
            "account_id": "acct_42",
            "subscription_id": "sub_7",
            "user_id": "u_8f14e45f",  # pseudonymous, never an email or a name
            "labels": {"feature": "support-chat"},
        }
    },
):
    ...  # Relay-managed LLM and tool calls
  • environment, user_id, account_id, subscription_id and labels are read. Child scopes inherit the snapshot taken when their root started.
  • A scope with no environment, or a production scope with no account_id, is skipped with a warning rather than billed to a guess. Omit environment from attribution_defaults to require it on every root scope.
  • The reference states how evicted or completed ancestors are handled.

Records

Relay operation resource.operation usage
Each nemo_relay.llm.execute call with a response codec generation llm tokens, requests: 1
Each nemo_relay.tools.execute call tool_execution tool: { type: 'invocation', call_count: 1 }

Cache tokens are counted apart from input_tokens, and a provider-reported cost becomes cost.total_cost. Not metered: LLM calls without a response codec, whose end events carry no normalized usage.

Documentation

  • Reference: options, attribution, record fields, errors, diagnostics, operational bounds
  • Examples: agent_scope.py runs on a canned response, without network access
  • Changelog
  • AUDR specification, which defines every record field

License

Apache-2.0. Contributions follow CONTRIBUTING.md.

Metadata

Release files for audr-adapter-nemo-relay 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for audr-adapter-nemo-relay 0.3.0
File Size Uploaded
audr_adapter_nemo_relay-0.3.0.tar.gz 21.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for audr-adapter-nemo-relay 0.3.0
File Interpreter ABI Platform
audr_adapter_nemo_relay-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 45.4 kB

Release files / audr_adapter_nemo_relay-0.3.0.tar.gz

Download URL audr_adapter_nemo_relay-0.3.0.tar.gz
Size 21.1 kB
Tags Source
SHA-256 checksum
How to use checksums
bae720374f6f05f14a084898e5d7fbde5ea4d65888cb7d01f7bd93bbd37d1508
BLAKE2b-256 checksum
How to use checksums
3d163494391290e0106b989505d0cb67bb3c551328418ea0fa534f289b6d9ecc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release files / audr_adapter_nemo_relay-0.3.0-py3-none-any.whl

Download URL audr_adapter_nemo_relay-0.3.0-py3-none-any.whl
Size 24.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
76c8202389cac9aba2cf8e75bc43c63d640a017f6d2f7be1051b2131c552407e
BLAKE2b-256 checksum
How to use checksums
7f14b5aaefab599eb80a1c4e9b40f54f32058af7c07eb0c2b4b8b7b83ed6f98d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page