audr-adapter-nemo-relay
The AUDR plugin for NVIDIA NeMo Relay turns
every completed Relay LLM call and tool execution into one
AUDR record for an audr.Client your application
owns. It reads usage, identifiers and timings only: never prompts, responses, tool
arguments or tool results.
Status: alpha. The record model tracks AUDR v1.0.0; until 1.0.0, a minor release may change the public API.
Setup
pip install "audr-adapter-nemo-relay[runtime]"
Requires Python 3.11 or later and nemo-relay 0.8 within 0.8.x, which the runtime extra
installs. Relay is imported at activation, so importing this package alone leaves it out of
the process.
Usage
Register the plugin once, then run Relay-managed work inside the plugin context. Every LLM call made with a response codec, and every tool execution, is then metered:
import asyncio
import nemo_relay
from audr import Attribution, Client, FileSink
from nemo_relay import plugin as relay_plugin
from audr_adapter_nemo_relay import PLUGIN_KIND, NeMoRelayConfig, NeMoRelayPlugin
async def call_model(request: nemo_relay.LLMRequest) -> nemo_relay.JsonValue:
# Call your provider here; this returns a canned OpenAI Chat Completions response.
return {
"model": "gpt-4o-mini",
"choices": [{"message": {"role": "assistant", "content": "It shipped."}}],
"usage": {"prompt_tokens": 12, "completion_tokens": 7},
}
async def main() -> None:
async with Client(FileSink("audr.jsonl")) as client:
usage_plugin = NeMoRelayPlugin(client=client) # on the loop that owns the client
config = NeMoRelayConfig(attribution_defaults=Attribution(environment="production"))
relay_config = relay_plugin.PluginConfig(
components=[relay_plugin.ComponentSpec(kind=PLUGIN_KIND, config=config.to_dict())]
)
relay_plugin.register(PLUGIN_KIND, usage_plugin)
try:
async with relay_plugin.plugin(relay_config):
with nemo_relay.scope.scope(
"support-agent",
nemo_relay.ScopeType.Agent,
metadata={"audr": {"account_id": "acct_42", "subscription_id": "sub_7"}},
):
request = nemo_relay.LLMRequest(
{},
{
"model": "gpt-4o-mini",
"messages": [{"role": "user", "content": "Where is order 42?"}],
},
)
await nemo_relay.llm.execute(
"openai",
request,
call_model,
model_name="gpt-4o-mini",
response_codec=nemo_relay.codecs.OpenAIChatCodec(),
)
await usage_plugin.drain(timeout=5)
finally:
relay_plugin.deregister(PLUGIN_KIND)
asyncio.run(main())
Leaving the async with block drains the client's queue and closes its sink. A runnable
version with a canned model response, without network access or provider keys, is
examples/agent_scope.py.
examples/nemo_relay_chat.py
is an interactive terminal chat against an OpenAI-compatible endpoint; install the
example extra to run it.
Attribution
Attribution is resolved when a root scope starts, from its metadata["audr"] merged field
by field over attribution_defaults. The scope wins, and labels merge by key:
import nemo_relay
with nemo_relay.scope.scope(
"support-agent",
nemo_relay.ScopeType.Agent,
metadata={
"audr": {
"account_id": "acct_42",
"subscription_id": "sub_7",
"user_id": "u_8f14e45f", # pseudonymous, never an email or a name
"labels": {"feature": "support-chat"},
}
},
):
... # Relay-managed LLM and tool calls
environment,user_id,account_id,subscription_idandlabelsare read. Child scopes inherit the snapshot taken when their root started.- A scope with no
environment, or aproductionscope with noaccount_id, is skipped with a warning rather than billed to a guess. Omitenvironmentfromattribution_defaultsto require it on every root scope. - The reference states how evicted or completed ancestors are handled.
Records
| Relay operation | resource.operation |
usage |
|---|---|---|
Each nemo_relay.llm.execute call with a response codec |
generation |
llm tokens, requests: 1 |
Each nemo_relay.tools.execute call |
tool_execution |
tool: { type: 'invocation', call_count: 1 } |
Cache tokens are counted apart from input_tokens, and a provider-reported cost becomes
cost.total_cost. Not metered: LLM calls without a response codec, whose end events carry
no normalized usage.
Documentation
- Reference: options, attribution, record fields, errors, diagnostics, operational bounds
- Examples:
agent_scope.pyruns on a canned response, without network access - Changelog
- AUDR specification, which defines every record field
License
Apache-2.0. Contributions follow CONTRIBUTING.md.
Metadata
Release files for audr-adapter-nemo-relay 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| audr_adapter_nemo_relay-0.3.0.tar.gz | 21.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| audr_adapter_nemo_relay-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 45.4 kB
Release files / audr_adapter_nemo_relay-0.3.0.tar.gz
| Download URL | audr_adapter_nemo_relay-0.3.0.tar.gz |
|---|---|
| Size | 21.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
bae720374f6f05f14a084898e5d7fbde5ea4d65888cb7d01f7bd93bbd37d1508
|
|
BLAKE2b-256 checksum How to use checksums |
3d163494391290e0106b989505d0cb67bb3c551328418ea0fa534f289b6d9ecc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency logRelease files / audr_adapter_nemo_relay-0.3.0-py3-none-any.whl
| Download URL | audr_adapter_nemo_relay-0.3.0-py3-none-any.whl |
|---|---|
| Size | 24.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
76c8202389cac9aba2cf8e75bc43c63d640a017f6d2f7be1051b2131c552407e
|
|
BLAKE2b-256 checksum How to use checksums |
7f14b5aaefab599eb80a1c4e9b40f54f32058af7c07eb0c2b4b8b7b83ed6f98d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency log