agentmetrics-llamaindex
AgentMetrics integration for LlamaIndex. Call instrument() once and every agent run and query engine response reports back to your dashboard showing latency, cost, token counts, tool calls, and errors, via LlamaIndex's native instrumentation API with zero changes to your agent code.
Install
pip install agentmetrics-llamaindex
Quickstart
from agentmetrics_llamaindex import instrument
# Register on the global root dispatcher once at startup
span_handler = instrument(
agent_id="my-llamaindex-agent",
base_url="http://localhost:8099",
)
# Run your agents and query engines as normal
response = agent.chat("Summarize this document")
span_handler.flush()
API
instrument(agent_id, base_url)
Registers AgentMetricsSpanHandler and AgentMetricsEventHandler on the global LlamaIndex root dispatcher. Returns the span handler for flushing.
| Parameter | Default | Description |
|---|---|---|
agent_id |
"llamaindex-agent" |
Fallback label if the agent has no name attribute |
base_url |
"http://localhost:8099" |
AgentMetrics server address |
AgentMetricsSpanHandler
LlamaIndex BaseSpanHandler that tracks top-level agent/engine spans. Emits a run summary on span completion or error.
AgentMetricsEventHandler
LlamaIndex BaseEventHandler that accumulates token counts and tool calls from LLMChatEndEvent, LLMCompletionEndEvent, and AgentToolCallEvent.
.flush(timeout=10.0)
Blocks until all in-flight HTTP requests complete.
What gets tracked
Each top-level agent or query engine span emits one event to /v1/events:
| Field | Description |
|---|---|
status |
success or failed |
duration_ms |
Wall-clock span duration |
input_tokens / output_tokens |
Aggregated across all LLM calls |
cache_read_tokens / cache_write_tokens |
Cache token counts (Anthropic) |
llm_calls |
Number of LLM requests in the span |
tool_calls |
Tool call count from AgentToolCallEvent |
tool_names |
Set of tools invoked |
model |
Model name extracted from raw LLM response |
estimated_cost_usd |
Computed from token counts and model pricing |
error |
First 500 chars of the error message on failure |
The handler detects top-level spans by checking whether the span has no parent and whether the owning instance is an agent, engine, runner, or query object.
License
Metadata
Release files for agentmetrics-llamaindex 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| agentmetrics_llamaindex-0.2.0.tar.gz | 5.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| agentmetrics_llamaindex-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 11.6 kB
Release files / agentmetrics_llamaindex-0.2.0.tar.gz
| Download URL | agentmetrics_llamaindex-0.2.0.tar.gz |
|---|---|
| Size | 5.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
69a06e12e2bbc6d8c495e9db4dd655feb5f9c4766a13a56a5eefe5193476a7f9
|
|
BLAKE2b-256 checksum How to use checksums |
3b7c360bb2b896b45d80ccd455d0420a36a76bec827a07ba1bb76a139865a320
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.5
|
Release files / agentmetrics_llamaindex-0.2.0-py3-none-any.whl
| Download URL | agentmetrics_llamaindex-0.2.0-py3-none-any.whl |
|---|---|
| Size | 6.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f784f9599b3c8cb388f45a7a6fe35bf6c33b7f7db38af78605785c785d65dc26
|
|
BLAKE2b-256 checksum How to use checksums |
94a03a58efa6ebf05b105f4da7409f8f121da6c4dbd321dbf1a549d848aadb2e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.13.5
|