adk-finops
Universal FinOps cost, token usage, and grounding fee tracking for the Google Agent Development Kit (ADK) and LLM workflows.
Table of Contents
- Why adk-finops?
- Key Features
- Installation
- Quickstart with Google ADK
- Dual-Scope Telemetry: Turn vs. Session
- Budget Guards & Circuit Breakers
- Task Outcome & Wasted Spend Analytics (Success vs. Failure)
- Context Caching Savings Analytics (ROI Tracker)
- Multi-Agent Cost Attribution & Delegation Tracking
- Automated FinOps Optimization Advisor
- Rich Terminal Summary Box
- Decoupled Rate Cards (Custom & Enterprise Pricing)
- 1. Custom JSON Rate Card
- 2. Environment Variable
- 3. Remote URL / Enterprise Pricing Endpoint
- 4. Enterprise Negotiated Discounts
- 5. Programmatic Model & Tool Registration
- 6. Regional (
non_global) & Date-Tiered (2027) Pricing - 7. Dynamic Remote Rate Card Syncing & Google Cloud Pricing Extraction Framework
- Smart Tool Classification & Explicit Tool Billing (
@billable) - Gemini 2.5 & 3.x Thinking Tokens Billing
- Standalone Usage (Without ADK)
- Configuration Reference
- Built-In Model Rate Cards
- One-Line BigQuery Exporter
- Limitations & Roadmap (Next Release)
- License
Why adk-finops?
Building production AI agents with Google ADK involves multi-step tool-calling loops, reasoning models, and external search tools. However, monitoring real financial costs in production presents several challenges:
- Missing Cumulative Turn Tokens: In a tool-calling loop (e.g., Call 1 decides to call a tool $\rightarrow$ Tool runs $\rightarrow$ Call 2 generates final response), default event telemetry only displays the token count of Call 2, hiding the cost of earlier reasoning calls.
- Thinking / Reasoning Tokens: Models like Gemini 2.5 Pro separate
thoughtsTokenCountfromcandidatesTokenCount. If thinking tokens are not explicitly extracted, your token usage will not match your Google Cloud bill. - MCP vs. Paid Grounding: Generic keyword matching often misclassifies free local/MCP tools (like
search_documents) as paid Vertex AI Grounding queries ($35/1k or $2.50/1k). - Hardcoded Pricing: Model prices change frequently, and enterprises negotiate custom volume discounts. Hardcoded rates in application code require constant code updates.
adk-finops solves all of these problems with a single lightweight, decoupled plugin.
Key Features
- Native Google ADK Integration: Intercepts model and tool invocations via the ADK
BasePluginlifecycle (before_run,before_model,after_model,on_event,after_tool,after_run). - Dual-Scope Accounting: Simultaneously tracks metrics for both the active turn (all calls within a user message) and the cumulative session (entire conversation history).
- Pre-Flight & Post-Call Budget Guards: Estimate input tokens and projected cost in
before_model_callback(< 0.5ms) using deterministic provider token profiles (google,openai,anthropic,deepseek). Block massive 200K+ token dumps (max_prompt_tokens) at$0.00cloud cost or automatically downgrade models (gemini-2.5-pro➔gemini-2.5-flash) in-place before the request leaves the process. - Automated FinOps Optimization Advisor: Zero-LLM, deterministic rule engine that analyzes completed session telemetry in
< 1msand calculates concrete$and%savings across Context Caching Opportunities, Thinking Token Alerts (thinking_budget=0), and Model Right-Sizing (gemini-3.5-flash➔gemini-3.5-flash-lite,gemini-2.5-pro➔gemini-2.5-flash,gpt-4o➔gpt-4o-mini). Easily toggled on/off viaenable_optimization_advisor=True/False. - Task Outcome & Wasted Spend Analytics: Distinguishes productive spend (
status="success") from wasted capital burned on failed retry loops or runtime exceptions (status="failed","error"). Computes average spend per successful task vs. average wasted spend per failed loop, capital loss percentage, and auto-exports crashed sessions to BigQuery. - Context Caching Savings ROI: Demonstrates financial value by tracking gross cost (cost without caching) vs. actual net cost, reporting exact dollars and percentage saved (e.g. up to 90% savings via Gemini Context Caching).
- Decoupled Rate Cards: Pricing data is stored in clean JSON. Override rates via local file, remote URL, environment variable, or code without modifying the engine.
- Enterprise Volume Discounts: Configure global or provider-specific discount multipliers (e.g., 15% Google Cloud negotiated discount).
- Accurate Thinking Tokens: Automatically captures and bills
thoughts_token_countat the output rate while displaying thinking tokens separately in reports. - Explicit Tool Billing (
@billable) & Per-Tool Attribution (breakdown_by_tool): Decorate custom functions/classes with@billable(fee=...)or@billable(fee_fn=...)(with automaticcharge_on_error=Falseprotection), register rates viaCostTracker.register_tool_rate/tool_rates, and track per-agent tool call counts and fees (Agent -> Tool -> #Calls -> Cost) across the Rich terminal box, BigQuery, local logs, OpenTelemetry, and the web dashboard. - Real-Time UI Streaming: Streams cost metrics to
event.actions.state_delta["finops_cost"]for live updates in the ADK Web UI, with formatted terminal stdout logging. - Zero Heavy Dependencies: Pure Python standard library for the core tracker.
Installation
From PyPI
pip install adk-finops
With Rich Terminal Summary
pip install "adk-finops[rich]"
With BigQuery Exporter
pip install "adk-finops[bigquery]"
With Google ADK
pip install "adk-finops[adk]"
Full Enterprise Suite (ADK + Rich + BigQuery)
pip install "adk-finops[all]"
From GitHub (Direct Git Dependency)
pip install git+https://github.com/dmoonat/adk-finops.git
Or in requirements.txt:
adk-finops @ git+https://github.com/dmoonat/adk-finops.git@main
Quickstart with Google ADK
Add FinOpsCostPlugin to your App in agent.py:
from google.adk.agents import Agent
from google.adk.apps import App
from adk_finops import FinOpsCostPlugin
# 1. Define your agent
root_agent = Agent(
model="gemini-2.5-pro",
name="my_agent",
instruction="You are a helpful assistant.",
tools=[...]
)
# 2. Instantiate the FinOps plugin
finops_plugin = FinOpsCostPlugin(
name="finops_cost_tracker",
default_model="gemini-2.5-pro",
)
# 3. Attach plugin to your App
app = App(
name="my_agent",
root_agent=root_agent,
plugins=[finops_plugin],
)
⚠️ Note:
FinOpsCostPluginis a nativeApp/Runner-levelBasePlugin. Register it only viaApp(plugins=[finops_plugin])(orRunner(plugins=[finops_plugin])) — do not also passplugin.after_model_callbacktoAgent(after_model_callback=...), or ADK will invoke the callback twice.
Run your agent with adk web or adk run. Telemetry will log directly to the terminal and appear in the Web UI session state under finops_cost.
Dual-Scope Telemetry: Turn vs. Session
adk-finops resolves the mismatch between single-event inspection and session-level totals by reporting both scopes concurrently:
| Scope | Description | Matches in ADK Web UI |
|---|---|---|
turn |
Cumulative metrics for the current user interaction (Call 1 + Tool Execution + Call 2) | Full cost of the current turn |
session |
Cumulative metrics across all turns in the chat session (Turn 1 + Turn 2 + ...) | ADK "Usage Summary for Session" |
State Delta Example
{
"total_tokens": 18671,
"prompt_tokens": 16617,
"completion_tokens": 1110,
"thoughts_tokens": 944,
"total_cost_usd": 0.0135785,
"currency": "USD",
"turn": {
"total_calls": 2,
"total_tool_calls": 0,
"prompt_tokens": 14222,
"completion_tokens": 1079,
"thoughts_tokens": 908,
"total_tokens": 16209,
"llm_cost_usd": 0.0092341,
"tool_cost_usd": 0.0,
"total_cost_usd": 0.0092341,
"breakdown_by_model": {
"gemini-2.5-pro": {
"calls": 2,
"prompt_tokens": 14222,
"completion_tokens": 1079,
"thoughts_tokens": 908,
"total_cost_usd": 0.0092341
}
}
},
"session": {
"total_calls": 3,
"total_tool_calls": 0,
"prompt_tokens": 16617,
"completion_tokens": 1110,
"thoughts_tokens": 944,
"total_tokens": 18671,
"llm_cost_usd": 0.0135785,
"tool_cost_usd": 0.0,
"total_cost_usd": 0.0135785
}
}
Budget Guards & Circuit Breakers
Prevent runaway agent loops and surprise bills with both Pre-Flight (before_model_callback) and Post-Call (after_model_callback) budget enforcement. Set hard USD spending limits per session, per turn, or per sub-agent, plus an optional hard ceiling on input prompt tokens (max_prompt_tokens):
from adk_finops import FinOpsCostPlugin
finops_plugin = FinOpsCostPlugin(
default_model="gemini-2.5-pro",
budget_limit_usd=1.00, # Hard stop at $1.00 per chat session
turn_budget_limit_usd=0.25, # Maximum $0.25 on any single turn
max_prompt_tokens=200_000, # Pre-flight token cap (e.g. prevent 2x >200K tier charges)
preflight_budget_guard=True, # Estimate input tokens & cost BEFORE the network call (default: True)
on_budget_exceeded="halt", # Action: "halt", "warn", or "downgrade"
fallback_model="gemini-2.5-flash", # Target model when using "downgrade"
)
Pre-Flight Estimation (before_model_callback)
Before every outgoing LLM request, adk-finops estimates input tokens across system instructions, conversation history, tool schemas, function call/response payloads, and inline multimodal parts in < 0.5ms with zero network overhead (estimate_request_tokens(llm_request, model_name=...)):
- Automatically resolves the provider (
google,openai,anthropic,deepseek) fromllm_request.modeland applies provider-specific tokenization profiles (PROVIDER_TOKEN_PROFILES). - Calculates
projected_cost_usd = current_cost_usd + estimated_call_cost_usd(accounting for>200Kcontext tier pricing and cached tokens). - If
projected_cost_usdwould exceed your Agent, Turn, or Session budget (or ifestimated_prompt_tokens > max_prompt_tokens),before_model_callbackintervenes before the request is sent to the LLM provider — incurring $0.00 in cloud charges.
| Component | Google (gemini-*) |
OpenAI (gpt-4o, o1, o3) |
Anthropic (claude-*) |
DeepSeek (deepseek-*) |
|---|---|---|---|---|
| ASCII Chars / Token | 4.0 (SentencePiece 256k) |
4.4 (o200k_base) / 4.0 (cl100k_base) (or exact tiktoken if installed) |
3.6 (Claude 65k BPE vocab) |
3.7 (128k BPE vocab) |
| Turn Framing | +4 tokens / msg |
+4 tokens / msg |
+5 tokens / msg |
+4 tokens / msg |
| Tool Call / Resp Envelope | +8 tokens |
+12 tokens |
+16 tokens |
+10 tokens |
| Hidden Tool System Preamble | 0 tokens |
+16 tokens |
Dynamic by model & tool_choice:• Opus: 286 (auto/none) / 406 (forced)• Sonnet: 354 (auto/none) / 474 (forced)• Haiku: 264 (auto/none) / 340 (forced) |
+12 tokens |
| Per-Tool Schema Overhead | +36 tokens / tool |
+42 tokens / tool |
+45 tokens / tool |
+40 tokens / tool |
Images (image/*) |
258 tokens |
425 tokens (85 + 2×170 tiles) |
1,334 tokens ((1024×1024)/750) |
512 tokens |
PDFs (application/pdf) |
258 tokens / page |
800 tokens / page |
2,250 tokens / page (text + page image) |
512 tokens / page |
Customizing or Registering Provider Token Profiles
Just like pricing rate cards, you can register a new provider/model tokenization profile or override specific fields on existing providers programmatically, via FinOpsCostPlugin(token_profiles=..., token_profiles_path=...), or via ADK_FINOPS_TOKEN_PROFILES_PATH:
from adk_finops import CostTracker, register_token_profile, update_token_profile
# 1. Register a brand-new provider profile (inherits missing defaults from base_provider)
register_token_profile("mistral", {
"ascii_chars_per_token": 3.8,
"tool_System_preamble_tokens": 24,
"image_tokens": 512,
"pdf_page_tokens": 750,
})
# 2. Partially override an existing provider (merges nested dicts automatically)
update_token_profile("anthropic", {
"pdf_page_tokens": 2500,
"tool_System_preamble_tokens": {"opus_5.5_auto_or_none": 300},
..
})
# 3. Or via CostTracker / FinOpsCostPlugin
CostTracker.register_token_profile("meta-llama", {"ascii_chars_per_token": 3.9, ..})
Enforcement Modes
| Mode | Pre-Flight (before_model_callback) & Post-Call Behavior |
Best Used For |
|---|---|---|
"halt" (default) |
Short-circuits before_model_callback by returning a safe LlmResponse before the network request is sent ($0.00 cloud cost), or halts subsequent calls if breached mid-turn. |
Production safeguards, blocking massive 200K+ token dumps before they are billed. |
"downgrade" |
Mutates llm_request.model = fallback_model (e.g. gemini-2.5-pro ➔ gemini-2.5-flash) in-place before the request leaves the process so the current call immediately executes at the lower rate. |
Graceful service degradation with zero user downtime. |
"warn" |
Logs a [FinOps Pre-Flight Guard] Warning with projected cost and marks exceeded: True in session state (finops_cost.budget) without interrupting execution. |
Soft monitoring and alerting. |
Task Outcome & Wasted Spend Analytics (Success vs. Failure)
In production AI systems, a failed agent loop (e.g., an agent repeatedly failing schema validation across 3 retries, or crashing due to a downstream API outage after generating a complex plan) often burns 5x–20x more tokens than a successful task while delivering zero business value.
adk-finops tracks every LLM call in real time and classifies outcomes into Productive Spend vs. Wasted Capital, allowing teams to compare the average spend of a successful task against the wasted spend of failed loops.
Understanding status vs. is_failure
| Field | Type | Values | Purpose |
|---|---|---|---|
status |
str |
"success", "pending", "failed", "error", "aborted", "budget_exceeded" |
Operational Root Cause: Identifies how the task ended (e.g., validation loop exhausted "failed", downstream tool crash "error", or circuit breaker halt "budget_exceeded"). |
is_failure |
bool |
True or False |
Financial Bucket: Binary flag (True when status is failed, error, aborted, or budget_exceeded) used to separate Wasted Spend (True) from Productive Spend (False). |
1. Automatic Workflow State Detection (Google ADK)
If any ADK workflow node or critic sets ctx.state["failed"] = True and ctx.state["error_reason"] = "...", FinOpsCostPlugin automatically detects the failure in after_run_callback, tags the root cause, and marks the run's cost as wasted spend:
def strict_validator(node_input, ctx):
attempts = ctx.state.get("attempts", 0) + 1
ctx.state["attempts"] = attempts
if attempts >= 3:
ctx.state["failed"] = True
ctx.state["error_reason"] = "Exhausted 3 retry attempts without passing validation"
return Event(output="Aborted", actions=EventActions(route="abort"))
2. Explicit Recording & Crash Recovery (Auto-Export to BigQuery)
When an unhandled exception crashes an agent run, ADK skips after_run_callback. Calling finops_plugin.record_task_status() inside your except block ensures the wasted tokens are logged and automatically streamed to BigQuery:
from adk_finops import CostTracker, print_summary, print_task_efficiency_summary
try:
async for event in runner.run_async(user_id="u1", session_id=session_id, new_message=msg):
...
except Exception as e:
# Logs error outcome AND automatically flushes the crashed session to BigQuery
finops_plugin.record_task_status(
session_id=session_id,
status="error",
error=f"{type(e).__name__}: {str(e)}",
)
print_summary(session_id=session_id)
# Render comparative efficiency report across all tasks
print_task_efficiency_summary()
Comparative Task Efficiency Report
╭───────────── 📊 FinOps Task Efficiency & Wasted Spend Analysis ──────────────╮
│ │
│ Metric Successful Tasks Failed Loops (Wasted) Total / Impact │
│ ────────────────────────────────────────────────────────────────────────── │
│ Task Count 2 2 4 (50.0% fail rate) │
│ Total Spend $0.0021 $0.0080 $0.0101 (79.6% was) │
│ Avg Spend / Task $0.0010 $0.0040 Wasted is 3.9x avg │
│ Total Tokens 864 3,383 4,247 │
│ │
│ ⚠️ Capital Loss: $0.0080 (79.6% of total spend) was burned on uncompleted │
│ or failed tasks. │
│ 🛡️ Budget Context: Total spend is $0.0101 / $5.0000 (0.2% limit used; │
│ waste is 0.16% of budget limit). │
╰─────────────── adk-finops • Successful vs. Wasted Agent Loops ───────────────╯
Context Caching Savings Analytics (ROI Tracker)
Context caching can reduce prompt token costs by up to 90%. adk-finops automatically calculates Gross Cost (what you would have paid without caching), Actual Net Cost, and Total Savings:
{
"total_cost_usd": 0.00334,
"gross_cost_usd": 0.00550,
"savings_usd": 0.00216,
"savings_pct": 39.3,
"cached_tokens": 8000
}
Real-time stdout log highlight:
[FinOps LLM] turn=turn_1 model=gemini-2.5-flash tokens=11000 cost=$0.003340 | 💰 Saved $0.002160 (39.3%) via Context Caching
[FinOps Summary] Turn cost=$0.003340 | Session cost=$0.003340 | 💰 Total Saved: $0.002160 (39.3%) via Caching
Multi-Agent Cost Attribution & Delegation Tracking
In hierarchical multi-agent architectures (e.g., a supervisor delegating subtasks to a researcher and a coder), each agent makes distinct model calls, executes different tools, and consumes different context windows. Without granular attribution, teams cannot identify which sub-agent is driving 80% of costs or entering a costly reasoning loop.
adk-finops automatically attributes every LLM invocation, thinking token, context caching savings, and tool fee to the specific agent executing the task, with optional per-agent budget limits:
from adk_finops import FinOpsCostPlugin
finops_plugin = FinOpsCostPlugin(
default_model="gemini-2.5-pro",
budget_limit_usd=2.00, # Total session cap: $2.00
agent_budgets={
"researcher": 0.50, # Cap researcher at $0.50
"coder": 1.00, # Cap coder at $1.00
},
on_budget_exceeded="halt", # Halt if any agent breaches its limit
)
Hierarchical Root (Parent) Agent ➔ Sub-Agent Attribution
FinOpsCostPlugin automatically inspects the ADK agent tree (agent.root_agent, agent.parent_agent, and agent.sub_agents) during execution and tags every agent entry with:
root_agent_name: The top-level orchestrator / parent agent for the run (e.g.,"supervisor").parent_agent_name: The immediate parent agent (nullfor the root orchestrator,"supervisor"for delegated sub-agents).agent_role:"root_self"(the root orchestrator's own direct LLM/tool calls),"sub_agent"(a delegated specialist sub-agent), or"root_rollup"(the overall parent-level session rollup combining the root agent + all its sub-agents).
Every turn and session summary includes this hierarchy inside breakdown_by_agent:
{
"root_agent_name": "supervisor",
"breakdown_by_agent": {
"supervisor": {
"calls": 1,
"prompt_tokens": 1000,
"completion_tokens": 200,
"thoughts_tokens": 0,
"cached_tokens": 0,
"total_tokens": 1200,
"llm_cost_usd": 0.00225,
"tool_cost_usd": 0.0,
"total_cost_usd": 0.00225,
"savings_usd": 0.0,
"root_agent_name": "supervisor",
"parent_agent_name": null,
"agent_role": "root_self"
},
"researcher": {
"calls": 2,
"prompt_tokens": 4000,
"completion_tokens": 500,
"thoughts_tokens": 0,
"cached_tokens": 2000,
"total_tokens": 4500,
"llm_cost_usd": 0.00189,
"tool_cost_usd": 0.028,
"total_cost_usd": 0.02989,
"savings_usd": 0.00054,
"savings_pct": 22.2,
"root_agent_name": "supervisor",
"parent_agent_name": "supervisor",
"agent_role": "sub_agent"
},
"coder": {
"calls": 1,
"prompt_tokens": 8000,
"completion_tokens": 1500,
"thoughts_tokens": 300,
"cached_tokens": 0,
"total_tokens": 9800,
"llm_cost_usd": 0.01275,
"tool_cost_usd": 0.0,
"total_cost_usd": 0.01275,
"savings_usd": 0.0,
"root_agent_name": "supervisor",
"parent_agent_name": "supervisor",
"agent_role": "sub_agent",
"tools": {
"patent_search": {
"calls": 2,
"total_cost_usd": 0.03
}
}
}
},
"breakdown_by_tool": {
"researcher": {
"google_search": {
"calls": 2,
"total_cost_usd": 0.028
}
},
"coder": {
"patent_search": {
"calls": 2,
"total_cost_usd": 0.03
}
}
}
}
Real-Time Terminal Logs
[FinOps LLM] turn=turn_1 session=sess_1 agent=supervisor model=gemini-2.5-pro tokens=1200 cost=$0.002250
[FinOps Grounding] turn=turn_1 session=sess_1 agent=researcher tool=google_search fee=$0.014000
[FinOps Tool] turn=turn_1 session=sess_1 agent=coder tool=patent_search fee=$0.015000
[FinOps LLM] turn=turn_1 session=sess_1 agent=researcher model=gemini-2.5-flash tokens=4500 cost=$0.001890 | 💰 Saved $0.000540 (22.2%) via Context Caching
[FinOps Agents] supervisor: $0.0023 (1200 tok) | researcher: $0.0299 (4500 tok) | coder: $0.0428 (9800 tok)
[FinOps Tools] researcher -> google_search: 2 calls ($0.0280) | coder -> patent_search: 2 calls ($0.0300)
Automated FinOps Optimization Advisor
Beyond raw telemetry, adk-finops includes an Automated FinOps Optimization Advisor (src/adk_finops/advisor.py) that analyzes completed sessions in < 1ms with zero LLM calls and $0.00 overhead, computing concrete dollar and percentage savings directly from your active RateCardRegistry:
- ⚡ Context Caching Opportunity:
- Detects agents sending $\ge 2,000$ uncached prompt tokens across $\ge 2$ turns and calculates exact savings from enabling Context Caching (
cached_input_per_1mvs.input_per_1m).
- Detects agents sending $\ge 2,000$ uncached prompt tokens across $\ge 2$ turns and calculates exact savings from enabling Context Caching (
- 🧠 Thinking Token Alert:
- Detects agents where
thoughts_tokens >= 500account for $\ge 60%$ of total output token cost, recommendingthinking_budget=0(or a lowerThinkingConfigcap) for routing/classification steps.
- Detects agents where
- 🎯 Model Right-Sizing:
- Detects agents using flagship models (
gemini-3.5-pro$\rightarrow$gemini-3.5-flash$\rightarrow$gemini-3.5-flash-lite,gemini-2.5-pro$\rightarrow$gemini-2.5-flash,gpt-4o$\rightarrow$gpt-4o-mini,claude-3-5-sonnet$\rightarrow$claude-3-5-haiku) for short responses (< 200average output tokens) with0tool calls, calculating exact savings from switching to the lighter tier.
- Detects agents using flagship models (
Enabling or Disabling the Optimization Advisor
The advisor is enabled by default and can be toggled on or off in FinOpsCostPlugin:
from adk_finops import FinOpsCostPlugin
# Enabled by default (renders in Terminal Box & attaches to get_summary()['optimization_insights'])
finops_plugin = FinOpsCostPlugin(
enable_optimization_advisor=True,
)
# Disable the advisor if you only want raw telemetry
finops_plugin = FinOpsCostPlugin(
enable_optimization_advisor=False,
)
Or toggle globally via environment variable:
export ADK_FINOPS_OPTIMIZATION_ADVISOR="false"
Rich Terminal Summary Box
adk-finops includes an out-of-the-box, color-coded, border-styled terminal summary box. When running in a terminal, it provides instant financial visibility after every turn, displaying turn vs. session costs, context caching ROI, model breakdowns, sub-agent attributions, and per-tool cost breakdowns (Agent -> Tool -> Calls -> Tool Cost):
╭───────────────────────── 💸 ADK FinOps Cost Summary ─────────────────────────╮
│ │
│ Scope Calls Tokens LLM Cost Tool Fees Total Cost │
│ ────────────────────────────────────────────────────────────────── │
│ Current Turn 2 10,100 $0.0053 $0.0290 $0.0343 │
│ Session Total 2 10,100 $0.0053 $0.0290 $0.0343 │
│ │
│ 💰 Context Caching Savings: $0.0016 saved (4.5% reduction from $0.0359 gross)│
│ │
│ Model Calls Tokens (In/Out) Cost (USD) Savings │
│ ───────────────────────────────────────────────────────────────────────── │
│ gemini-2.5-pro 1 1,200 / 300 $0.0030 — │
│ gemini-2.5-flash 1 8,000 / 600 $0.0023 $0.0016 (41.5%) │
│ │
│ Agent Calls Tokens LLM Cost Tool Fees Total Cost │
│ ──────────────────────────────────────────────────────────────────── │
│ 🤖 researcher 4 14,400 $0.0078 $0.0290 $0.0368 │
│ 🤖 router_agent 1 5,900 $0.0130 $0.00 $0.0130 │
│ 🤖 formatter 1 1,610 $0.0024 $0.00 $0.0024 │
│ │
│ Agent Tool Calls Tool Cost │
│ ──────────────────────────────────────────────────────────────────── │
│ 🤖 researcher 🔧 patent_search 1 $0.0150 │
│ 🤖 researcher 🔧 google_search 1 $0.0140 │
│ │
│ 🛡️ Budget Guard: $0.0343 / $1.0000 (3.4% utilized) │
│ │
│ 💡 Optimization Insights (Est. Savings: $0.0148 | 63.6% cut available) │
│ ⚡ Context Caching Opportunity: Agent 'researcher' sent >12k uncached │
│ prompt tokens across 4 turns. Enabling Context Caching would save $0.0026 │
│ (90%). │
│ 🧠 Thinking Token Alert: Thinking tokens (4,200) were 82% of │
│ 'router_agent' output cost ($0.0105); consider setting thinking_budget=0. │
│ 🎯 Model Right-Sizing: Agent 'formatter' used gemini-2.5-pro for <150 │
│ output tokens with 0 tool calls; switching to gemini-2.5-flash saves 70% │
│ ($0.0017). │
│ ℹ️ Disclaimer: Insights are deterministic hints; validate against your │
│ use-case, eval data & business requirements. │
╰───────────────── adk-finops • Universal Token & Cost Engine ─────────────────╯
Terminal Box Features
- Rich 24-Bit Color Styling: When
richis installed (pip install "adk-finops[rich]"), it renders full color highlights, styled headers, and rounded boxes. - Pure-Python Unicode Fallback: If
richis not installed, it falls back seamlessly to an aligned pure-Python Unicode box drawing (╭─╮│╰─╯) with zero external dependencies. - Responsive 80-Column Layout: Designed to fit standard terminal windows without clipping or wrapping.
Automatic & On-Demand Integration
adk-finops supports both hands-off automated terminal reporting in Google ADK and explicit on-demand rendering for any Python workflow:
1. Automatic Integration (Google ADK)
When using FinOpsCostPlugin, the summary box is rendered automatically to stdout at the conclusion of every turn (after_run_callback):
from google.adk.agents import Agent
from google.adk.apps import App
from adk_finops import FinOpsCostPlugin
# Automatically renders the Rich box and Optimization Advisor at the end of each turn
finops_plugin = FinOpsCostPlugin(
default_model="gemini-2.5-pro",
render_terminal_box=True, # Enabled by default
enable_optimization_advisor=True, # Enabled by default
)
app = App(
name="support_agent",
root_agent=Agent(name="support_agent", model="gemini-2.5-pro"),
plugins=[finops_plugin],
)
2. On-Demand Integration (Standalone & Custom Workflows)
You can trigger terminal rendering programmatically at any time across your batch scripts, agent loops, or background tasks:
from adk_finops import CostTracker, format_summary_box, print_summary
# 1. Print current turn/session summary box directly to terminal stdout
print_summary()
# Or specify an explicit turn and session ID:
print_summary(run_id="turn_123", session_id="session_abc")
# 2. Or call directly from CostTracker
CostTracker.print_summary()
# 3. Export formatted box string (ANSI or plain text) for custom loggers or webhooks (Slack/Discord)
summary_data = CostTracker.get_summary()
box_text = format_summary_box(summary_data)
logger.info("\n" + box_text)
Decoupled Rate Cards (Custom & Enterprise Pricing)
Pricing is completely decoupled from the tracking engine. You can configure rates via:
1. Custom JSON Rate Card
Create a my_rates.json file:
{
"models": {
"custom-fine-tuned-model": {
"provider": "google",
"input_per_1m": 0.50,
"output_per_1m": 2.00,
"cached_input_per_1m": 0.05
}
},
"tools": {
"paid_external_api": 0.01
}
}
Pass the file path when initializing the plugin:
finops_plugin = FinOpsCostPlugin(rate_card_path="my_rates.json")
2. Environment Variable
Set the rate card path globally without touching any code:
export ADK_FINOPS_RATE_CARD_PATH="/etc/finops/rates.json"
3. Remote URL / Enterprise Pricing Endpoint
Fetch central rate cards from an internal endpoint or cloud storage bucket (GCS/S3):
from adk_finops import CostTracker
CostTracker.load_rate_card_url("https://internal.corp.com/finops/rates.json")
4. Enterprise Negotiated Discounts
Apply your organization's contracted cloud discounts:
# Apply a 15% discount across all models and tools
finops_plugin = FinOpsCostPlugin(discount_percent=15.0)
# Or configure dynamically via CostTracker:
CostTracker.set_discount(15.0) # Global 15% discount
CostTracker.set_discount(20.0, provider="google") # 20% discount on Google Cloud models
Or set via environment variable:
export ADK_FINOPS_DISCOUNT_PERCENT="15.0"
5. Programmatic Model & Tool Registration
If a model is not present in the bundled default_rates.json, adk-finops applies the default fallback rate ($0.30 / $2.50 / $0.03 per 1M tokens), sets "is_fallback_rate": True on the recorded usage/summary, and logs a one-time warning prompting you to register the model's exact pricing via CostTracker.register_rate_card:
from adk_finops import CostTracker
CostTracker.register_rate_card("claude-opus-5", {
"provider": "anthropic",
"input_per_1m": 5.00,
"output_per_1m": 25.00,
"cached_input_per_1m": 0.50,
})
CostTracker.register_tool_rate("internal_vector_db", 0.0005)
6. Regional (non_global) & Date-Tiered (2027) Pricing
adk-finops automatically detects your Google Cloud region and applies non_global regional rates (e.g., us-central1, europe-west1) as well as promotional-to-standard pricing transitions (standard_pricing_2027 starting Jan 1, 2027):
- Region Resolution Precedence:
- Explicit
regionpassed toFinOpsCostPlugin(region="us-central1")orCostTracker.set_region("us-central1") GOOGLE_CLOUD_LOCATIONenvironment variable (standard Google ADK.envconfiguration)ADK_FINOPS_REGIONenvironment variable- Defaults to
"global"
- Explicit
from adk_finops import CostTracker, FinOpsCostPlugin
# Automatically uses GOOGLE_CLOUD_LOCATION from .env if set, or specify explicitly:
finops_plugin = FinOpsCostPlugin(
region="us-central1", # Applies non_global rates (+10% regional pricing where applicable)
effective_date="2027-01-01", # Optional: simulate or enforce 2027 standard pricing
)
7. Dynamic Remote Rate Card Syncing & Google Cloud Pricing Extraction Framework
To ensure your agents always calculate costs against the newest model pricing without waiting for a package upgrade, adk-finops provides opt-in Dynamic Remote Rate Card Syncing (with a 24-hour local disk cache at ~/.cache/adk-finops/remote_rates.json and graceful offline fallback) as well as a Google Cloud & Multi-Provider Pricing Extraction Framework (src/adk_finops/pricing_extractor.py).
Canonical Pricing Sources (source_urls & sources in default_rates.json)
Every rate card generated by adk-finops records all upstream sources queried in its top-level "source_urls" and "sources" metadata:
| Source Name | Canonical URL | Scope & Coverage |
|---|---|---|
gemini_enterprise_pricing |
https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing |
Google Cloud Gemini Enterprise Agent Platform (global, non_global regional +10%, standard_pricing_2027, >200K context tiers, and Grounding Tools) |
litellm_registry |
https://raw.githubusercontent.com/BerriAI/litellm/main/model_prices_and_context_window.json |
Multi-Provider (openai, anthropic, deepseek, and cross-validation) |
gcp_cloud_billing_catalog_api (optional) |
https://cloudbilling.googleapis.com/v1/services/C7E2-9256-1C43/skus |
Official Google Cloud Billing Catalog REST API (when GOOGLE_CLOUD_BILLING_API_KEY or --gcp-api-key is supplied; sent via X-Goog-Api-Key header) |
API Key Reference (Where to Get & How to Use Each Key)
adk-finops uses two optional keys for different purposes — neither is required for standard local cost tracking:
| Key / Environment Variable | CLI Flag / Python Arg | Where to Get It | Purpose & How It Works |
|---|---|---|---|
GOOGLE_CLOUD_BILLING_API_KEY (optional) |
--gcp-api-key / gcp_billing_api_key |
Issued by Google Cloud: Google Cloud Console $\rightarrow$ APIs & Services $\rightarrow$ Credentials $\rightarrow$ Create Credentials $\rightarrow$ API key (enable Cloud Billing API). | Used only by adk-finops extract-pricing to query Google's cloudbilling.googleapis.com REST API. Sent via the X-Goog-Api-Key HTTP header. |
ADK_FINOPS_DASHBOARD_API_KEY (optional) |
--api-key / dashboard_api_key |
Pre-Shared Secret (PSK) you create: Because you run the adk-finops dashboard FastAPI server, you create any random token yourself (e.g. python3 -c "import secrets; print(secrets.token_urlsafe(24))"). |
You set the same secret string in the environment (.env, Cloud Run env var, or Secret Manager) of both the Dashboard Server (which verifies it) and your Agents (HTTPExporter, which sends it in the X-FinOps-Key header). |
Note:
pricing_extractor.pyseeds Google Cloud model definitions directly fromdefault_rates.jsonvia_load_google_baseline_rates()so curatednon_global,standard_pricing_2027, and_gt_200kfields are never overwritten or duplicated.
Runtime Remote Rate Card Syncing (FinOpsCostPlugin & CostTracker)
from adk_finops import CostTracker, FinOpsCostPlugin
# Option A: Enable via FinOpsCostPlugin (or set ADK_FINOPS_SYNC_REMOTE_RATES=1 in .env)
finops_plugin = FinOpsCostPlugin(
sync_remote_rates=True, # Syncs latest default_rates.json (cached locally for 24h)
remote_cache_ttl_seconds=86400, # Optional custom cache TTL
)
# Option B: Programmatic sync on demand
status = CostTracker.sync_remote_rate_card(force=True)
Running the Pricing Extraction Framework (4 Ways)
-
Via the
adk-finopsCLI (extract-pricing&sync-rates):# 1. Sync local rate card cache (~/.cache/adk-finops/remote_rates.json) from canonical remote repo adk-finops sync-rates --force # 2. Dry-run pricing extraction report (prints Rich table comparing live sources vs default_rates.json) adk-finops extract-pricing # 3. Export merged rate card (including newly discovered models like gpt-4.1, o3, claude-4) to a custom file adk-finops extract-pricing --include-new-models --output ./custom_rates.json # 4. Update bundled src/adk_finops/rates/default_rates.json in-place adk-finops extract-pricing --update-default --include-new-models
-
Via the Python Module (
python -m adk_finops.pricing_extractor):python -m adk_finops.pricing_extractor --include-new-models --output ./custom_rates.json
-
Programmatically in Python (
extract_and_sync_pricing):from adk_finops import extract_and_sync_pricing report = extract_and_sync_pricing( update_default_rates=False, output_path="./custom_rates.json", include_new_models=True, quiet=False, ) print("Sources synced:", report.sources_succeeded) print("Added models:", list(report.added_models.keys()))
-
Automated Weekly GitHub Actions Workflow (
.github/workflows/sync_pricing.yml): Runs every Monday at06:00 UTC(or manually viaworkflow_dispatchwithinclude_new_modelsenabled by default) to extract live rates and commit any changes tosrc/adk_finops/rates/default_rates.json.
Smart Tool Classification & Explicit Tool Billing (@billable)
adk-finops distinguishes between paid cloud grounding services ([FinOps Grounding]), explicitly billed custom/MCP tools ([FinOps Tool]), and free tools:
- Explicit
@billableDecorator (from adk_finops import billable) & Tool Rate Registration: Attach fixed per-call fees or dynamic per-result fees (fee_fn) directly to Python tool functions or custom ADKBaseToolclasses, with built-in error protection (charge_on_error=Falseby default so failed API calls return$0.00fee), or register tool pricing viaFinOpsCostPlugin(tool_rates=...)andCostTracker.register_tool_rate(...):from adk_finops import CostTracker, FinOpsCostPlugin, billable # 1. Fixed per-call fee ($0.015 per invocation) @billable(fee=0.015, provider="serpapi") def patent_search(query: str) -> dict: """Searches external patent database.""" return {"hits": 10} # 2. Dynamic per-unit fee ($0.002 per page OCR'd; $0.00 if result has {"status": "error"}) @billable( fee_fn=lambda args, res: 0.002 * len(res.get("pages", [])), charge_on_error=False, ) def ocr_document(gcs_uri: str) -> dict: """Extracts text from PDF pages via paid OCR service.""" return {"pages": ["page1", "page2", "page3"]} # 3. Global programmatic tool rate registration CostTracker.register_tool_rate("enterprise_risk_score_api", 0.025) # 4. Plugin-level tool_rates (ideal for paid third-party MCP tools) finops_plugin = FinOpsCostPlugin( default_model="gemini-2.5-pro", tool_rates={"bloomberg_mcp_terminal": 0.05}, )
- Per-Agent Tool Cost Attribution (
breakdown_by_tool): Every turn and session tracks{agent_name: {tool_name: {"calls": count, "total_cost_usd": fee}}}, rendering a dedicatedAgent | Tool | Calls | Tool Costtable in the Rich Terminal Summary Box, logging[FinOps Tools]at turn end, and exportingbreakdown_by_toolto JSONL, CSV, HTTP, BigQuery, OpenTelemetry, and the Web Dashboard. - Google Search Grounding & Web Grounding: Tools named
google_search,GoogleSearchTool,EnterpriseWebSearchTool, or carryinggoogle_searchconfig attributes incur $0.014 / query ($14.00 per 1,000 queries; first 5,000 queries/month free) and log as[FinOps Grounding]. - Google Maps Grounding: Incurs $0.014 / query ($14.00 per 1,000 queries; first 5,000 queries/month free).
- Grounding with your data (Vertex AI Search / Datastores): Tools matching
vertex_search,vertex_ai_search,GroundingTool, or carryingvertex_ai_search/data_store_idconfig attributes incur $0.0025 / prompt ($2.50 per 1,000 prompts) and log as[FinOps Grounding]. - Free MCP, Database & Local Function Tools: Un-decorated
McpToolset/MCPToolendpoints, BigQuery toolsets, and local Python functions incur $0.00 tool fee.
Gemini 2.5 & 3.x Thinking Tokens Billing
Gemini 2.5 and Gemini 3.x models separate reasoning/thinking tokens into thoughtsTokenCount.
Per Google Cloud pricing rules:
Total Billable Output Tokens =
candidates_token_count+thoughts_token_count
adk-finops automatically incorporates thinking tokens into the output token rate while reporting them as a separate field in breakdown_by_model so you can monitor your reasoning overhead.
Standalone Usage (Without ADK)
adk-finops is designed to be used in any Python application — FastAPI/Flask backends, Celery/Ray background pipelines, LangChain/LlamaIndex workflows, or raw Google GenAI SDK scripts — without requiring Google ADK.
from adk_finops import CostTracker, BigQueryExporter
# Wrap any execution block with automatic lifecycle cleanup
with CostTracker.track_run("request_123"):
# Record Gemini call with thoughts/reasoning tokens
CostTracker.record_usage(
run_id="request_123",
model_name="gemini-3.7-flash",
prompt_tokens=1500,
completion_tokens=250,
thoughts_tokens=100,
cached_tokens=0,
)
# Record grounding or tool fee ($0.014 / query)
CostTracker.record_tool_call(
run_id="request_123",
tool_name="google_search",
)
# Print color-coded terminal summary box
CostTracker.print_summary("request_123")
# Get structured metrics dictionary
summary = CostTracker.get_summary("request_123")
print(f"Total Cost: ${summary['total_cost_usd']:.6f}")
print(f"Total Tokens: {summary['total_tokens']}")
# (Optional) Export to BigQuery in 1 line
# exporter = BigQueryExporter("my-project.finops.agent_costs")
# exporter.export_summary(summary)
Configuration Reference
Environment Variables
| Variable | Type | Description |
|---|---|---|
GOOGLE_CLOUD_LOCATION |
str |
Primary ADK environment variable for region detection (e.g., "global", "us-central1"). Non-global regions automatically apply non_global rates. |
ADK_FINOPS_REGION |
str |
Fallback pricing region if GOOGLE_CLOUD_LOCATION is not set (defaults to "global"). |
ADK_FINOPS_EFFECTIVE_DATE |
str |
ISO date (YYYY-MM-DD) to evaluate date-tiered pricing such as standard_pricing_2027 (defaults to today's date). |
ADK_FINOPS_RATE_CARD_PATH |
str |
Absolute or relative path to a custom JSON rate card file. |
ADK_FINOPS_DISCOUNT_PERCENT |
float |
Global enterprise discount percentage (e.g., 15.0 for 15%). |
ADK_FINOPS_MAX_PROMPT_TOKENS |
int |
Hard pre-flight ceiling on estimated input prompt tokens per LLM call (e.g., 200000). |
ADK_FINOPS_PREFLIGHT_GUARD |
bool |
Enable or disable pre-flight token & budget estimation in before_model_callback ("true" or "false", default "true"). |
ADK_FINOPS_OPTIMIZATION_ADVISOR |
bool |
Toggle the Automated FinOps Optimization Advisor ("true" or "false"). |
Plugin Initialization Parameters
FinOpsCostPlugin(
name: str = "finops_cost_tracker",
default_model: str = "gemini-2.5-flash",
rate_card_path: str | Path | None = None,
rate_card: dict[str, Any] | None = None,
tool_rates: dict[str, float] | None = None,
discount_percent: float | None = None,
region: str | None = None,
effective_date: str | date | None = None,
budget_limit_usd: float | None = None,
turn_budget_limit_usd: float | None = None,
agent_budgets: dict[str, float] | None = None,
max_prompt_tokens: int | None = None,
preflight_budget_guard: bool = True,
on_budget_exceeded: str = "halt", # "halt", "warn", or "downgrade"
fallback_model: str = "gemini-2.5-flash",
render_terminal_box: bool = True,
enable_optimization_advisor: bool = True,
)
| Parameter | Type | Default | Description |
|---|---|---|---|
default_model |
str |
"gemini-2.5-flash" |
Fallback model name if not reported by the LLM response. |
rate_card_path |
str | Path |
None |
Path to custom rate card JSON file. |
tool_rates |
dict[str, float] |
None |
Custom per-call tool fees in USD (e.g. {"bloomberg_mcp": 0.05}). |
discount_percent |
float |
None |
Enterprise volume discount percentage (e.g. 15.0 for 15%). |
region |
str | None |
None |
Pricing region override. When None, auto-detects from GOOGLE_CLOUD_LOCATION $\rightarrow$ ADK_FINOPS_REGION $\rightarrow$ "global". |
effective_date |
str | date |
None |
Optional ISO date ("YYYY-MM-DD") for date-tiered pricing (e.g. "2027-01-01" for 2027 standard rates). |
budget_limit_usd |
float |
None |
Maximum cumulative spending limit in USD for the entire chat session. |
turn_budget_limit_usd |
float |
None |
Maximum spending limit in USD for any single user turn. |
agent_budgets |
dict[str, float] |
None |
Per-agent spending caps in USD (e.g. {"researcher": 0.50, "coder": 1.00}). |
max_prompt_tokens |
int |
None |
Optional hard cap on estimated input prompt tokens per LLM call (e.g. 200_000 to block >200K tier charges). |
preflight_budget_guard |
bool |
True |
Estimates input tokens and projected cost in before_model_callback before sending the LLM request. |
on_budget_exceeded |
str |
"halt" |
Action on budget breach: "halt" (short-circuit/block), "warn", or "downgrade" (in-place model switch). |
fallback_model |
str |
"gemini-2.5-flash" |
Target model when using "downgrade" mode. |
render_terminal_box |
bool |
True |
Renders a beautiful color-coded summary box to stdout at the end of each turn. |
enable_optimization_advisor |
bool |
True |
Enables deterministic FinOps Optimization Insights in the summary box and get_summary(). |
Built-In Model Rate Cards
The bundled default_rates.json contains official public pricing (USD per 1M tokens, sourced from Google Cloud Generative AI Pricing):
| Model | Provider | Input / 1M (<=200k) |
Output / 1M (<=200k) |
Cached / 1M (<=200k) |
Context >200k (_gt_200k) / Regional (non_global) / 2027 Notes |
|---|---|---|---|---|---|
gemini-3.1-pro-preview |
$2.00 | $12.00 | $0.20 | >200k: In $4.00, Out $18.00, Cached $0.40 |
|
gemini-3.8-flash-cyber |
$1.50 | $7.50 | $0.15 | Non-global: In $1.65, Out $8.25, Cached $0.165 | |
gemini-3.8-flash |
$0.75 | $3.75 | $0.075 | Non-global: $0.825 / $4.125 • 2027 Standard: $1.50 / $7.50 (Non-global: $1.65 / $8.25) | |
gemini-3.7-flash |
$0.75 | $3.75 | $0.075 | Non-global: $0.825 / $4.125 • 2027 Standard: $1.50 / $7.50 (Non-global: $1.65 / $8.25) | |
gemini-3.6-flash |
$0.75 | $3.75 | $0.075 | Non-global: $0.825 / $4.125 • 2027 Standard: $1.50 / $7.50 (Non-global: $1.65 / $8.25) | |
gemini-3.5-flash |
$1.50 | $9.00 | $0.15 | Non-global: In $1.65, Out $9.90, Cached $0.165 | |
gemini-3.5-flash-lite |
$0.30 | $2.50 | $0.03 | Non-global: In $0.33, Out $2.75, Cached $0.033 | |
gemini-3.1-flash-lite |
$0.25 | $1.50 | $0.025 | Non-global: In $0.275, Out $1.65, Cached $0.0275 | |
gemini-3-flash-preview |
$0.50 | $3.00 | $0.05 | — | |
gemini-3-pro-image |
$2.00 | $12.00 | $0.20 | >200k: In $4.00, Out $18.00, Cached $0.40 |
|
gemini-3.1-flash-image |
$0.50 | $3.00 | $0.05 | Nano Banana 2 image generation | |
gemini-3.1-flash-lite-image |
$0.25 | $1.50 | $0.025 | Nano Banana Lite image generation | |
gemini-2.5-pro |
$1.25 | $10.00 | $0.125 | >200k: In $2.50, Out $15.00, Cached $0.25 |
|
gemini-2.5-pro-computer-use-preview |
$1.25 | $10.00 | $0.125 | >200k: In $2.50, Out $15.00, Cached $0.25 |
|
gemini-2.5-flash |
$0.30 | $2.50 | $0.03 | >200k: In $0.30, Out $2.50, Cached $0.03 |
|
gemini-2.5-flash-lite |
$0.10 | $0.40 | $0.01 | >200k: In $0.10, Out $0.40, Cached $0.01 |
|
gemini-2.5-flash-image |
$0.30 | $2.50 | $0.03 | — | |
gemini-2.5-flash-live |
$0.50 | $2.00 | $0.05 | Live API text rates | |
gemini-2.0-flash |
$0.15 | $0.60 | $0.0375 | — | |
codemender |
$0.75 | $3.75 | $0.075 | 2027 Standard: In $1.50, Out $7.50, Cached $0.15 | |
gpt-4o |
OpenAI | $2.50 | $10.00 | $1.25 | — |
gpt-4o-mini |
OpenAI | $0.15 | $0.60 | $0.075 | — |
o1 |
OpenAI | $15.00 | $60.00 | $7.50 | — |
o3-mini |
OpenAI | $1.10 | $4.40 | $0.55 | — |
claude-3-7-sonnet |
Anthropic | $3.00 | $15.00 | $0.30 | — |
claude-3-5-sonnet |
Anthropic | $3.00 | $15.00 | $0.30 | — |
claude-3-5-haiku |
Anthropic | $0.80 | $4.00 | $0.08 | — |
deepseek-v3 |
DeepSeek | $0.14 | $0.28 | $0.014 | — |
deepseek-r1 |
DeepSeek | $0.55 | $2.19 | $0.14 | — |
One-Line BigQuery Exporter
Stream turn, session, model, and sub-agent FinOps telemetry directly into a Google BigQuery dataset with a single configuration parameter:
from adk_finops import FinOpsCostPlugin
finops_plugin = FinOpsCostPlugin(
default_model="gemini-2.5-flash",
bigquery_table="my-gcp-project.finops.agent_costs",
bigquery_export_scope="both", # "session" (default), "turn", or "both"
bigquery_tags={"env": "production", "service": "support-agent"},
)
Zero-Boilerplate Schema Management
When bigquery_table is specified, adk-finops automatically inspects and provisions the dataset and table with enterprise best practices:
- Partitioning: Day-partitioned on
timestampto optimize query performance and reduce scan costs. - Clustering: Clustered by
[session_id, agent_name, model_name]for sub-second filtering in Looker Studio and BI tools. - Automatic Schema Evolution: If the target BigQuery table already exists from an earlier version of
adk-finops,ensure_table_exists()automatically adds any newly introducedNULLABLEcolumns (such asbreakdown_by_tool) in-place with zero downtime. - Non-Blocking Background Streaming: Ingestion runs asynchronously in a background thread pool, adding zero latency to agent responses.
Table Schema Reference
| Field Name | Type | Description |
|---|---|---|
timestamp |
TIMESTAMP |
Event timestamp in UTC (Partition Key) |
session_id |
STRING |
ADK Session ID (Clustering Key #1) |
turn_id |
STRING |
ADK Turn / Invocation ID |
scope |
STRING |
Record scope: "turn" or "session" |
agent_name |
STRING |
Attributed Agent name (Clustering Key #2) |
model_name |
STRING |
Model name / version (Clustering Key #3) |
prompt_tokens |
INTEGER |
Input prompt token count |
completion_tokens |
INTEGER |
Output candidate token count |
thoughts_tokens |
INTEGER |
Gemini 2.5 thinking token count |
cached_tokens |
INTEGER |
Context cached token count |
total_tokens |
INTEGER |
Total billable tokens |
llm_cost_usd |
FLOAT |
Net LLM API cost in USD |
tool_cost_usd |
FLOAT |
Search, Grounding & @billable tool fees in USD |
total_cost_usd |
FLOAT |
Total net cost in USD |
gross_cost_usd |
FLOAT |
Gross cost before caching discount |
savings_usd |
FLOAT |
Dollars saved via context caching |
savings_pct |
FLOAT |
Percentage saved via context caching |
tool_calls_count |
INTEGER |
Number of billable grounding/tool calls |
budget_limit_usd |
FLOAT |
Configured budget threshold |
budget_utilization_pct |
FLOAT |
Budget utilization percentage |
budget_exceeded |
BOOLEAN |
Whether budget guard was tripped |
breakdown_by_agent |
JSON |
Multi-agent attribution snapshot |
breakdown_by_model |
JSON |
Model distribution snapshot |
breakdown_by_tool |
JSON |
Per-agent tool call counts & fee attribution ({agent: {tool: {calls, total_cost_usd}}}) |
tags |
JSON |
User-provided tags (e.g. env, tenant_id) |
Understanding Streamed Rows & Dimensions (Rollup vs. Attributed Rows)
When streaming telemetry to BigQuery, adk-finops records both overall aggregates (for high-level session/turn reporting) and attributed breakdown rows (for granular drill-downs by agent or model).
Why are agent_name or model_name NULL in some rows?
In data warehousing and BI rollup patterns, NULL is assigned to dimension columns in summary/rollup rows to distinguish between overall totals and specific entity breakdowns:
- Overall Aggregate Rows (
agent_name IS NULL):- Emitted once per turn or session representing the cumulative total across all agents and models.
agent_nameisNULL(can be displayed as'OVERALL'or'TOTAL'viaCOALESCE(agent_name, 'OVERALL')).- If multiple models were used in the run,
model_nameisNULL(the complete distribution is preserved in thebreakdown_by_modelJSON column). If only a single model was used,model_nameis set to that model.
- Attributed Agent Rows (
agent_name IS NOT NULL):- Emitted for each sub-agent participating in the turn/session (
agent_name = 'research_agent', etc.). model_namecontains the primary model invoked by that agent (e.g.gemini-2.5-pro,gemini-2.5-flash).
- Emitted for each sub-agent participating in the turn/session (
Querying Best Practices (Avoiding Double-Counting)
Because both aggregate rows and attributed breakdown rows coexist in the same table, write your SQL queries according to the level of granularity you need:
-
To query overall totals (e.g. total spend per session):
SELECT session_id, total_cost_usd, total_tokens FROM `my-gcp-project.finops.agent_costs` WHERE scope = 'session' AND agent_name IS NULL;
-
To query per-agent breakdown (without double-counting with the aggregate):
SELECT agent_name, SUM(total_cost_usd) AS agent_spend FROM `my-gcp-project.finops.agent_costs` WHERE scope = 'session' AND agent_name IS NOT NULL GROUP BY 1;
-
To query by model:
SELECT model_name, SUM(total_cost_usd) AS model_spend FROM `my-gcp-project.finops.agent_costs` WHERE scope = 'session' AND model_name IS NOT NULL GROUP BY 1;
Near-Live FinOps Web Dashboard (adk-finops dashboard)
adk-finops includes a built-in, zero-extra-dependency FastAPI + Chart.js Near-Live Web Dashboard that auto-refreshes every 2 seconds, aggregating:
- Live In-Memory
CostTrackerState: Watch tokens and spend accumulate in real time while an agent is mid-execution. - Local Timestamped
.jsonl&.csvLogs: Automatically scanslogs/<YYYYMMDD_HHMMSS>/*.jsonland*.csv. - Remote HTTP Push (
POST /api/ingest): Receives telemetry pushed over HTTP from remote agent containers (HTTPExporter/dashboard_endpoint) and persists it to<log_dir>/ingested_costs.jsonl. - Optional BigQuery Live Sync: Toggle the Include BigQuery switch in the UI (
15scache TTL) to merge cloud warehouse records (ADK_FINOPS_BIGQUERY_TABLE).
Hierarchical Root (Parent) Agent ➔ Sub-Agent Drilldown & Cascading Filters
- 5 Cascading Hierarchy Filters:
1. 👑 Root (Parent) Agent: Filter by a top-level orchestrator/parent agent (e.g.,👑 coordinator_agent). Selecting a Root Agent automatically computes the Overall Parent-Level Spend & Tokens (∑ Parent + All Sub-Agents) in the KPI cards while displaying all of its sub-agents in the breakdown charts and hierarchy tree.2. ↳ Sub-Agent Drilldown: Dynamically cascades to list only the children belonging to the selected Root Agent (∑ Overall Parent Total,👑 Root Orchestrator Direct Only, or individual↳ 🤖 Sub-Agents).3. Session Filter (Scoped): Automatically scopes the session dropdown so it only lists sessions belonging to the selected Root Agent (and Sub-Agent).4. Model Filter: Scoped to the models invoked by the selected Root / Sub-Agent.5. Task Outcome: Filter between✅ Effective Spend Only (Success)and🔥 Wasted Spend Only (Failed / Error / Budget).
👑 Root (Parent) Agent ➔ Sub-Agents Hierarchy RollupExplorer: Interactive parent-to-child cards showing each Root Agent's overall parent rollup (∑ Overall Parent SpendandOverall Parent Tokens) alongside each child's spend, token breakdown (In / Out / Think), model, tool badges (🔧 <tool> ×N ($cost)), and percentage share of parent spend with click-to-filter support.🔧 Tool Cost Attribution (Agent ➔ Tool ➔ #Calls ➔ Cost)Panel & Live Telemetry Table: Dedicated tool cost breakdown table showingAgent | Tool | # Calls | Avg Cost / Call | Total Tool Costplus aTools (Calls / Fee)column in the Live Telemetry Table.- 5 Executive KPI Cards & 3 Interactive Charts: Total Spend ($ with active scope badge), Effective Spend ($), Wasted Spend ($ & %), Context Caching Savings ($), Total Tokens (
In / Out / Think), Spend Efficiency Doughnut, Sub-Agent Stacked Cost Bar (LLM vs Tool/Grounding), and Per-Model Token Composition.
Developer Mode: Embedded Background Server in FinOpsCostPlugin
Start the live dashboard automatically in a background daemon thread inside your agent process:
from adk_finops import FinOpsCostPlugin
finops_plugin = FinOpsCostPlugin(
enable_dashboard=True, # or set ADK_FINOPS_ENABLE_DASHBOARD=true
dashboard_port=8088, # default: 8088 (auto-selects next free port if busy)
jsonl_path="logs/finops_costs.jsonl", # saves into logs/<YYYYMMDD_HHMMSS>/finops_costs.jsonl
)
Enterprise & Org Admin Mode: Standalone Centralized Dashboard (3 Patterns)
An organization admin can run adk-finops dashboard as a single centralized FinOps control plane (with zero agents running inside the dashboard process) to monitor dozens of independent root agents and sub-agent teams across an organization:
Pattern A: Central Cloud Warehouse Mode (BigQuery Only)
All agents across the org stream to a shared BigQuery table (ADK_FINOPS_BIGQUERY_TABLE="org-project.finops.agent_costs"). The admin runs the standalone dashboard pointing strictly to BigQuery:
adk-finops dashboard \
--bigquery-table org-project.finops.agent_costs \
--log-dir "" \
--host 0.0.0.0 \
--port 8088
Pattern B: Multi-Project Shared Directory / Volume Mode (--log-dir)
Pass comma-separated directories to watch timestamped .jsonl / .csv files across multiple agent repositories or shared volumes simultaneously:
adk-finops dashboard \
--log-dir "/srv/agents/retail_bot/logs,/srv/agents/finance_bot/logs" \
--port 8088
Pattern C: Direct HTTP Push Mode (dashboard_endpoint $\rightarrow$ POST /api/ingest)
When an organization does not use BigQuery and agents run in isolated containers/VMs without a shared disk, agents push telemetry over HTTP via HTTPExporter to the central dashboard's POST /api/ingest endpoint.
Option 1: Local Dev Mode (No API Key Required)
If neither --api-key nor ADK_FINOPS_DASHBOARD_API_KEY is configured on the dashboard server, /api/ingest accepts local telemetry pushes without requiring a token:
- Start the central dashboard server:
adk-finops dashboard --host 127.0.0.1 --port 8088 --log-dir central_logs
- Configure
dashboard_endpointon your agent:from adk_finops import FinOpsCostPlugin finops_plugin = FinOpsCostPlugin( dashboard_endpoint="http://127.0.0.1:8088", # Pushes via HTTPExporter -> POST /api/ingest export_tags={"team": "payments", "service": "refund_agent"}, )
Option 2: Authenticated Mode (Pre-Shared Secret ADK_FINOPS_DASHBOARD_API_KEY)
Because adk-finops dashboard is a self-hosted FastAPI server (not a third-party SaaS), you control the password/token that protects its POST /api/ingest endpoint. It works like a standard Pre-Shared Key (PSK) or webhook secret in 3 steps:
- Step 1 — Generate a random secret string once:
python3 -c "import secrets; print(secrets.token_urlsafe(24))" # Example output: finops_live_9xK2mP7vQ4wL8nR1zT5yB3cF
- Step 2 — Share that secret with BOTH the Dashboard Server and your Agent containers (via a shared
.envfile locally, or via GCP Secret Manager / Kubernetes Secrets / Cloud Run env vars in production):# On the Dashboard Server VM / Container: export ADK_FINOPS_DASHBOARD_API_KEY="finops_live_9xK2mP7vQ4wL8nR1zT5yB3cF" adk-finops dashboard --host 0.0.0.0 --port 8088 --log-dir central_logs
- Step 3 — Your Agents send the exact same secret automatically in the
X-FinOps-Keyheader:import os from adk_finops import FinOpsCostPlugin, HTTPExporter # If ADK_FINOPS_DASHBOARD_API_KEY is in your agent's .env or environment, # FinOpsCostPlugin and HTTPExporter pick it up automatically! finops_plugin = FinOpsCostPlugin( dashboard_endpoint="http://finops-dash.internal:8088", dashboard_api_key=os.getenv("ADK_FINOPS_DASHBOARD_API_KEY"), export_tags={"team": "payments", "service": "refund_agent"}, )
When the agent finishes a turn/session,HTTPExportersendsX-FinOps-Key: finops_live_9xK2mP7vQ4wL8nR1zT5yB3cFonPOST /api/ingest. The dashboard server compares that header against its ownADK_FINOPS_DASHBOARD_API_KEYusing constant-timehmac.compare_digest— accepting matching requests (200 OK) and rejecting anyone else (401 Unauthorized).- Real-Time Memory + Automatic Disk Persistence (
<log_dir>/ingested_costs.jsonl): Every HTTP-pushed batch is immediately served from bounded in-memory LRU cache (MAX_INGESTED_ROWS = 5,000,source="http_ingest") and appended tocentral_logs/ingested_costs.jsonlon the dashboard server. If the dashboard server is restarted later with--log-dir central_logs, all previously pushed sessions and agent hierarchies are automatically restored fromingested_costs.jsonlwithout any data loss or double-counting.
- Real-Time Memory + Automatic Disk Persistence (
How Hybrid Local + BigQuery Deduplication Works
When both local logs (jsonl_path / csv_path) and bigquery_table are enabled, the dashboard deduplicates every record by (session_id, scope, agent_name):
- Zero Double-Counting: Active and local sessions load in
0msfrom memory/disk; if the samesession_idalso exists in BigQuery, the duplicate cloud row is ignored. - Historical Backfill: Older runs or teammate sessions that exist only in BigQuery are seamlessly merged into the dashboard when Include BigQuery is checked.
Unified CLI & Programmatic Reporting (adk-finops report — Local Files + BigQuery)
You can generate a unified, deduplicated FinOps report directly from the terminal (or export it as a GitHub PR Markdown table or JSON file) across local .jsonl/.csv logs, BigQuery, or both combined:
1. CLI Usage (adk-finops report)
# 1. Unified Rich Terminal Report (merges & deduplicates local ./logs + BigQuery if ADK_FINOPS_BIGQUERY_TABLE is set)
adk-finops report --log-dir ./logs --bigquery-table my-project.finops.agent_costs
# 2. Local Files Only (skip BigQuery)
adk-finops report --log-dir ./logs --no-bigquery
# 3. BigQuery Only (skip local files) and export as a GitHub PR Markdown report
adk-finops report --log-dir none --bigquery-table my-project.finops.agent_costs --format markdown --output finops_report.md
# 4. Filter by Root Agent, Sub-Agent, or Session Outcome (--status success|failed) and export as JSON
adk-finops report --log-dir ./logs --root-agent coordinator_agent --status failed --format json --output failed_runs.json
2. Programmatic Python API (generate_finops_report)
from adk_finops import generate_finops_report
report = generate_finops_report(
log_dir="logs",
bigquery_table="my-project.finops.agent_costs",
include_bigquery=True,
output_format="markdown", # "table" (Rich CLI), "markdown", or "json"
output_path="finops_summary.md",
print_report=True,
)
print("Total Net Spend:", report["kpis"]["total_cost_usd"])
Cloud-Agnostic & Local Exporters (JSONL, CSV, OpenTelemetry)
You don't need a cloud warehouse to persist FinOps telemetry. adk-finops includes built-in local file exporters (JSONLExporter, CSVExporter) and an OpenTelemetryExporter that work completely offline or with any observability backend (DuckDB, Pandas, Jaeger, Datadog, Arize Phoenix, Honeycomb).
1. Configure Local & OTEL Exporters on FinOpsCostPlugin
from adk_finops import FinOpsCostPlugin
finops_plugin = FinOpsCostPlugin(
jsonl_path="logs/finops_costs.jsonl", # or set ADK_FINOPS_JSONL_PATH
csv_path="logs/finops_costs.csv", # or set ADK_FINOPS_CSV_PATH
enable_otel=True, # or set ADK_FINOPS_ENABLE_OTEL=true
export_scope="session", # "session", "turn", or "both"
export_tags={"env": "local_dev"},
)
Each exported row in .jsonl and .csv automatically includes first-class task outcome columns (status, is_failure, error) alongside token counts, USD costs, context caching savings, and agent/model/tool breakdowns (breakdown_by_agent, breakdown_by_model, breakdown_by_tool).
2. Query Local .jsonl / .csv Logs Instantly with DuckDB or Pandas
-- Query local JSONL file directly using DuckDB CLI
SELECT
status AS task_outcome,
is_failure AS is_wasted_spend,
COUNT(*) AS total_runs,
ROUND(SUM(total_cost_usd), 4) AS total_spend_usd
FROM read_json_auto('logs/finops_costs.jsonl')
WHERE scope = 'session' AND agent_name IS NULL
GROUP BY 1, 2;
3. OpenTelemetry Semantic Conventions (OpenTelemetryExporter)
When enable_otel=True (or OpenTelemetryExporter is used), adk-finops enriches the active span and emits gen_ai.finops.<scope> spans with standard attributes:
gen_ai.usage.input_tokens,gen_ai.usage.output_tokens,gen_ai.usage.thoughts_tokens,gen_ai.usage.cached_tokens,gen_ai.usage.total_tokensgen_ai.usage.cost_usd,gen_ai.usage.llm_cost_usd,gen_ai.usage.tool_cost_usd,gen_ai.usage.savings_usd,gen_ai.usage.tool_calls_countgen_ai.finops.breakdown_by_tool(JSON-serializedAgent -> Tool -> Calls & Costmap)gen_ai.finops.task_outcome(success,failed,error,budget_exceeded)gen_ai.finops.is_wasted_spend(true/false)
4. Standalone Exporter Usage (BigQueryExporter, JSONLExporter, CSVExporter, OpenTelemetryExporter)
You can also use any exporter directly in standalone Python scripts or custom agent frameworks:
from adk_finops import CostTracker
from adk_finops.exporters import (
BigQueryExporter,
CSVExporter,
JSONLExporter,
OpenTelemetryExporter,
)
exporters = [
JSONLExporter("finops_costs.jsonl"),
CSVExporter("finops_costs.csv"),
OpenTelemetryExporter(),
# BigQueryExporter(table_id="my-gcp-project.finops.agent_costs"),
]
with CostTracker.track_run("batch_job_42"):
CostTracker.record_usage(
run_id="batch_job_42",
model_name="gemini-2.5-pro",
prompt_tokens=15000,
completion_tokens=800,
cached_tokens=12000,
agent_name="data_extractor",
)
CostTracker.record_task_status(session_id="batch_job_42", status="success")
summary = CostTracker.get_summary("batch_job_42")
for exporter in exporters:
exporter.export_summary(summary, scope="session", tags={"env": "prod", "pipeline": "etl"})
Sample SQL Queries for Looker Studio
Once data streams into BigQuery, power executive dashboards and chargeback reports with standard SQL:
1. Top 5 Most Expensive Agents by LLM & Grounding Spend
SELECT
agent_name,
COUNT(DISTINCT session_id) AS total_sessions,
SUM(total_tokens) AS total_tokens,
ROUND(SUM(llm_cost_usd), 4) AS llm_cost,
ROUND(SUM(tool_cost_usd), 4) AS grounding_fees,
ROUND(SUM(total_cost_usd), 4) AS total_spend,
ROUND(SUM(savings_usd), 4) AS caching_dollars_saved
FROM `my-gcp-project.finops.agent_costs`
WHERE timestamp >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 30 DAY)
AND agent_name IS NOT NULL
GROUP BY 1
ORDER BY total_spend DESC
LIMIT 5;
2. Context Caching Savings & ROI by Model
SELECT
model_name,
SUM(cached_tokens) AS total_cached_tokens,
ROUND(SUM(gross_cost_usd), 4) AS gross_spend_without_cache,
ROUND(SUM(total_cost_usd), 4) AS actual_net_spend,
ROUND(SUM(savings_usd), 4) AS net_dollars_saved,
ROUND(SAFE_DIVIDE(SUM(savings_usd), SUM(gross_cost_usd)) * 100, 1) AS overall_savings_pct
FROM `my-gcp-project.finops.agent_costs`
WHERE timestamp >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 7 DAY)
AND scope = 'session'
AND agent_name IS NULL
GROUP BY 1
ORDER BY net_dollars_saved DESC;
3. Successful vs. Wasted Spend & Failure Root-Cause Analysis
SELECT
COALESCE(JSON_VALUE(tags, '$.status'), 'success') AS task_outcome,
COALESCE(JSON_VALUE(tags, '$.is_failure'), 'false') AS is_wasted_spend,
COUNT(*) AS total_runs,
ROUND(SUM(total_cost_usd), 4) AS total_spend_usd,
ROUND(AVG(total_cost_usd), 4) AS avg_cost_per_task,
ROUND(AVG(total_tokens), 0) AS avg_tokens_per_task
FROM `my-gcp-project.finops.agent_costs`
WHERE scope = 'session'
AND agent_name IS NULL
GROUP BY 1, 2
ORDER BY total_spend_usd DESC;
License
Distributed under the Apache License 2.0. See LICENSE for details.
Release files for adk-finops 1.0.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| adk_finops-1.0.1.tar.gz | 135.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| adk_finops-1.0.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 263.6 kB
Release files / adk_finops-1.0.1.tar.gz
| Download URL | adk_finops-1.0.1.tar.gz |
|---|---|
| Size | 135.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
bebae9ab235646b3ed9df6c0ced387836b793533825387b8d7a46246b5316e02
|
|
BLAKE2b-256 checksum How to use checksums |
fda7349196b6e35b19f1d0d6182a2409c9c1e18232c18f55fdb722a8bf34104d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / adk_finops-1.0.1-py3-none-any.whl
| Download URL | adk_finops-1.0.1-py3-none-any.whl |
|---|---|
| Size | 128.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
52321038d372d1db5390350ef69e5f48714673ca868953b3a598613fdc87e0c0
|
|
BLAKE2b-256 checksum How to use checksums |
f9382bea021af8cc0e905776231c708c84b3cd301d9a3549be466b1402e0fce0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|