Skip to main content

CLI and shared client library for Applied Labs AI support agents

Project description

Applied Labs CLI

CLI and client library for Applied Labs AI support agents.

Installation

pip install applied-cli

CLI Usage

# Authenticate
applied login

# List agents
applied agents
applied agents --format json

# Find and inspect conversations
applied conversation-find --query "refund" --limit 5
applied conversations --resolution escalated --view detail --limit 10
applied conversation <id> --messages --format json

# Query tickets
applied tickets --status open

# Knowledge base
applied knowledge --type qa --search "refund"
applied knowledge <id> --format json
applied knowledge --source "https://help.example.com" --format csv
applied knowledge-diff --source "https://help.example.com"
applied content list --source "https://help.example.com"
applied content-update <content_id> --text "Corrected source text"
applied content-resync <content_id> --wait
applied knowledge-pin <id>
applied knowledge-link <response_id> --content <content_id>
applied knowledge-protect <id>
applied knowledge-unprotect <id>

# Taxonomy
applied taxonomy --type topics
applied taxonomy-counts --start 2026-03-01 --end 2026-03-18 --format csv

# Analytics
applied analytics-report --view overview_aggregate_metrics --model conversation --start 2026-04-01 --end 2026-04-30 --format json
applied analytics --group-by topic --metrics count --start 2026-04-01 --end 2026-04-30 --format json
applied analytics --group-by intent --metrics count --start 2026-04-01 --end 2026-04-30 --format json
applied metrics --metric-name conversation.resolve --start 2026-04-01 --end 2026-04-30 --period day --format json

analytics-report returns the selected report payload, not necessarily a { "rows": [...] } object. analytics returns grouped rows and currently supports --metrics count. Raw analytics SQL is not available through the public CLI surface.

Benchmarks & Scenarios

LEGACY surface. Benchmarks and scenarios are the older regression-suite system. The current evaluation system is the /evals surface (feedback + simulations + test data), driven by the simulation_test_*, test_customer_*, and simulation_run_* tools (MCP / /v1/simulation-tests/, /v1/test-customers/, /v1/simulation-runs/). The feedback field on scenario_update / scenario_run_update is a legacy free-text note, NOT the new Judgment+Comment feedback surface. Use the /evals tools for new evaluation work; the tools below are for maintaining existing benchmark suites.

A benchmark is a named regression suite; a scenario is one test conversation (built from a real input_conversation_id) that can belong to one or more benchmarks. The typical loop is: build a suite → run it → review the pass rate → fix → re-run.

# Inspect benchmarks and their scenarios
applied benchmarks --agent-id <agent_id> --format json
applied benchmark <benchmark_id> --format json
applied scenarios --benchmark-id <benchmark_id> --format json

# Build a suite
applied benchmark-create --agent-id <agent_id> --name "Cancel Regression"
applied scenario-create --input-conversation-id <conversation_id> --name "<name>" \
  --benchmark-id <benchmark_id>

# Build a suite fast from several real conversations at once. Each scenario is
# named from its conversation's title (or "<prefix> N" with --name-prefix).
applied scenario-create-bulk --conversation-ids <id1>,<id2>,<id3> \
  --benchmark-id <benchmark_id>

# Port a suite to another agent (e.g. email -> chat). Cross-agent recreates the
# scenarios under the destination agent; same-agent just tags them in.
# Dry-run by default; add --apply to write.
applied benchmark-clone <source_benchmark_id> --dest-benchmark-name "Chat Regression" \
  --target-agent-id <chat_agent_id> --apply

# Run a benchmark and wait for results in one command.
# --contact-email runs as a contact that has an email, fixing
# "Email is not present in the conversation" on test conversations.
applied scenario-bulk-run --benchmark-id <benchmark_id> \
  --contact-email test@example.com --wait
applied scenario-bulk-status <job_id> --include-runs --format json

# Kill a stuck bulk run (deletes its queued/running runs; finished runs preserved)
applied scenario-bulk-cancel <job_id> --apply

# Review pass/fail health (pass_status reflects the latest run per scenario)
applied benchmark-results <benchmark_id> --format json

# Portfolio go/no-go: pass rates across all of an agent's benchmarks at a glance
applied benchmarks --agent-id <agent_id> --with-results --format json

# Rate scenarios as you evaluate
applied scenario-update <scenario_id> --pass-status pass --feedback "<note>"

# Safe delete — refuses to wipe scenarios unless you opt in
applied benchmark-delete <benchmark_id> --detach-scenarios   # preserve scenarios
applied benchmark-delete <benchmark_id> --force              # cascade delete

# Recover deleted benchmark/scenario rows from a local PITR export
applied scenario-recover-catalog --recovery-dir <dir> --apply

Deleting a benchmark cascades and permanently deletes its scenarios and runs, so benchmark-delete refuses a non-empty benchmark unless you pass --detach-scenarios (unlink the scenarios first so they survive under their agent) or --force.

Library Usage

from applied_cli import AppliedClient, tools

# Create client
client = AppliedClient(token="al_xxx")

# Use tools
agents = await tools.agent_list(client, output_format="csv")
conversations = await tools.conversation_query(
    client,
    filters={"resolution": "escalated"},
    view="search",
    limit=20,
)

Tools

Tool Description
agent_list List all AI agents
conversation_find Find conversations with compact agent-friendly fields
conversation_get Get single conversation with messages
conversation_query Search/filter conversations
ticket_query Search/filter tickets
knowledge_list List knowledge base items
taxonomy_list List topics, intents, and flags
taxonomy_counts Aggregate conversation counts by topic and intent
analytics_report Read standard dashboard/report analytics views
analytics_query Aggregate supported conversation dimensions with count
metrics_query Roll up named metric events
benchmark_clone Copy all scenarios from one benchmark into another
benchmark_delete Delete a benchmark (guards against wiping scenarios)
benchmark_results Pass/fail/unrated tally and pass rate for a benchmark
benchmark_list List benchmarks (with per-benchmark pass rates via with_results)
scenario_create_bulk Build scenarios from several conversations at once
scenario_bulk_run Run scenarios (contact override + wait-to-completion)
scenario_bulk_cancel Cancel a stuck bulk run's queued/running scenario runs
review_run Grade real conversations with the AI Reviewer (pass/fail)
review_status Verdict pairs + the Reviewer's reason for a set of conversations
review_set_verdict Record the human verdict with a teaching reason
review_teach Teach the Reviewer a rule directly
review_make_test Turn a flagged conversation into a proposed regression test
review_proposals_list The suite change-inbox (proposed adds/edits/retires/merges)
review_proposals_decide Accept or dismiss a proposed test

The full review_* namespace also includes review_get, review_clear_verdict, review_add_note, review_trust, and review_lessons — the outer review loop: grade real traffic, confirm/overrule with a reason that teaches the Reviewer, and promote failures into the simulation suite as tests.

Examples

# Find conversations using the default compact search view
await tools.conversation_find(
    client,
    query="refund",
    output_format="csv"
)

# Find escalated conversations
await tools.conversation_query(
    client,
    filters={"resolution": "escalated"},
    view="detail",
    output_format="csv"
)

# Get conversation with transcript
await tools.conversation_get(
    client,
    conversation_id="abc-123",
    include_messages=True,
    message_limit=50,
    output_format="json"
)

# Search knowledge base
await tools.knowledge_list(
    client,
    kb_type="qa",
    search="refund policy",
    output_format="json"
)

Development

# Install dev dependencies
pip install -e ".[dev]"

# Run tests
pytest

# Lint
ruff check --fix .
ruff format .

License

MIT

Project details


Release history Release notifications | RSS feed

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

applied_cli-0.7.1.tar.gz (221.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

applied_cli-0.7.1-py3-none-any.whl (177.4 kB view details)

Uploaded Python 3

File details

Details for the file applied_cli-0.7.1.tar.gz.

File metadata

  • Download URL: applied_cli-0.7.1.tar.gz
  • Upload date:
  • Size: 221.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for applied_cli-0.7.1.tar.gz
Algorithm Hash digest
SHA256 8cf1636c03fef501e9fa9fa14b3083d4f26611b718006c27a1505f26508f3690
MD5 163404775c883e4b3a269d2b796928a7
BLAKE2b-256 c30f9d88952a8396d896030c5347220092c610a6d0be3d3e8f67e2bb2b87957a

See more details on using hashes here.

Provenance

The following attestation bundles were made for applied_cli-0.7.1.tar.gz:

Publisher: publish.yml on AppliedLabsAI/applied-cli

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file applied_cli-0.7.1-py3-none-any.whl.

File metadata

  • Download URL: applied_cli-0.7.1-py3-none-any.whl
  • Upload date:
  • Size: 177.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for applied_cli-0.7.1-py3-none-any.whl
Algorithm Hash digest
SHA256 0ae7c65d78ee4c0cc5a30e6f389e6c73827e2c7eb60a78e9e15f0c331c2179f6
MD5 13d6d66404ae1cabef6f914f229f6d35
BLAKE2b-256 4e06fdb742d7d7c9d64c959d7c84381bbf2d2a6eb3759a62f95d01e40ad2ee49

See more details on using hashes here.

Provenance

The following attestation bundles were made for applied_cli-0.7.1-py3-none-any.whl:

Publisher: publish.yml on AppliedLabsAI/applied-cli

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page