Skip to main content

LLM agent policy for Inspect Robots: frontier LLMs (Claude, GPT, anything OpenAI-compatible) drive any registered embodiment through tool calls.

Project description

inspect-robots-agent

LLM agent policy for Inspect Robots: frontier LLMs (Claude, GPT, anything behind an OpenAI-compatible API) drive any registered embodiment through tool calls, as a first-class Policy named agent. The same policy runs ad-hoc instructions and scores on registered tasks next to fine-tuned VLAs.

Install

pip install inspect-robots inspect-robots-agent

Quickstart (no hardware)

export ANTHROPIC_API_KEY=sk-ant-...

inspect-robots "pick up the cube" --policy agent \
    -P model=anthropic/claude-fable-5 --embodiment cubepick

Model strings are OpenRouter-style provider/model, resolved from -P model=... or $INSPECT_ROBOTS_MODEL. API keys come from the environment:

  1. -P base_url=... (with -P api_key_env=NAME): any OpenAI-compatible endpoint
  2. anthropic/* model with ANTHROPIC_API_KEY: the Anthropic compat endpoint
  3. openai/* model with OPENAI_API_KEY: OpenAI
  4. OPENROUTER_API_KEY: OpenRouter, any model string

How it works

Motion tool calls state where to go, not how long to move. For absolute modes, the move tool (move_joints for joint spaces, move_to for Cartesian pose modes) interpolates named partial targets from the observed state at a fixed safe speed. The default max_speed_frac=0.1 allows a tenth of each dimension's range per second, subject to a 5%-of-range per-step ceiling that matches the core's default delta backstop. At that default a near-full-range move exceeds the 10 s per-call playout cap, so the agent receives a split-the-move error and issues it as two smaller motions; raise the fraction (up to 0.5 before the ceiling binds at 10 Hz) for faster arms. The tool result reports the computed step count and, when the embodiment declares control_hz, the corresponding playout time. duration_s is not part of either motion tool.

For displacement modes, move_by splits the requested total so every action fits the box side in that direction. The action box is the embodiment author's per-step speed statement, so max_speed_frac does not apply to displacement modes. done and give_up end the trial through the core's policy-stop channel.

When control_hz is None, the plugin uses a 10 Hz fallback to compute step counts and the per-call playout cap, but leaves the emitted chunk rate unset. The embodiment then plays the chunk at its native rate. In this case the speed and playout caps are step-count constructs, not wall-clock guarantees, and the tool result does not report seconds.

When the embodiment publishes operating notes via EmbodimentInfo.docs (joint layout, sign conventions, gripper polarity), the policy appends them to the system prompt as an Embodiment notes: section. The per-step observation also labels the proprioceptive state vector with the action dimension names (left_j0=0.01 ...) whenever the mapping is unambiguous.

Every action still passes the CLI's default safety approvers (bounds clamp plus per-step delta limit); the plugin contains no safety-critical code path of its own. An explicit --max-action-delta tighter than 5% of range can truncate absolute interpolants. In displacement modes, a value tighter than the action box can truncate each move_by step. Either setting can make the executed motion fall short of the tool's requested total.

Warning: Guardrails are on by default at the CLI. Never pass --disable-guardrails on real hardware unless you fully trust the policy and the rig.

Configuration knobs (all -P key=value): model, base_url, api_key_env, max_llm_calls (default 100), temperature, effort, max_speed_frac. The speed fraction defaults to 0.1 and applies only to absolute modes.

LLMAgentPolicy.transcript() returns the current conversation as a deep copy with streamed camera frames replaced by omission markers, ready for core eval-log persistence.

Reasoning effort defaults to low: robot control is latency-sensitive (the arm stands still while the model thinks), safety guardrails sit below the model either way, and frontier models at low effort remain strong at this task shape. Raise it for hard manipulation problems (-P effort=high) or pass -P effort=none to omit the parameter for endpoints that reject it (the CLI reads a bare none as null). To send the literal wire value none and disable reasoning, quote it: -P effort="'none'". GPT-5.x on chat completions requires the literal none when function tools are in play (any other value, or omitting the field, is a 400). In Python, effort=None omits the field and effort="none" sends the wire value.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

inspect_robots_agent-0.5.0.tar.gz (32.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

inspect_robots_agent-0.5.0-py3-none-any.whl (18.8 kB view details)

Uploaded Python 3

File details

Details for the file inspect_robots_agent-0.5.0.tar.gz.

File metadata

  • Download URL: inspect_robots_agent-0.5.0.tar.gz
  • Upload date:
  • Size: 32.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for inspect_robots_agent-0.5.0.tar.gz
Algorithm Hash digest
SHA256 a288a4666a51f8ff3b91cd8ffce549e546baa21c592ca2b5ee2baa02aca8687f
MD5 205bb785b6b2f0f4ef7a17166f11c572
BLAKE2b-256 a34963f9a4276778aa8029be66194373fa3497ef0f4a7913761970a6c82879b0

See more details on using hashes here.

Provenance

The following attestation bundles were made for inspect_robots_agent-0.5.0.tar.gz:

Publisher: release.yml on robocurve/inspect-robots

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file inspect_robots_agent-0.5.0-py3-none-any.whl.

File metadata

File hashes

Hashes for inspect_robots_agent-0.5.0-py3-none-any.whl
Algorithm Hash digest
SHA256 78eefb89c4dad66371c32427b6d695216bbb9ce99b99262477266a64f1cfc9d9
MD5 401fc8108a0dbc945a1e6d4d4ef43325
BLAKE2b-256 a14b8335f85f23a91871a36c8518d48ef079d3cfbb5c4c3f501fcd75625959ac

See more details on using hashes here.

Provenance

The following attestation bundles were made for inspect_robots_agent-0.5.0-py3-none-any.whl:

Publisher: release.yml on robocurve/inspect-robots

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page