Skip to main content

sequence-ai

Cloud inference for robot policies.

pip install sequence-ai
import sequence_ai

with sequence_ai.connect(model="pi05-droid") as policy:
    out = policy.run(
        observe=robot.read_observation,
        act=robot.apply_action,
        validate=robot.is_safe,
        hold=robot.hold,
        max_actions=250,
    )

What it does for you

Four things, and each of them is something a loop written straight against the HTTP endpoints has to get right before it behaves properly on a robot:

Without it
Keeps the connection open 185 ms of handshake on every call — 60% of a bare request
Refills before the buffer empties The arm stops once per chunk, for a full round trip
Classifies errors No way to tell "wait 118 s" from "stop and page someone"
One driver per buffer Two loops on one handle interleave, and both report success

None of the four is guesswork. Every one is a defect that existed in this client, was measured, and was fixed — the numbers below are those measurements, not estimates.

It keeps the connection open

Measured against the production gateway, six samples each:

median min max
new connection per call 305.9 ms 246.7 766.7
one reused connection 121.0 ms 116.8 133.4

184.9 ms per call — 60% of the total — is TCP and TLS handshake. That is what with sequence_ai.connect(...) removes. Nothing proprietary: any HTTP client that reuses a connection gets the same result. This package just makes it the default rather than something you have to remember.

Why it matters: a chunk is a deadline. pi05-droid returns 1.00 s of motion per call, and a warm end-to-end /v1/act through this gateway measured 793 ms — the next chunk has to arrive before the current one finishes playing, so 185 ms of avoidable handshake is a fifth of the entire budget.

Actions come in chunks

One call returns a block of future actions, not a single command — between 0.5 s and 2.1 s of motion depending on the model. That is why cloud inference works at all: the control loop does not need a network round trip per control step.

with sequence_ai.connect(model="pi05-droid") as policy:
    pred = policy.predict(observation)
    print(len(pred.action_chunk), "steps covering",
          pred.action_chunk.covers_seconds, "s at",
          pred.action_chunk.control_frequency_hz, "Hz")

The gap between chunks is the hard part

A loop that waits for the buffer to empty before asking for the next chunk stops the arm once per chunk, every chunk, for a full round trip. It is not an occasional hiccup — it is structural, and on the faster models it dominates:

model chunk covers refill arm actually moving
cosmos3-edge-policy-droid 2.13 s ~0.79 s 73%
pi05-droid 1.00 s ~0.79 s 56%
lingbot-va-5b 0.64 s ~0.79 s 45%
lingbot-vla-v2-6b 0.50 s ~0.79 s 39%

So run() starts the next inference while the current chunk is still playing, and reports what happened rather than hoping:

out = policy.run(observe=..., act=..., max_actions=250, on_underrun=robot.hold_position)

print(out)        # RunOutcome(250 actions over 17 chunks in 16.8s, completed)
out.underruns     # times the buffer ran dry before the next chunk arrived
out.underrun_s    # total seconds the arm spent with no command
out.max_seam_jump # largest per-dimension step across a chunk boundary

The trade, stated plainly: a prefetched chunk is computed from an observation taken before the previous chunk finished, so the overlap window is open-loop. Freshness and continuity are in direct opposition here and no setting gets both. prefetch=False restores strictly closed-loop behaviour, stall included — right for bench work, wrong for a moving arm.

max_seam_jump is a measurement, not a correction, and it stays one unless you ask otherwise. smooth_seam=N ramps the first N steps of each new chunk out of the last executed action, and it is the only setting in this library that changes a number on its way to the motors — so it is off by default and guarded twice: it applies only when the chunk declares an absolute action space (action_chunk.action_space), and never to the first chunk of a run. Blending deltas is not smoothing; it rescales the increments and moves the arm somewhere the model never asked for, so a delta chunk is passed through untouched even when you ask. max_seam_jump is measured before any blending, so turning it on cannot hide what it smooths.

Many robots at once

One Policy per control loop. A Policy holds one action buffer, and two loops popping from it do not take turns — they interleave. Driving the same handle from a second thread raises, because both quieter options are worse: interleaving sends each robot a shuffled half of the other's plan while both loops report success, and a blocking lock would make the second loop run at half rate, underrunning on every chunk.

def drive(robot, model):
    with sequence_ai.connect(model=model) as policy:      # one each
        return policy.run(observe=robot.read, act=robot.apply,
                          validate=robot.is_safe, max_actions=250)

with ThreadPoolExecutor() as pool:
    left, right = pool.map(drive, [arm_l, arm_r], ["pi05-droid"] * 2)

This costs nothing to follow. The handshake is paid once per Policy, not once per call. Measured with eight loops running concurrently against one gateway: eight connections, ten requests each, zero sequence breaks, zero underruns. Requests arriving while the gateway is busy queue on the server rather than displacing work already in flight.

Cold starts

A worker that is not loaded takes about 100 seconds to become ready. The gateway does not stall on it — it answers 503 immediately with the number:

try:
    policy.predict(observation)
except sequence_ai.Unavailable as exc:
    if exc.warming:
        print(f"loading; ready in ~{exc.retry_after_s}s")   # 118

run() can wait that out for you, but only before the first action:

policy.run(..., startup_timeout_s=180)

That line is the whole design. Before the first action nothing is moving, so waiting is free. Once the arm is in motion, silently pausing it for two minutes and resuming from a two-minute-old plan is worse than stopping — so mid-run warming is raised, and hold fires.

Three ways to drive it

From most control to least. They are the same request underneath; the difference is who owns the loop.

policy.predict(obs)          # the whole chunk, you do everything
policy.next_action(obs)      # we hold the buffer, you own the cadence
policy.run(observe=, act=)   # we own the loop

Safety

An action returned by any model is model output, not a safe robot command.

This library does not check joint limits, reachability, collisions, or whether a step is safe at the robot's current velocity. Bounds checking, a watchdog and an e-stop belong between this library and your motors.

policy.run(
    observe=robot.read_observation,
    act=robot.apply_action,
    validate=robot.is_safe,   # return False to stop the loop
    hold=robot.hold,          # called if it stops early, or if act() raises
    max_actions=250,
)

max_actions is required and keyword-only. There is no run_forever() — an unbounded loop that moves a robot should not be startable by accident.

validate=None is allowed for bench and simulation work, and warns once so it cannot happen silently on real hardware.

Errors are typed

Because a controller reacts differently to each:

meaning what to do
AuthError key missing, revoked, expired stop; retrying will not help
OutOfCredit balance exhausted stop and hold; top up
InvalidRequest bad model, malformed observation, body too large fix it; deterministic
Unavailable upstream blip, or a cold worker check .warming — see below
ChunkExhausted asked for an action with an empty buffer and no observation pass observation= every call

Configuration

export SEQUENCES_API_KEY=seq_live_...      # or pass api_key= to connect()
export SEQUENCES_BASE_URL=...              # for staging; defaults to production

Get a key at app.generalsequences.com.

Install footprint

One dependency: httpx. Python 3.9+.

A robot controller is often on a Jetson with a pinned, fragile Python environment, and ROS 2 Humble ships Python 3.10. Every transitive dependency is another chance for the install to fail on the machine that actually matters.

If your controller is not Python

The endpoints underneath are public and documented at generalsequences.com/docs. There is no private control plane and nothing this package reaches that you cannot:

curl https://api.generalsequences.com/v1/act \
  -H "Authorization: Bearer $SEQUENCES_API_KEY" \
  -d '{"model":"accounts/sequences/models/pi05-droid","observation":{...}}'

Keeping that door open is deliberate — a vendor SDK that is the only supported way in locks you to one language, and robot controllers are very often C++. But it is a door, not the front entrance. The four items in the table at the top are work your client then has to do itself, and the reference documents each precisely enough to reimplement. Reaching for Python first is simply cheaper.

Metadata

Release files for sequence-ai 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sequence-ai 0.1.0
File Size Uploaded
sequence_ai-0.1.0.tar.gz 48.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sequence-ai 0.1.0
File Interpreter ABI Platform
sequence_ai-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 88.2 kB

Release files / sequence_ai-0.1.0.tar.gz

Download URL sequence_ai-0.1.0.tar.gz
Size 48.5 kB
Tags Source
SHA-256 checksum
How to use checksums
a03940fa70c7550897c089b62cd5159dd269f1ba9a9cff12088d92b9a8e5f2c6
BLAKE2b-256 checksum
How to use checksums
f3559daa440b278e2231947e37e291126fbb4feb1671e7a1dcd480500cd3f3a5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.2

Release files / sequence_ai-0.1.0-py3-none-any.whl

Download URL sequence_ai-0.1.0-py3-none-any.whl
Size 39.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4db0258ff32e19d0281a4de13c8dc35f3782456c202087c1cc223a1d30fb3c57
BLAKE2b-256 checksum
How to use checksums
f332c5097eb000025dd04ecf3ae309045f8927a6adaa1f73e0664bd37c64cab0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.2

Release history Release notifications | RSS feed

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page