sequence-ai
Cloud inference for robot policies.
pip install sequence-ai
import sequence_ai
with sequence_ai.connect(model="pi05-droid") as policy:
out = policy.act(
"pick up the cup",
observe=robot.read_observation, # cameras + joints
act=robot.apply_action, # one joint target
validate=robot.is_safe,
hold=robot.hold,
max_seconds=30,
)
policy.interrupt() # from any thread, at any moment
Try the loop with no account and no network. Pass base_url="mock" (or set
SEQUENCES_BASE_URL=mock) and connect() runs the whole connect → act loop against an in-process
gateway and a self-signed HTTPS worker — the real direct-connect transport, cert-pin and all, so
what you run offline is the code path production takes. No key required:
import sequence_ai
with sequence_ai.connect(model="pi05-droid", base_url="mock") as policy:
for _ in range(100):
chunk = policy.predict(my_observation).action_chunk # a real ActionChunk, every call
Two brains, one string between them
A VLM looks at a snapshot, decides what to do next, and says it in a sentence. This package is the how: it takes that sentence with the current cameras and joints, asks a vision-language-action model for the next second of joint targets, and plays them out while fetching the next second.
The two halves run at completely different speeds — the brain thinks between subtasks, the arm
needs a chunk every second — and instruction is their entire interface.
| endpoint | who ships it | |
|---|---|---|
| slow brain | POST /v1/chat |
Messages in, one message out |
| fast brain | this package | policy.act("pick up the cup", ...) |
act() is shaped to be called as a tool. It overwrites the observation's instruction on every
call, so your eyes never have to know what the brain last decided and a stale sentence cannot
leak into a chunk after the brain has moved on. max_seconds is required and has no default:
this drives hardware, and a tool call a language model can start but nothing bounds is one an
arm can be left running by a dropped conversation.
Wiring the two together — retries, failure detection, choosing the next sub-goal — is yours. This package provides inference and a control loop, not orchestration.
Changing your mind
policy.interrupt()
Thread-safe, and the only method meant to be called from a thread other than the one driving.
It stops within one control period, not one chunk. The loop checks between every action, so at 15 Hz the arm stops in about 67 ms. Waiting for the current chunk to finish playing would be up to a full second of an arm still reaching for something the brain has already given up on.
It also cuts through a stalled fetch. If the buffer has run dry and the loop is blocked on a chunk that is taking five seconds, an interrupt does not wait that out — stop latency is a property of this loop and never of the service.
Interruption unwinds through hold() like any other early stop, and is not raised: a stop you
asked for is an outcome, not an error. RunOutcome.reason says which ending happened.
reason |
meaning |
|---|---|
completed |
ran out of actions to play |
until |
your until() returned True — you judged the subtask done |
timeout |
max_seconds elapsed |
interrupted |
someone called interrupt() |
validate |
a validate callback refused an action |
raised:<Class> |
your callback or the network raised |
What it does for you
Four things, and each of them is something a loop written straight against the HTTP endpoints has to get right before it behaves properly on a robot:
| Without it | |
|---|---|
| Keeps the connection open | 185 ms of handshake on every call — 60% of a bare request |
| Refills before the buffer empties | The arm stops once per chunk, for a full round trip |
| Classifies errors | No way to tell "wait 118 s" from "stop and page someone" |
| One driver per buffer | Two loops on one handle interleave, and both report success |
None of the four is guesswork. Every one is a defect that existed in this client, was measured, and was fixed — the numbers below are those measurements, not estimates.
It keeps the connection open
Measured against the production gateway, six samples each:
| median | min | max | |
|---|---|---|---|
| new connection per call | 305.9 ms | 246.7 | 766.7 |
| one reused connection | 121.0 ms | 116.8 | 133.4 |
184.9 ms per call — 60% of the total — is TCP and TLS handshake. That is what
with sequence_ai.connect(...) removes. Nothing proprietary: any HTTP client that reuses a
connection gets the same result. This package just makes it the default rather than
something you have to remember.
Why it matters: a chunk is a deadline. pi05-droid returns 1.00 s of motion per call, and a
warm end-to-end /v1/act through this gateway measured 793 ms — the next chunk has to
arrive before the current one finishes playing, so 185 ms of avoidable handshake is a fifth
of the entire budget.
Actions come in chunks
One call returns a block of future actions, not a single command — between 0.5 s and 2.1 s of motion depending on the model. That is why cloud inference works at all: the control loop does not need a network round trip per control step.
with sequence_ai.connect(model="pi05-droid") as policy:
pred = policy.predict(observation)
print(pred.action_chunk)
# ActionChunk(15 steps x 8 dims, joint_delta, droid_franka, 15 Hz, covers 1.00s, replan after 8)
Read joint_delta in that line before you use the rows. It is the field that decides what
the numbers mean, and pi0.5-DROID's are increments — feeding them to a controller that expects
absolute joint targets is invisible: same shape, same dtype, values in the normal range, no
error, and the arm never completes the task. replan after 8 is the second one: executing all
15 rows when the model asks to be replanned after 8 fails the same quiet way.
You do not have to handle either one. Declare what your controller takes and commands() applies
both, plus the gripper exception and the value clip:
policy = sequence_ai.connect(model="pi05-droid", controller="joint_position")
for target in policy.commands(observation): # already absolute, already 8 rows
arm.move_to_joint_positions(target)
What to do with the rows covers both controllers and
the by-hand route. If you are wiring this into an existing harness rather than starting from
run(), read that section first — it is the only part of this document where getting it wrong
produces no error at all.
The gap between chunks is the hard part
A loop that waits for the buffer to empty before asking for the next chunk stops the arm once per chunk, every chunk, for a full round trip. It is not an occasional hiccup — it is structural, and on the faster models it dominates:
| model | chunk covers | refill | arm actually moving |
|---|---|---|---|
cosmos3-edge-policy-droid |
2.13 s | ~0.79 s | 73% |
pi05-droid |
1.00 s | ~0.79 s | 56% |
lingbot-vla-v2-6b |
0.50 s | ~0.79 s | 39% |
So run() starts the next inference while the current chunk is still playing, and reports
what happened rather than hoping:
out = policy.run(observe=..., act=..., max_actions=250, on_underrun=robot.hold_position)
print(out) # RunOutcome(250 actions over 17 chunks in 16.8s, completed)
out.underruns # times the buffer ran dry before the next chunk arrived
out.underrun_s # total seconds the arm spent with no command
out.max_seam_jump # largest per-dimension step across a chunk boundary
The trade, stated plainly: a prefetched chunk is computed from an observation taken
before the previous chunk finished, so the overlap window is open-loop. Freshness and
continuity are in direct opposition here and no setting gets both. prefetch=False restores
strictly closed-loop behaviour, stall included — right for bench work, wrong for a moving arm.
max_seam_jump is a measurement, not a correction, and it stays one unless you ask
otherwise. smooth_seam=N ramps the first N steps of each new chunk out of the last executed
action, and it is the only setting in this library that changes a number on its way to the
motors — so it is off by default and guarded twice: it applies only when the chunk declares an
absolute action space (action_chunk.action_space), and never to the first chunk of a run.
Blending deltas is not smoothing; it rescales the increments and moves the arm somewhere the
model never asked for, so a delta chunk is passed through untouched even when you ask.
max_seam_jump is measured before any blending, so turning it on cannot hide what it smooths.
Many robots at once
One Policy per control loop. A Policy holds one action buffer, and two loops popping
from it do not take turns — they interleave. Driving the same handle from a second thread
raises, because both quieter options are worse: interleaving sends each robot a shuffled half
of the other's plan while both loops report success, and a blocking lock would make the second
loop run at half rate, underrunning on every chunk.
def drive(robot, model):
with sequence_ai.connect(model=model) as policy: # one each
return policy.run(observe=robot.read, act=robot.apply,
validate=robot.is_safe, max_actions=250)
with ThreadPoolExecutor() as pool:
left, right = pool.map(drive, [arm_l, arm_r], ["pi05-droid"] * 2)
This costs nothing to follow. The handshake is paid once per Policy, not once per call.
Measured with eight loops running concurrently against one gateway: eight connections, ten
requests each, zero sequence breaks, zero underruns. Requests arriving while the gateway is
busy queue on the server rather than displacing work already in flight.
What the loop says while it runs
run() reports lifecycle, not telemetry:
starting… first chunk on its way, or a cold worker loading — nothing is moving yet
running a command has reached the robot
done
Waiting out a cold worker (three to six minutes) and losing a few hundred milliseconds at a
chunk boundary are the library's problems, not yours — startup_timeout_s defaults to 1800 s,
so the first call on a cold worker waits rather than failing. Every number is still in
RunOutcome afterwards if you want it.
policy.run(..., on_status=log.info) # programmatic
policy.run(..., progress=False) # silent
The default progress="auto" writes a single line to stderr only when stderr is a
terminal — nothing in a script, a pipe, or a log file.
Cold starts
A worker that is not loaded takes three to six minutes to become ready — tens of gigabytes of weights fetched, extracted and JIT-compiled. A control loop's budget for one chunk is 533 ms. Those two numbers are why connecting is a separate call from controlling.
Measured end to end, request to first action:
| model | cold start | warm call |
|---|---|---|
pi05-droid |
180–215 s | 600–800 ms |
groot-n1-7-3b |
~299 s | ~400 ms |
cosmos3-edge-policy-droid |
~346 s | 950–1200 ms |
The worker stays warm for 10 minutes
This is the number that decides how you structure a session. After it becomes ready, a worker is held for 10 minutes past its last request. Every call inside that window is warm — the table's right-hand column. The clock restarts on each request, so a loop that calls even once a minute never goes cold.
Cross that 10-minute gap with no calls and the worker is released, and the next request pays the full cold start again.
call ──► 180-215 s cold start ──► warm ◄──────── 10 min idle window ────────► released
▲ │
└────── any call restarts the 10 min ────────┘
Why ten and not thirty. A held GPU bills whether or not you call it, and the shape of that bill is brutal for a session with thinking time in it: over one real day of use, inference was 3% of the GPU seconds paid for, cold starts 19%, and idle holding 77%. Thirty minutes made the worst case a session that calls once every twenty-five minutes — never cold, almost never working, and billed for all of it. Ten covers a loop and a short pause between runs without paying for a coffee break. {measured 2026-09-16 across 2,147 calls: 322 s of inference against 7,200 s of idle holding}
The practical consequence, if you are running a batch of experiments: the gap between them is what costs you, not the number of them. Ten tasks run back to back pay one cold start. The same ten with a 20-minute analysis gap between each pay ten — that is up to half an hour of waiting that looks like the model being slow and is actually the fleet being released and rebuilt.
There is no keep-alive call yet: wait_until_ready() reports on a worker and starts a cold one, but it
does not hold a running one — only acting does. Keeping a worker warm (min_containers) is coming.
Connect first, then drive
policy.wait_until_ready() # once, before the robot needs to move
while running:
policy.next_action(observe()) # every one of these is warm by construction
wait_until_ready() needs no observation — a robot should be able to bring its model up before
it is in position, which is exactly when it has no frame worth sending. It is authenticated but
not billed: it touches no GPU. It is also what starts the worker, so polling it is the
thing that brings the model up, not merely a way to watch.
ready() is the non-blocking form, returning (ready, eta_seconds).
Skipping it does not fail — it silently costs you the loop. Measured on one run against a cold endpoint, the client starting anyway:
first control request 13,533 ms <- the cold start, now inside the loop
p50 415 ms <- the steady state was always fine
mean 1,096 ms -> 0.49 arms sustainable
mean without that one 441 ms -> 1.21 arms sustainable
One request that should not have been in the loop is the entire difference between sustainable and not.
If you skip it anyway
The gateway does not stall on a cold worker — it answers 503 immediately with the number:
try:
policy.predict(observation)
except sequence_ai.Unavailable as exc:
if exc.warming:
print(f"loading; ready in ~{exc.retry_after_s}s") # 174
run() waits that out for you by default (startup_timeout_s=1800), and only before the
first action. It reports starting… while it does, so a wait is never mistakable for a hang:
policy.run(..., startup_timeout_s=0) # opt out: fail immediately on a cold worker
That line is the whole design. Before the first action nothing is moving, so waiting is free.
Once the arm is in motion, silently pausing it for two minutes and resuming from a
two-minute-old plan is worse than stopping — so mid-run warming is raised, and hold fires.
How big your frames are is how fast you go
Latency scales with the bytes you send, at roughly 16 ms per kB. This is the single largest thing under your control, and it is larger than the model:
observation p50 mean sustainable arms
14.6 kB 586 ms 623 ms 0.86
3.4 kB 384 ms 433 ms 1.23
Measured through the gateway, alternating A/B against one warm worker over one connection so
that drift in the service cancels. inference_ms did not move between the two (119 ms against
111 ms) — every millisecond of the difference was transport. A direct measurement against the
worker agreed: 181 ms for the same 11.2 kB.
Send JPEG, not arrays. A frame as a list of integers is roughly 16x the bytes of the same frame as base64 JPEG, and it is the most common way to land on the slow side of that table.
Use quality 95. Do not go below 90. With the sampling noise pinned, on a real DROID frame:
| quality | max|Δ action| | share of action amplitude |
|---|---|---|
| 95 | 0.0048 | 0.51% |
| 90 | 0.0019 | 0.20% |
| 85 | 0.0177 | 1.86% |
| 75 | 0.0269 | 2.86% |
Quality 85 and below is where the error jumps by an order of magnitude. This section previously
read "quality 85 is safe" and quoted the 0.51% figure next to it — but 0.51% is q95's number,
and q85 is nearly four times worse. sequence_ai.encode() defaults to 95 so the decision is not
yours to get wrong.
Hand it the native frame and let it size it
Do not downscale before calling encode(). Give it whatever your camera produces. It reads
the catalogue and does one of two things, and which one is the model's decision, not a setting:
| model | what encode() sends |
why |
|---|---|---|
pi05-droid |
224x224 JPEG |
the padded square from openpi's own resize_with_pad, the identical call its official DROID client makes, so the server's copy short-circuits and the frame reaches the model untouched |
cosmos3-edge-policy-droid |
640x360 JPEG |
three of those compose to exactly the 540x640 its server expects |
groot-n1-7-3b |
short edge to 256 JPEG |
its own first step is SmallestMaxSize(256), so a frame already there makes that step an identity |
Each model's target is the size its transform produces, so the server's resize is a no-op on
arrival — one resample instead of two. Downscaling yourself to some other size is two resamples
where one would do, and the intermediate has no claim to being right: training_frame_size is
what the dataset stored, not what the model reads.
It runs the server's own operation, so the two agree. For pi05 that is resize_with_pad with
the server's filter, bit-identical at every aspect ratio; a frame outside the trained ratio is
refused locally with the same message the gateway would 400, rather than cropped or squashed. For
GR00T the short-edge resize is shrink-only — a frame already smaller is left alone.
Always JPEG, because that is the format every served handler asks for. Measured in action space against the model's own call-to-call spread, quality 95 is indistinguishable from lossless; the cliff is far below, so there is no reason to spend 3-4x the payload on PNG.
| what you send | payload per view | error at the model's input |
|---|---|---|
native 1280x720 PNG, uploaded whole |
3,607 kB | 0.00 — the baseline |
encode() for pi05 |
24 kB | inside the sampler's own run-to-run spread |
| the same, with a different resize filter | 24 kB | 0.78 — a different image |
Fidelity below that is not a gentle slope. Degrading the frame further does not degrade the policy proportionally; it stops working. Quantising each channel to four levels — layout, edges, shadows and hues all preserved, a milder change than any sane JPEG setting — took pi0.5-DROID from placing the cube in the bowl in 5.3 s to never lifting it at all, with its closest approach 20x worse than the successful run's final distance. Measured in simulation, 2026-09-11.
Four numbers you should never have to remember
Which camera views a model wants, how wide its state vector is, what aspect ratio it trained on, and what its returned floats mean. All four are per-model, none is guessable, and every one of them fails silently — an unknown view name is not rejected, it arrives as an empty camera; a short state vector is not rejected, it is zero-padded. The robot moves either way; it just moves on something other than what you meant.
So read them off the catalogue instead of remembering them:
model = {m.short_id: m for m in sequence_ai.models(endpoint="act")}["pi05-droid"]
obs = sequence_ai.observation(
model,
images={"exterior_1": cam.read(), "wrist_left": wrist.read()}, # ndarray or PIL
joints=arm.joint_positions(), gripper=[arm.gripper_position()],
instruction="pick up the cup",
)
That sizes each frame the way the model's own server would — resize_with_pad for pi05, short
edge for GR00T — encodes JPEG at quality 95, and raises locally, before anything is billed, if
the views, the state width, or the aspect ratio do not match what the model declares. policy.predict() runs the same check and warns when you build an observation
yourself; connect(validate="strict") makes it raise, "off" skips it.
state_dim is the input width and is not action_dim. Across the served catalogue it is 8, 17 and 8.
Read it off the model, not off this sentence.
state_layout says how that width splits — {"joint_positions": 7, "gripper": 1} for
pi0.5-DROID — and the halves are not interchangeable. Eight joints with no gripper sums to 8 and
is accepted; the eighth is then dropped and the gripper zero-filled, which holds the hand open
for the whole episode. observation() checks each half, not the sum.
What to do with the rows that come back
Tell it what your controller takes, once, and stop deciding.
policy = sequence_ai.connect(model="pi05-droid", controller="joint_position")
for target in policy.commands(observation):
arm.move_to_joint_positions(target)
That is the whole loop body. commands() applies the model's action space, its conversion
reference, its gripper exception and its replan horizon, and hands back rows that are already
what your controller expects — the list it returns is what should be executed before observing
again, not the whole chunk with a footnote.
controller has two values and no default, because guessing is exactly the failure this removes:
| your controller takes | pass | example |
|---|---|---|
| absolute joint targets | controller="joint_position" |
Isaac Lab's DROID scene |
| the model's own space | controller="native" |
the DROID stack's RobotEnv |
The reference for the increments is taken from the observation you just sent — that is "state_at_request" — so there is no second value to read at the wrong moment.
The same thing by hand, if you want the chunk
predict() is unchanged and returns everything. Four things are yours if you take this route,
and each has a wrong answer that raises nothing and produces an array of the right shape:
Send them raw. Every model reports pass_raw: true, and that is what the checkpoint's own
reference loop does on a real robot — take the row, binarise the gripper, clip, env.step(),
nothing added:
env = RobotEnv(action_space="joint_velocity", gripper_action_space="position") # DROID stack
chunk = policy.predict(obs).action_chunk
for i in range(chunk.open_loop_horizon or len(chunk)):
env.step(sequence_ai.prepare(chunk, model, i))
model.action_semantics tells you which controller interface those rows are for, so you can
find out before the first request rather than after. The catalogue spans four action spaces —
joint_delta, joint_absolute, ee_absolute, ee_delta — and reading one as another is
invisible in the shape — the array has the same width and dtype either way.
Convert only if your controller cannot take that interface. Isaac Lab's DROID scene has a joint-position action term and no velocity one, which is the case this exists for:
q_ref = arm.joint_positions() # ONCE, before the chunk is requested
for target in sequence_ai.to_joint_positions(chunk, q_ref, model):
arm.move_to_joint_positions(target)
Every row offsets the same q_ref. They are cumulative displacements, so re-reading the live
pose each step turns the chunk into an integrator and the arm overshoots. to_joint_positions()
refuses any model that does not declare a verified conversion rather than inventing one.
sequence_ai.recipes holds one module per model — pi0.5-DROID and Cosmos3-Edge — with the
worked loop for that model and a check() that compares it against the live catalogue.
for_model() returns None for the rest.
Three ways to drive it
From most control to least. They are the same request underneath; the difference is who owns the loop.
policy.predict(obs) # the whole chunk, you do everything
policy.next_action(obs) # we hold the buffer, you own the cadence
policy.run(observe=, act=) # we own the loop
policy.act("...", observe=, act=) # we own the loop and the instruction
act() is run() with the instruction pinned and a time bound required — the shape a VLM calls
as a tool. Everything run() accepts, act() accepts.
The other two endpoints: chat and perceive
Act is a loop; these are one-shot calls. Same key, same catalogue, same typed errors.
r = sequence_ai.chat("claude-opus-4-8", [{"role": "user", "content": "what should the arm do?"}])
r.content # the reply — no choices[0] wrapper
r.tool_calls # set instead when finish_reason == "tool_calls"
chat() is the slow brain: a sentence in, a sentence (or a tool call) out. It is answered by an
external provider with no cold start, so a 503 is an outage rather than a warm-up.
boxes = sequence_ai.detect("grounding-dino-base", frame, labels=["block", "bowl"])
vecs = sequence_ai.embed("siglip-so400m", images=[frame], text=["a mug on a plate"])
depth = sequence_ai.depth("depth-anything-v2-small", frame) # depth.maps[0] is a numpy array
detect/embed/depth are the three perception tasks. The function you call is the task, so
the shape is enforced by the signature: detect requires labels, embed needs images or text
or both, depth takes images. A wrong pairing — asking the embedder for depth — is refused
locally, before a request is billed, rather than coming back as a 400. Images go through the same
encode() the act path uses; a cold perceive worker's warming 503 is waited out.
Perception is not served yet: the platform answers 501 for all three models until it opens. The
functions and their local checks are here so code written against them keeps working when it does.
Benchmarks & eval
The other half of the SDK: instead of driving a robot, score a deployed policy on a set of
simulation tasks. You write a benchmark the way you write a policy — a decorated class — bridge its
observations and actions to the policy with an adapter, and call seq.eval(...). Before spending
GPUs you run it with dry_run=True and see the estimated wall-clock, the dollar cost, and how many
GPUs you will actually be granted.
This surface is heavier than the inference client (it needs pydantic and the shared contract
package), so it installs separately:
pip install 'sequence-ai[authoring]' # the `seq` CLI + @seq.benchmark / @seq.adapter
pip install -e /path/to/sequence-base # the shared wire contract (private sibling; editable)
export SEQUENCES_API_KEY=seq_test_... # identifies your tenant (owns deployments, billed for evals)
sequence-base is not on PyPI yet — install it editable from its checkout, or the seq CLI raises
ModuleNotFoundError: No module named 'sequence_base' on first import.
seq init toy-policy && seq init toy-reach && seq init toy-adapter # scaffold three runnable examples
seq deploy toy_policy:ToyReacher --name toy-reacher # deploy a @seq.policy → an eval id
seq register toy_sim:ToyReach # register a @seq.benchmark
seq list # discover the ids eval expects
The toy benchmark's 1-D action needs the adapter that bridges it to the policy's 8-D contract, and the adapter is the one eval knob the three-flag CLI does not carry — so the dry-run estimate and the full run go through Python:
import seq
from toy_policy import ToyReacher
from toy_sim import ToyReach
from toy_adapter import ToyReacherAdapter
model = seq.deploy(ToyReacher, name="toy-reacher") # returns the deployment id string
seq.register(ToyReach)
est = seq.eval(model, ToyReach, adapter=ToyReacherAdapter, gpus=8, dry_run=True)
est.gpus_granted, est.wall_clock_s, est.cost_usd # (8, 60.0, 0.096) — estimate, reserves nothing
report = seq.eval(model, ToyReach, adapter=ToyReacherAdapter, gpus=8)
report.metrics["success"].overall # "success" is YOUR metric key from @seq.check — no platform success_rate
report.metrics["success"].per_task # [(reach_left, 1.0, n=10), (reach_right, 1.0, n=10)]
report.cost_usd # actual dollars, reconciled from real GPU-seconds
report.dashboard_url # https://app.generalsequences.com/eval/eval_...
report.videos # mp4 pointers the platform recorder produced (if render= was declared)
report.adapter_lossy # every lossy transform, with its reason
The seq eval <model> --benchmark <name> --gpus N [--dry-run] CLI prints the same estimate/report for
any benchmark whose contract already matches the policy; when an adapter is required it refuses with
action space mismatch … provide adapter rather than dropping the knob silently.
More GPUs buy speed, not a bigger bill — conditions shard across the granted cards. If your gpus=
request exceeds min(account.max_gpus_per_eval, shared_pool_remaining), it is clamped and
report.clamped_by names which bound did it, never a silent reduction. If the policy and benchmark
disagree on obs/action shape and you pass no adapter, eval refuses up front with
action space mismatch … provide adapter rather than handing you a silently-wrong score.
The full reference — the benchmark lifecycle (@seq.setup / @seq.reset / @seq.check / @seq.score,
the platform-owned reset→step→check loop), assets= + render=, adapter loss labelling, billing,
composite (VLM) evals, and every seq.eval parameter — is in eval.md.
Safety
An action returned by any model is model output, not a safe robot command.
This library does not check joint limits, reachability, collisions, or whether a step is safe at the robot's current velocity. Bounds checking, a watchdog and an e-stop belong between this library and your motors.
policy.run(
observe=robot.read_observation,
act=robot.apply_action,
validate=robot.is_safe, # return False to stop the loop
hold=robot.hold, # called if it stops early, or if act() raises
max_actions=250,
)
max_actions is required and keyword-only. There is no run_forever() — an unbounded loop
that moves a robot should not be startable by accident.
validate=None is allowed for bench and simulation work, and warns once so it cannot happen
silently on real hardware.
Errors are typed
Because a controller reacts differently to each:
| meaning | what to do | |
|---|---|---|
AuthError |
key missing, revoked, expired | stop; retrying will not help |
OutOfCredit |
balance exhausted | stop and hold; top up |
InvalidRequest |
bad model, malformed observation, body too large | fix it; deterministic |
Unavailable |
upstream blip, or a cold worker | check .warming — see below |
ChunkExhausted |
asked for an action with an empty buffer and no observation | pass observation= every call |
Debugging a deployment
When a deployment misbehaves — some control requests slower than expected, or one that just failed — two read-only commands localize the fault and pull the record for a single request by its id:
seq doctor <deployment> # probe the direct-connect chain segment by segment
seq logs <deployment> # structured logs, connect + worker sinks merged, time-ordered
seq logs <deployment> --rid r-… # one request: worker phase joined to client timing, attributed
seq doctor reports each hop separately (initial-connect, network leg, cert-pin, worker /ping, /act)
and, when healthy, splits the round-trip into network vs. worker-handler so you can tell a slow
network from slow policy inference. A cold deployment reports WARMING with an eta rather than failing.
seq logs --rid retrieves a request that failed on the worker (surfacing its act.fail line) or one
that failed before reaching the worker (402/401/503/connect — attributed from the connect-side record).
Authorization is enforced server-side against your SEQUENCES_API_KEY; you only see your own tenant.
Against the platform both answer 503 today — the log and diagnostics backends are not connected
yet. The offline fixtures below work. For the account side, sequence-ai doctor checks the gateway,
your key, your credit and your account's GPU limits.
Both run offline against shipped samples — no account, network, or GPU needed:
seq logs --fixture list # demo, warming, worker-down
seq doctor pick-place-pi05 --fixture demo # healthy, with the RTT split
seq logs pick-place-pi05 --fixture demo --rid r-worker01 # pull the failed request
Full walkthrough, the segment catalog, the redaction guarantee, and the fixture schema for your own
offline samples are in observability.md.
Authoring & deploying your own policy
Everything above drives a model that already runs on the platform. To put your own policy up —
declare the environment it needs, point it at your weights, and have the platform validate that
environment — use the authoring surface. It is a second import (seq) and a second console command
(seq), separate from the inference client so a robot install stays httpx-only:
pip install 'sequence-ai[authoring]'
A policy is a class decorated with @seq.policy. You declare what it needs — the GPU, the image,
where the weights live — and mark its lifecycle methods; nothing is built or downloaded on your
machine, the declaration is pure data the gateway builds from.
import seq
@seq.policy(
gpu="L4", # the card offered today; a list is accepted, only its first card is used
image=(
seq.Image.debian_slim() # a concrete Debian base, pinned
.apt_install("git", "libgl1", "libglib2.0-0") # system libraries
.run_commands( # openpi installs from source: its PyPI name is a placeholder
"git clone https://github.com/Physical-Intelligence/openpi /opt/openpi",
"cd /opt/openpi && uv pip install --system -e .",
)
),
weights="gs://my-bucket/pi05-droid/", # your checkpoint, in a publicly readable bucket (today)
)
class Pi05(seq.Policy):
@seq.load
def load(self, weights_dir): # cold start: pull the checkpoint into VRAM
self.model = ...
@seq.infer
def infer(self, obs, *, seed=None): # per step: observation -> ActionChunk
return ...
The image is declarative
Every seq.Image method returns a new image and builds nothing — the chain is a spec, and the
platform builds it into a layer. Start from one base, then extend it; the step order you
write is the build order, one for one.
seq.Image.debian_slim() |
a small, pinned Debian base — the default when you declare no image |
seq.Image.from_registry("nvidia/cuda:12.4.1-runtime-ubuntu22.04") |
start from any registry tag, carried through verbatim |
seq.Image.from_dockerfile("Dockerfile") |
build from your own Dockerfile (read at declaration time) |
.apt_install(*pkgs) |
system libraries (apt-get install) |
.uv_pip_install(*pkgs) |
Python packages (uv pip install) |
.run_commands(*cmds) |
arbitrary build commands, run in order |
.env({...}) |
build-time environment variables |
@seq.policy(deps=[...]) is shorthand for a trailing .uv_pip_install(...), so a quick policy needs
no explicit image= at all. Build-time env carries ordinary variables only — a credential put there
is rejected, because the image is not a place for secrets (see below).
Weights live on a volume, never in the image
weights="hf://org/model" or weights="gs://…/" is copied once, at deploy, into a volume the platform
mounts at /data — never baked into an image layer (so a multi-GB checkpoint never bloats the build), and
every cold start reads that copy.
- Hugging Face —
hf://org/modelorhf://org/model@<revision>. At deploy the revision is resolved to a commit and exactly that commit is copied, so a re-deploy is deterministic. A private or gated repo needs a token: store it once (seq secret set hf --env HF_TOKEN) and declare it,secrets=[seq.Secret.from_name("hf")]. The copy runs with only that token; declaring it also handsHF_TOKENto your policy at runtime. A repo the platform cannot read — private without a token, gated, a wrong revision — is refused when you deploy, in seconds, with what to fix. - Your bucket — a
gs://bucket has to be publicly readable today. Private buckets and uploading a local checkpoint are coming.
The URI is validated as you write it: an empty bucket, a single-slash gs:/x, a .. traversal, or any
other scheme is rejected with a message that names the problem.
seq.Secret.from_name(name) references a credential by name; the plaintext never enters the SDK or
the deploy payload. Explicit volumes= (seq.Volume.from_name(...)) is not supported yet: a deploy
that declares one is refused rather than run without it.
Have the platform validate it
seq init pi05-droid # scaffold a working template to edit
seq validate pi05_droid:Pi05Droid # or: seq validate my_policy.py:Pi05
seq validate introspects the policy into the deployment request, validates it against the platform's
contract, and echoes back exactly what will be installed and where the weights mount — with nothing
built or pulled locally:
policy Pi05Droid
platform accepted (validated against deployment contract 0.3.0)
image debian:12-slim
apt-get install libgl1, libglib2.0-0
uv pip install openpi, jax[cuda12]
weights gs://general-sequences-models/pi05-droid/
→ mounted at /data (pulled by the platform onto a volume, not into the image)
When it reports accepted, hand the build to the gateway with seq deploy --remote (which needs
SEQUENCES_API_KEY and a reachable gateway). A malformed declaration fails here, at authoring time,
with a readable error and a non-zero exit — not as an opaque build failure minutes later.
Configuration
export SEQUENCES_API_KEY=seq_live_... # or pass api_key= to connect()
export SEQUENCES_BASE_URL=... # staging URL, or `mock` for the offline local tier
Get a key at app.generalsequences.com.
Secrets — bring your own keys (BYOK)
Some policies need a credential you own: a Hugging Face token for private hf:// weights, or your own
OpenRouter key for the VLM in a composite policy (we run the low-level policy on GPU and bill for that; the
VLM calls go to your OpenRouter account and never touch us). You store those keys once, by name, and refer
to them from a policy by that name alone — the key value never goes into your code, your logs, or any
output.
A secret is one name bound to a set of environment variables. The two canonical bundles are expressible by name alone:
seq secret set hf --env HF_TOKEN # private hf:// weights (and your policy's own use)
seq secret set openrouter --env OPENROUTER_API_KEY # your VLM key for a composite policy
Values are read from stdin, never the command line
There is no --env KEY=VALUE and no --value flag — a value never lands in argv, your shell history, or
the process listing. Pipe it in, or type it at the hidden prompt:
printf '%s' "$OPENROUTER_API_KEY" | seq secret set openrouter --env OPENROUTER_API_KEY
# or, interactively, you are prompted per key with the input hidden:
seq secret set openrouter --env OPENROUTER_API_KEY
# Value for OPENROUTER_API_KEY (input hidden):
A bundle with several env vars reads one line per --env, in order:
printf '%s\n%s\n' "$AWS_ID" "$AWS_SECRET" | seq secret set aws --env AWS_ACCESS_KEY_ID --env AWS_SECRET_ACCESS_KEY
List and remove — names only, never values
seq secret ls # name + env-var NAMES + timestamps; the value is never shown, ever
seq secret rm openrouter
ls is metadata only. There is no command that reads a stored value back — from the SDK a secret is
write-only. Re-running seq secret set <name> rotates the key (it replaces the whole bundle); no delete
first. On the live tier a rotated value goes live on the next worker start, or immediately with
seq restart <deployment>.
Use it from a policy — by name only
import seq
@seq.policy(gpu="L4", secrets=[seq.Secret.from_name("openrouter")])
class MyPolicy(seq.Policy):
...
Secret.from_name carries the name and nothing else — no value parameter exists, and it does no network I/O.
seq deploy checks at deploy time that the name exists and is yours.
Try it without an account — the mock tier
seq secret talks to the gateway, which needs an API key. To exercise the whole workflow offline with no
account and no network, point it at the mock tier — a local store that keeps only the metadata (name +
env-var names), never the value:
export SEQUENCES_BASE_URL=mock # or pass --base-url mock to each command
printf '%s' "$OPENROUTER_API_KEY" | seq secret set openrouter --env OPENROUTER_API_KEY
seq secret ls
# name env created last used
# openrouter OPENROUTER_API_KEY 2026-09-28T... never
The mock store lives at ~/.config/sequences/mock-secrets.json (override with SEQ_MOCK_SECRETS_PATH, or
mock:///abs/path.json). It never contains a secret value — only the names ls shows.
On the live tier, set/ls/rm each require your key (export SEQUENCES_API_KEY=seq_live_...) and are
scoped to it: another key never sees your secrets. A missing or invalid key fails closed — nothing is stored,
listed, or deleted.
Installing the authoring CLI
seq is the authoring surface and pulls pydantic (the robot-side sequence_ai client stays httpx-only). It
also needs the shared contract package sequence-base, which is a private sibling not published to PyPI, so
install it editable from its checkout:
pip install -e /path/to/sequence-base # shared wire contract (private; not on PyPI)
pip install -e '/path/to/sequence-sdk[authoring]'
If sequence_base is missing, seq tells you exactly this instead of a stack trace.
Install footprint
One dependency: httpx. Python 3.9+.
A robot controller is often on a Jetson with a pinned, fragile Python environment, and ROS 2 Humble ships Python 3.10. Every transitive dependency is another chance for the install to fail on the machine that actually matters.
Metadata
Release files for sequence-ai 0.14.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| sequence_ai-0.14.3.tar.gz | 470.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| sequence_ai-0.14.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 790.3 kB
Release files / sequence_ai-0.14.3.tar.gz
| Download URL | sequence_ai-0.14.3.tar.gz |
|---|---|
| Size | 470.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4f57d93548944179ca5e35e29e07b4c96653dac294718ba4d594ebeb9ce6db7f
|
|
BLAKE2b-256 checksum How to use checksums |
a088337d6015f279f4bd8a443f7c6eb471fa5a73d501db1eecb8af075f717e99
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency logRelease files / sequence_ai-0.14.3-py3-none-any.whl
| Download URL | sequence_ai-0.14.3-py3-none-any.whl |
|---|---|
| Size | 319.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
b44f07f4678d55225f5c9a8a2fb0085b776b7d85b1b41bbc681f2e3b8ab8f74c
|
|
BLAKE2b-256 checksum How to use checksums |
e8600c4b8a651bb3420432e7d95e0373eeda87ccc34898c1a4642f96de7c0594
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency log