Skip to main content

gi-ingest CLI Guide

Offload robotics recordings from TF cards and deliver them to GILabs Data Platform.

Reads ego rig cards, copies each session to a local disk and verifies the copy, delivers the lot to GILabs as one sealed batch, and wipes the cards only once the server confirms it holds the data. Built for the way an offload actually goes: someone swapping cards at a reader for half an hour, a link pushing bytes overnight with nobody watching, and a wipe the next morning.

Installation

macOS and Linux. Windows is not supported: cards are looked for under /Volumes and /media, so a Windows install finds nothing.

Prerequisite — uv, once per machine. It fetches its own Python, so nothing else needs installing first.

curl -LsSf https://astral.sh/uv/install.sh | sh

Reopen your terminal afterwards so uv is on your PATH, then install the CLI into an isolated environment of its own:

uv tool install gi-ingest
gi-ingest --version

Authentication

Get a token from the vendor portal at vendor.gilabs.xyz → Upload tokens. It is shown once and stored only as a hash, so copy it when it appears — the portal prints the whole command below with your token already in it.

gi-ingest login --api-key gik_...
gi-ingest doctor

doctor checks the four things an overnight run depends on: your version, your token, that a delivery destination has been granted, and that the staging disk has room. Run it before a session rather than discovering a problem at 2am.

Configuration

Everything is stored per machine in ~/.config/gi-ingest/config.json, which also holds your token and is written owner-readable only.

gi-ingest config show                                   # settings and where they live
gi-ingest config set staging-dir /Volumes/BigDisk/gilabs
Setting What it is Default
staging-dir Where cards are copied to. Rarely somewhere a multi-terabyte round fits by default ~/gilabs-staging
project-id Which project deliveries are filed against. Set it with gi-ingest destinations --select unset
ledger-path The local record of what is staged, in flight and delivered ~/.local/state/gi-ingest/ledger.db
api-url The GILabs backend production

Any of them can be overridden for a single run without changing what is stored:

gi-ingest stage --staging /Volumes/OtherDisk/gilabs
gi-ingest stage --project <id>
gi-ingest upload --ledger /Volumes/OtherDisk/other.db

--ledger is how you keep two independent offload runs on one machine from sharing state — each needs its own, or they will see each other's queues.

The ledger is worth knowing about even though you rarely touch it. It is what makes an interrupted upload resumable and lets reclaim know which cards are safe to wipe, so it is the one file where losing it costs you a re-stage rather than nothing. It survives upgrades, including the rename from gilabs-ingest.

Upload a batch

One loop per round of cards: copy them, push them overnight, wipe them the next morning. Three commands rather than one because they run on two different timescales, and fusing them turns a 30-minute attended task into an all-night one.

Phase When What it does
stage attended, minutes per card Copies cards to a local disk and verifies each copy
upload unattended, hours Pushes everything staged as one delivery, overnight
reclaim attended, next morning Wipes only what the server confirmed
# 1. Drain every card. Swap cards as each finishes; nothing is wiped yet.
gi-ingest stage
gi-ingest stage --watch                 # or let it auto-copy on insert

# 2. Leave this running. One delivery, resumable, safe to nohup.
gi-ingest upload --max-bandwidth 50M

# 3. Next morning, with the cards back in the reader.
gi-ingest reclaim

1 — stage copies each session to the staging disk, checksums it on the way back out, and tells you which project the round is bound for before it starts. The card is left completely intact, so it can be pulled the moment a copy finishes.

2 — upload opens a batch and receives AWS credentials scoped to that batch's folder alone; they cannot read or write anything else, including your own other deliveries. Files go straight to S3, and the credentials are re-minted as they expire, so a multi-hour run and a laptop that sleeps both just work. It then seals the batch with a manifest listing every episode, part, size and SHA-256, and GILabs verifies every declared file is present at its declared size before ingesting anything — a half-finished upload can never be processed as if it were complete.

3 — reclaim wipes only what the server has confirmed, checking each card's own metadata against its records first. It clears the staging copy at the same time, which is the one that actually fills a disk.

Commands

gi-ingest --version       # what you are running
gi-ingest whoami          # vendor, scopes, and where you may deliver
gi-ingest destinations    # the org/project list, with ids for --project
gi-ingest doctor          # version, token, destination, disk
gi-ingest queue           # what is still in play — staged, in flight, failed
gi-ingest queue --card ID # ...and that card's session names
gi-ingest history         # past deliveries
gi-ingest history --card ID
gi-ingest status          # deliveries and their QC verdicts
gi-ingest status --batch btch_...
gi-ingest retry           # re-queue anything that failed
gi-ingest reset --card ... # send staged episodes back a step (see below)

FAQ

Who recorded this, and where?

At the end of each card, stage asks — once per recording date, since a day's work is normally one place and one person:

2 session(s) on this card have no operator or environment.
Environment (name or id, blank to skip): warehouse
  env_00003  Manila warehouse · warehouse
Operator (name or id, blank to skip): maria
  operator_0001  Maria Santos · collector
Attributed 2 session(s) to env_00003, operator operator_0001.

Type any part of the name — the place, the person — and pick from what matches. The ids come from the environments and operators your team registered in the vendor portal; register them there first if the list is empty.

A card holding sessions from more than one date is asked about once per date, and says so before the first prompt. Each date is a separate answer: skipping one does not skip the others. Two sittings on the same date share an answer — a morning and an evening in the same place is the common case, and splitting it would ask twice for one answer.

Both prompts can be skipped with a blank line, and a stage with no terminal (--watch left running) never asks. The answers are written into each staged session's session.json, which is where the platform reads them from — the card itself is never modified.

Why was a session marked unusable?

It was shorter than the minimum the destination project sets. stage measures each recording before copying it — reading the index at the tail of each .mcap part, not the footage — and a session under that minimum is recorded as unusable and left on the card rather than copied and uploaded.

The minimum is per project, set by a project admin in the data centre, and defaults to 0: unless someone has set one, nothing is refused. stage prints the project's minimum when it starts if there is one.

It is not uploaded, and reclaim asks before wiping it:

3 session(s) on these cards were refused as shorter than 30s (412.0 MiB).
They were never uploaded, so wiping them deletes the only copy.
Wipe them too? [y/N]

Pressing Enter keeps them, and a reclaim with no terminal — scheduled, or piped — always keeps them. --wipe-unusable / --keep-unusable answers ahead of time. If you disagree with the refusal itself, gi-ingest reset --card <id> --to staged puts the session back in the upload queue instead.

A session whose length cannot be measured is not refused. A .umi session has no index to read, and a recording interrupted by a power cut has lost its — withholding real footage over a missing few kilobytes is the more expensive mistake.

Can I redo a card I have already staged?

Yes — gi-ingest reset:

gi-ingest queue                                   # lists the card ids
gi-ingest queue --card F0EB8630                   # ...and that card's session names
gi-ingest reset --card A1B2-C3D4                  # forget the copy; stage it again from the card
gi-ingest reset --card A1B2-C3D4 --to staged      # keep the copy; upload it again
gi-ingest reset --session session_00042 --dry-run # one session, no changes made

The card id is the volume UUID, which is not written on the card — gi-ingest queue lists it next to the volume label, which is what you actually recognise the card by. gi-ingest queue --card <id> then names that card's sessions.

Both --card and --session take a unique prefix, so --card F0EB8630 and --session session_00008 are usually enough. If a prefix matches two, the candidates are listed and nothing is changed.

queue shows only what is still in play, so a card that has already been delivered will not appear there — gi-ingest history --card <id> has it. That is also the answer when reset says "already delivered": the episode is the server's now, and there is nothing local left to redo.

--to discovered (the default) deletes the staged copy so stage takes it off the card again — use it when you have reason to distrust the copy, and only while the card still holds the footage. It tells you how much it is about to delete and asks before doing it; --dry-run shows the same without changing anything. --to staged keeps the copy and lets upload send it once more.

It will not touch an episode the server has already confirmed, and it will not send one back to discovered after its card has been reclaimed: in both cases the thing it would send you back to no longer exists. Failures do not need it — gi-ingest retry re-queues those.

The upload was interrupted — do I start again?

No. Re-run gi-ingest upload. It re-uses the same open delivery and skips files already sent, so a restart costs a listing rather than the bytes. That covers a kill, a crash, a dropped connection or a closed lid.

Why is wiping the cards a separate step?

upload finishes hours after the cards were pulled, so it cannot wipe them — they are back in the rig by then.

It also means the card keeps a second copy until reclaim runs, which is deliberate: between stage and confirmation the staging disk would otherwise be the sole home of a whole collection round.

Where do the cards get copied to?

~/gilabs-staging by default, which is rarely where a multi-terabyte round fits. Point it at the right disk once — see Configuration.

Which project does a delivery go to?

Every delivery is filed against one project, and destinations are granted by GILabs — you cannot add one yourself. If the portal's Destinations tab is empty, ask your GILabs contact.

With one granted destination there is nothing to choose and no flag to pass. With several, you have to say which, before staging rather than at upload:

gi-ingest destinations --select     # pick from a list; saved for this machine
gi-ingest destinations              # just show them
gi-ingest stage --project <id>      # or choose for a single run

You will also be asked directly the first time stage needs to know and the answer is not obvious. A project id is a uuid, and retyping one off a portal page is how a delivery ends up under the wrong project.

stage refuses to copy anything until this is settled, and both stage and upload print the destination before they start. That is deliberate rather than fussy: a delivery goes to one project, so cards staged for two of them in the same session would be filed together and only one of them correctly — and nothing about the upload would look wrong while it happened.

For the same reason upload refuses a queue holding episodes staged for different projects. Upload one destination's worth at a time.

What happens if a card was already delivered by someone else?

stage asks GILabs which episodes it already has before copying, so a card a colleague already delivered costs one request instead of its bytes. Offline, it stages anyway and says so; the duplicate is caught server-side at seal.

Upgrading

uv tool install --upgrade gi-ingest

Worth doing when doctor says so. Older versions are refused for upload and reclaim, because some of them lost data without reporting it — a card wiped against another card's confirmation, and an interrupted upload that could never be sealed.

Getting help

gi-ingest doctor first: it answers most questions, and its output is the useful thing to send on. Then your GILabs contact, with the version it printed.

Release files for gi-ingest 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for gi-ingest 0.3.0
File Size Uploaded
gi_ingest-0.3.0.tar.gz 150.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for gi-ingest 0.3.0
File Interpreter ABI Platform
gi_ingest-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 214.5 kB

Release files / gi_ingest-0.3.0.tar.gz

Download URL gi_ingest-0.3.0.tar.gz
Size 150.4 kB
Tags Source
SHA-256 checksum
How to use checksums
3009277fb5a887deb012a477b90d7783c914bff869a3cf2e3397bc4cf4d23158
BLAKE2b-256 checksum
How to use checksums
881bc665a2009bd4ebdd6a0ac2bbfaf336abf5a4e377f6dd36649d9f7b85816a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 21, 2026.

Transparency log

Release files / gi_ingest-0.3.0-py3-none-any.whl

Download URL gi_ingest-0.3.0-py3-none-any.whl
Size 64.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
58616c89cb2452d10ea1f2ad0b917edd1328c0c603fb1733d0e48c439b26f907
BLAKE2b-256 checksum
How to use checksums
c447573d901403c6c79108cba5388671151d6eaa17f3e4d791ab5589d232b222
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 21, 2026.

Transparency log

Release history Release notifications | RSS feed

0.9.0

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.2

2 release files

0.3.1

2 release files

This release

0.3.0 This release

2 release files

0.2.1

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page