idiotproof
Software for agents that create and edit video.
idiotproof is the IDEA ADK. ADK means Agent Development Kit: software
intended to be written and run by agents. It is designed for one-off Python scripts, CLI calls, and
headless agent workflows alongside FFmpeg. It is not an application framework
or a stable dependency for persistent software. Do not build a long-lived
service around its current interfaces; the package and dynamic tool schemas
will change as the product develops.
Bobbie plus FFmpeg is intended to be a complete headless media toolchain. Use FFmpeg for cutting, joining, transcoding, and audio. Use Bobbie for GPU-native vision and rendering.
Install and authenticate
pip install idiotproof
export IDEA_SECRET_KEY=idea_sk_test_...
The alpha client defaults to https://dev.67idea.com/v1. Override it with
IDEA_API_BASE_URL or Idea(base_url=...). A secret key belongs to one IDEA
project and environment. The API mints short-lived, workspace-scoped JWTs
internally; your code never creates or forwards those JWTs.
Required visual-review workflow
media_publish() returns signed URLs specifically so an agent can inspect its
work without downloading it. Signed URLs are transient credentials, not values
that must be ignored.
After publishing media, an agent should:
- Keep the signed URL only in the local viewer process's memory.
- Generate the viewer response dynamically. A page may contain several URLs to compare multiple edits, but no signed URL belongs in an on-disk HTML file or manifest.
- Serve only that directory with a server bound to
127.0.0.1, never0.0.0.0. - Open the localhost URL in the available browser. If browser control is unavailable, still create the viewer and provide its localhost URL.
- View the remote media through the page; do not download it merely to preview it.
- Do not print, paste into chat, persist, or commit the signed URL. Stop the viewer process after review or expiration.
Downloading is appropriate when saving an edit for continued local work or durable storage. It is not required for visual review.
Python
The public vocabulary is deliberately small: one Idea client, opaque
ChatId values, dynamic runtime tools, and structured IdeaError failures.
import asyncio
import json
from pathlib import Path
from idiotproof import Idea
async def main() -> None:
async with Idea() as idea:
chat_id = await idea.create_chat()
upload = await idea.upload_file(chat_id, "clip.mp4")
tools = await idea.get_tools()
print([definition["id"] for definition in tools])
# Functions and documentation come from the live MCP catalog.
print(tools.submit_bobbie_job.__doc__)
render_input = json.loads(Path("render.json").read_text())
render_input.setdefault("input_media", upload["workspace_uri"])
result = await tools.submit_bobbie_job(chat_id, **render_input)
published = await idea.media_publish(chat_id, result["output_media"])
# Give display_url directly to an in-memory viewer bound to 127.0.0.1.
# That viewer is agent workflow code, not an SDK side effect. Do not
# print the URL or write it into HTML, JSON, logs, or a manifest.
asyncio.run(main())
create_chat() returns a random opaque ChatId locally. The API lazily
materializes its project-scoped server workspace on the first upload or tool
call. Save the ID if another process must continue in the same workspace. Do
not put email addresses or other personal data in it.
upload_file(chat_id, path) probes video before network activity. An image or
video of at most 600 decoded frames follows the ordinary authorized TUS path
and returns one upload JSON object with workspace_uri. A longer video
automatically follows the ordered TUS metadata workflow described below and
returns an upload_asset JSON object with ordered parts and
workspace_uris; it never silently uploads a truncated single result.
For long videos or concurrent ordered uploads, use upload_batch() and
describe logical assets rather than a global upload queue:
uploads = await idea.upload_batch(
chat_id,
[
"long-video-a.mp4",
["video-b-001.mp4", "video-b-002.mp4"],
"reference.png",
],
concurrency=4,
)
Each top-level item is one logical asset. A bare path is one source asset; an
ordered path sequence supplies already separated pieces of one asset. Before
any request, the helper probes videos with ffprobe and uses FFmpeg to
re-encode every source longer than 600 frames into independently decodable
parts of at most 600 frames. It verifies that the parts preserve the complete
decoded frame count.
The helper flattens those parts in caller order and adds uploadBatchId,
uploadBatchIndex, and uploadBatchSize to each authenticated TUS creation.
The first creation reserves the complete contiguous block of upload_NNN
names; transfers may then proceed concurrently in any order. No reservation
endpoint or client-side completion throttle is involved. Results retain the
same nested asset/part order, and the helper verifies that returned basenames
form one contiguous block in that order. If an upload fails, it lets every
started operation settle and raises IdeaBatchError; its flat items retain
successes and failures in original asset/part traversal order.
media_publish(chat_id, media_path) calls the stable API-key-authenticated
media endpoint and returns short-lived viewing and download URLs. It is outside
the dynamic MCP catalog and does not depend on tool discovery.
download_media(chat_id, media_path, destination) is the durable counterpart.
It publishes internally, keeps the signed download URL in memory, verifies the
download, and atomically replaces the destination only after success. Small
objects and origins without safe range support use one streamed request. Large
objects use bounded parallel byte ranges only when the origin supplies a
content length, byte-range support, and a strong ETag. Range responses must
match their requested offsets and object identity; otherwise the helper safely
falls back or fails. resume=True verifies completed temporary ranges after
republishing the stable workspace path. Resume files never contain signed
URLs. In download_and_concat(), one shared transfer limit bounds publishing,
whole-object requests, and nested range requests so concurrency does not
multiply by the number of files.
get_tools() fetches each dynamic tool's name, description, and JSON Schema at
runtime. It returns an iterable catalog and generates documented async methods
such as tools.submit_bobbie_job(...). Punctuation becomes _; use
await tools.call(exact_name, chat_id, input_dict) for unusual or colliding
names. Tool methods return the tool output rather than a transport wrapper.
Every discovered tool is a callable RuntimeTool. Call one normally, bind
shared arguments once with partial(), or map it over varying inputs with
bounded concurrency:
render = tools.submit_bobbie_job.partial(
chat_id,
passes=effect,
timeout_seconds=300,
)
jobs = await render.map(
[
{"input_media": "workspace:segment-001.mp4", "output_media": "one.mp4"},
{"input_media": "workspace:segment-002.mp4", "output_media": "two.mp4"},
],
concurrency=4,
)
Results retain input order even when calls finish out of order. The entire
input iterable is validated before any request starts, and an input key may
not duplicate a bound/common key. tools.map(exact_name, chat_id, inputs) is
the escape hatch for exact MCP names. errors="collect" returns ordered
MapItem values; the default waits for every item to settle and then raises
IdeaBatchError, whose items preserve both successes and failures. Mapped
tool calls are not automatically retried because they may have side effects.
Cancelling map() prevents queued calls from starting and cancels local waits
for calls already in flight. It cannot retract a remote operation that the
service already accepted.
Use numbered_media() when a pipeline needs predictable ordered workspace or
output names:
from idiotproof import numbered_media
media = numbered_media(
"workspace:video_a_rendered_{index:03d}.mp4",
count=3,
)
# workspace:video_a_rendered_001.mp4, ...002.mp4, ...003.mp4
count must be non-negative, and the formatted values must be unique. Use a
stable sequence prefix when several source videos share a workspace. Generate
names before launching concurrent work and retain input order when collecting
results; completion time must never determine segment order. A workflow
manifest remains authoritative—zero-padding is convenient, not an ordering
guarantee. This helper is for predictable workspace and output names, never
media_publish() URLs: published URLs are signed and cannot be reconstructed
from a pattern.
Use download_and_concat() when ordered workspace videos are final and must
become one durable local video:
from idiotproof import download_and_concat, numbered_media
combined = await download_and_concat(
idea,
chat_id,
numbered_media(
"workspace:video_a_rendered_{index:03d}.mp4",
count=11,
),
output="output/final.mp4",
concurrency=4,
resume=True,
)
The input must be an ordered sequence; sets, mappings, and bare strings are
rejected. Downloads may complete in any order, but index-derived local names
and caller order control concatenation. Every part is probed with ffprobe.
All selected video streams must have exactly equal width and height—neither
mode scales, crops, pads, or rotates a mismatch. The default
concat_mode="copy" requires compatible streams and never re-encodes.
Explicit "encode" mode may normalize other stream properties while
preserving the dimension invariant. Supplying audio_source ignores segment
audio and remuxes that local audio once with stream copying. It preserves the
complete video sequence even when that audio ends slightly earlier. FFmpeg
and ffprobe must be installed. The result is a CombinedMedia containing the
final probe summary and SHA-256.
This is a durable-output helper, not the visual-review path. Iterate by
publishing a short representative clip and comparing remote display_url
variants through the localhost viewer. Download and concatenate only after an
effect is selected.
All individual failures are IdeaError; inspect code, status_code,
details, and request_id. Never log the secret key or upload token. Signed
media URLs may be retained transiently for the localhost viewer, but should
not appear in terminal history, chat messages, durable logs, or version
control.
CLI and MCP
Commands emit JSON so agents can compose them with scripts:
idea-adk create-chat
idea-adk upload-file "$CHAT_ID" clip.mp4
idea-adk get-tools
idea-adk call "$CHAT_ID" submit_bobbie_job --input @render.json
idea-adk map "$CHAT_ID" submit_bobbie_job --input @batch.json --concurrency 4
idea-adk media-publish "$CHAT_ID" workspace:result.mp4
batch.json contains an inputs array and an optional common object:
{
"common": {"passes": [{"kind": "fragment_shader", "shader_text": "..."}]},
"inputs": [
{"input_media": "workspace:one.mp4", "output_media": "one-out.mp4"},
{"input_media": "workspace:two.mp4", "output_media": "two-out.mp4"}
]
}
The CLI returns {"results": [...]} by default. With --errors collect, it
returns ordered items, each containing its index and exactly one of
result or structured error.
Use Python upload_batch() when a source may exceed 600 frames or the ADK must
reserve and upload a whole multi-asset manifest.
media-publish returns its requested JSON to stdout, including signed URLs;
do not use that command in a logged shell. A localhost review process should
call media_publish() internally and retain the response only in memory.
Run idea-adk mcp as a stdio MCP adapter and let it inherit
IDEA_SECRET_KEY. It exposes create_chat, upload_file, and get_tools,
followed by functions generated from the live server catalog. Media publishing
remains a separate stable Python/CLI operation and is intentionally not added
to the dynamic MCP catalog.
This README is the default advice agents should receive. Project-specific
instructions belong in AGENTS.md for Codex or CLAUDE.md for Claude Code.
Rules of thumb for agents
Uploads are currently normalized to exactly 600 frames at 3 megapixels,
slightly above 1080p. upload_file() detects a source beyond the input limit
and automatically creates independently decodable parts without dropping
source frames. Use upload_batch() directly for several logical assets or
when you want an explicitly nested result.
For a novel effect, first probe frame rate and time base and render one representative preview clip targeting about 20 seconds while staying safely below the service limit—590 frames is a useful cap when the maximum is 600. Use an explicit range when supplied; otherwise the source midpoint is a deterministic fallback, while a vision-capable agent may deliberately select a high-motion or representative region. Upload that preview once and map several effect variants over it. After choosing an effect, render the complete ordered segments once. Direct full-source rendering remains reasonable for a known or trivial effect.
Each TUS resource uses sequential, resumable chunks; core TUS requires ordered
offsets within that resource. For a long source, upload_file() delegates to
upload_batch(), which performs media-aware splitting before network activity,
reserves every output name in one transaction, and uploads those separate
resources concurrently. TUS chunks are transport details; they are not
independently decodable media segments. Upload completion order is never media
order.
Bobbie is a custom rendering engine that runs arbitrary GLSL. You describe the pipeline but do not control bindings; Bobbie assigns them programmatically. A pipeline can combine GLSL, smaller ML models, optimized CUDA kernels for classical computer vision, and optional Bayesian priors over color and shape. CUDA-GLSL interop keeps video on the GPU. Because Bobbie controls every tensor dimension, supported pipelines should not run out of GPU memory.
Invalid requests should fail before rendering. The expected render-time
failures are a video timeout and file too large. Incompressible output such as
raw static can exceed 100 MB; the service deletes it. Elaborate pipelines can
serve either image/video editing or advanced computer-vision work.
A separate tool extracts frames for a VLM. The current model is Qwen-VL 27B;
custom VLMs are not supported. Bobbie requires grounded objects and coordinates
normalized to the half-open interval [0, 1000).
Frame/VLM inspection plus Bobbie lets an agent move through any video. Use code to crop, zoom, and rotate. Use vision tools to measure camera and object motion and perform tracking, detection, and segmentation. Other catalog tools expose metadata about uploaded files, processes, and jobs.
IDEA deletes media aggressively. If an edit matters, download it for continued local editing or copy it immediately into durable storage such as S3.
Metadata
Release files for idiotproof 0.5.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| idiotproof-0.5.0.tar.gz | 74.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| idiotproof-0.5.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 110.0 kB
Release files / idiotproof-0.5.0.tar.gz
| Download URL | idiotproof-0.5.0.tar.gz |
|---|---|
| Size | 74.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1ba06596bd42051129f0177e9cc80e5514fbb5c9a837be7b5206c940c6b19216
|
|
BLAKE2b-256 checksum How to use checksums |
eee969bcbc377b617cdb6d01ae01589f8c4a8d6c734bb331ffdc71bc710db48b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 6, 2026.
Transparency logRelease files / idiotproof-0.5.0-py3-none-any.whl
| Download URL | idiotproof-0.5.0-py3-none-any.whl |
|---|---|
| Size | 35.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5df8b1c69ce690e69aeeea753c119da96c87e58212da9acf8b9ed8b6edd83c88
|
|
BLAKE2b-256 checksum How to use checksums |
9c7298ae0366666812941545acb617a374ff8831527888e56fad363af96bd10a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 6, 2026.
Transparency log