Skip to main content

idiotproof

Software for agents that create and edit video.

idiotproof is the IDEA ADK. ADK means Agent Development Kit: software intended to be written and run by agents. It is designed for one-off Python scripts, CLI calls, and headless agent workflows alongside FFmpeg. It is not an application framework or a stable dependency for persistent software. Do not build a long-lived service around its current interfaces; the package and dynamic tool schemas will change as the product develops.

Bobbie plus FFmpeg is intended to be a complete headless media toolchain. Use FFmpeg for cutting, joining, transcoding, and audio. Use Bobbie for GPU-native vision and rendering.

Install and authenticate

pip install idiotproof
export IDEA_SECRET_KEY=idea_sk_test_...

The alpha client defaults to https://dev.67idea.com/v1. Override it with IDEA_API_BASE_URL or Idea(base_url=...). A secret key belongs to one IDEA project and environment. The API mints short-lived, workspace-scoped JWTs internally; your code never creates or forwards those JWTs.

Required visual-review workflow

media_publish() returns signed URLs specifically so an agent can inspect its work without downloading it. Signed URLs are transient credentials, not values that must be ignored.

After publishing media, an agent should:

  1. Keep the signed URL only in the local viewer process's memory.
  2. Generate the viewer response dynamically. A page may contain several URLs to compare multiple edits, but no signed URL belongs in an on-disk HTML file or manifest.
  3. Serve only that directory with a server bound to 127.0.0.1, never 0.0.0.0.
  4. Open the localhost URL in the available browser. If browser control is unavailable, still create the viewer and provide its localhost URL.
  5. View the remote media through the page; do not download it merely to preview it.
  6. Do not print, paste into chat, persist, or commit the signed URL. Stop the viewer process after review or expiration.

Downloading is appropriate when saving an edit for continued local work or durable storage. It is not required for visual review.

Python

The public vocabulary is deliberately small: one Idea client, opaque ChatId values, dynamic runtime tools, and structured IdeaError failures.

import asyncio
import json
from pathlib import Path

from idiotproof import Idea


async def main() -> None:
    async with Idea() as idea:
        chat_id = await idea.create_chat()
        upload = await idea.upload_file(chat_id, "clip.mp4")

        tools = await idea.get_tools()
        print([definition["id"] for definition in tools])

        # Functions and documentation come from the live MCP catalog.
        print(tools.submit_bobbie_job.__doc__)
        render_input = json.loads(Path("render.json").read_text())
        render_input.setdefault("input_media", upload["workspace_uri"])
        result = await tools.submit_bobbie_job(chat_id, **render_input)

        published = await idea.media_publish(chat_id, result["output_media"])
        # Give display_url directly to an in-memory viewer bound to 127.0.0.1.
        # That viewer is agent workflow code, not an SDK side effect. Do not
        # print the URL or write it into HTML, JSON, logs, or a manifest.


asyncio.run(main())

create_chat() returns a random opaque ChatId locally. The API lazily materializes its project-scoped server workspace on the first upload or tool call. Save the ID if another process must continue in the same workspace. Do not put email addresses or other personal data in it.

upload_file(chat_id, path) probes video before network activity. An image or video of at most 600 decoded frames follows the ordinary authorized TUS path and returns one upload JSON object with workspace_uri. A longer video automatically follows the ordered TUS metadata workflow described below and returns an upload_asset JSON object with ordered parts and workspace_uris; it never silently uploads a truncated single result.

For long videos or concurrent ordered uploads, use upload_batch() and describe logical assets rather than a global upload queue:

uploads = await idea.upload_batch(
    chat_id,
    [
        "long-video-a.mp4",
        ["video-b-001.mp4", "video-b-002.mp4"],
        "reference.png",
    ],
    concurrency=4,
)

Each top-level item is one logical asset. A bare path is one source asset; an ordered path sequence supplies already separated pieces of one asset. Before any request, the helper probes videos with ffprobe and uses FFmpeg to re-encode every source longer than 600 frames into independently decodable parts of at most 600 frames. It verifies that the parts preserve the complete decoded frame count.

The helper flattens those parts in caller order and adds uploadBatchId, uploadBatchIndex, and uploadBatchSize to each authenticated TUS creation. The first creation reserves the complete contiguous block of upload_NNN names; transfers may then proceed concurrently in any order. No reservation endpoint or client-side completion throttle is involved. Results retain the same nested asset/part order, and the helper verifies that returned basenames form one contiguous block in that order. If an upload fails, it lets every started operation settle and raises IdeaBatchError; its flat items retain successes and failures in original asset/part traversal order.

media_publish(chat_id, media_path) calls the stable API-key-authenticated media endpoint and returns short-lived viewing and download URLs. It is outside the dynamic MCP catalog and does not depend on tool discovery.

download_media(chat_id, media_path, destination) is the durable counterpart. It publishes internally, keeps the signed download URL in memory, verifies the download, and atomically replaces the destination only after success. Small objects and origins without safe range support use one streamed request. Large objects use bounded parallel byte ranges only when the origin supplies a content length, byte-range support, and a strong ETag. Range responses must match their requested offsets and object identity; otherwise the helper safely falls back or fails. resume=True verifies completed temporary ranges after republishing the stable workspace path. Resume files never contain signed URLs. In download_and_concat(), one shared transfer limit bounds publishing, whole-object requests, and nested range requests so concurrency does not multiply by the number of files.

get_tools() fetches each dynamic tool's name, description, and JSON Schema at runtime. It returns an iterable catalog and generates documented async methods such as tools.submit_bobbie_job(...). Punctuation becomes _; use await tools.call(exact_name, chat_id, input_dict) for unusual or colliding names. Tool methods return the tool output rather than a transport wrapper.

Every discovered tool is a callable RuntimeTool. Call one normally, bind shared arguments once with partial(), or map it over varying inputs with bounded concurrency:

render = tools.submit_bobbie_job.partial(
    chat_id,
    passes=effect,
    timeout_seconds=300,
)
jobs = await render.map(
    [
        {"input_media": "workspace:segment-001.mp4", "output_media": "one.mp4"},
        {"input_media": "workspace:segment-002.mp4", "output_media": "two.mp4"},
    ],
    concurrency=4,
)

Results retain input order even when calls finish out of order. The entire input iterable is validated before any request starts, and an input key may not duplicate a bound/common key. tools.map(exact_name, chat_id, inputs) is the escape hatch for exact MCP names. errors="collect" returns ordered MapItem values; the default waits for every item to settle and then raises IdeaBatchError, whose items preserve both successes and failures. Mapped tool calls are not automatically retried because they may have side effects.

Cancelling map() prevents queued calls from starting and cancels local waits for calls already in flight. It cannot retract a remote operation that the service already accepted.

Use numbered_media() when a pipeline needs predictable ordered workspace or output names:

from idiotproof import numbered_media

media = numbered_media(
    "workspace:video_a_rendered_{index:03d}.mp4",
    count=3,
)
# workspace:video_a_rendered_001.mp4, ...002.mp4, ...003.mp4

count must be non-negative, and the formatted values must be unique. Use a stable sequence prefix when several source videos share a workspace. Generate names before launching concurrent work and retain input order when collecting results; completion time must never determine segment order. A workflow manifest remains authoritative—zero-padding is convenient, not an ordering guarantee. This helper is for predictable workspace and output names, never media_publish() URLs: published URLs are signed and cannot be reconstructed from a pattern.

Use download_and_concat() when ordered workspace videos are final and must become one durable local video:

from idiotproof import download_and_concat, numbered_media

combined = await download_and_concat(
    idea,
    chat_id,
    numbered_media(
        "workspace:video_a_rendered_{index:03d}.mp4",
        count=11,
    ),
    output="output/final.mp4",
    concurrency=4,
    resume=True,
)

The input must be an ordered sequence; sets, mappings, and bare strings are rejected. Downloads may complete in any order, but index-derived local names and caller order control concatenation. Every part is probed with ffprobe. All selected video streams must have exactly equal width and height—neither mode scales, crops, pads, or rotates a mismatch. The default concat_mode="copy" requires compatible streams and never re-encodes. Explicit "encode" mode may normalize other stream properties while preserving the dimension invariant. Supplying audio_source ignores segment audio and remuxes that local audio once with stream copying. It preserves the complete video sequence even when that audio ends slightly earlier. FFmpeg and ffprobe must be installed. The result is a CombinedMedia containing the final probe summary and SHA-256.

This is a durable-output helper, not the visual-review path. Iterate by publishing a short representative clip and comparing remote display_url variants through the localhost viewer. Download and concatenate only after an effect is selected.

All individual failures are IdeaError; inspect code, status_code, details, and request_id. Never log the secret key or upload token. Signed media URLs may be retained transiently for the localhost viewer, but should not appear in terminal history, chat messages, durable logs, or version control.

CLI and MCP

Commands emit JSON so agents can compose them with scripts:

idea-adk create-chat
idea-adk upload-file "$CHAT_ID" clip.mp4
idea-adk get-tools
idea-adk call "$CHAT_ID" submit_bobbie_job --input @render.json
idea-adk map "$CHAT_ID" submit_bobbie_job --input @batch.json --concurrency 4
idea-adk media-publish "$CHAT_ID" workspace:result.mp4

batch.json contains an inputs array and an optional common object:

{
  "common": {"passes": [{"kind": "fragment_shader", "shader_text": "..."}]},
  "inputs": [
    {"input_media": "workspace:one.mp4", "output_media": "one-out.mp4"},
    {"input_media": "workspace:two.mp4", "output_media": "two-out.mp4"}
  ]
}

The CLI returns {"results": [...]} by default. With --errors collect, it returns ordered items, each containing its index and exactly one of result or structured error.

Use Python upload_batch() when a source may exceed 600 frames or the ADK must reserve and upload a whole multi-asset manifest.

media-publish returns its requested JSON to stdout, including signed URLs; do not use that command in a logged shell. A localhost review process should call media_publish() internally and retain the response only in memory.

Run idea-adk mcp as a stdio MCP adapter and let it inherit IDEA_SECRET_KEY. It exposes create_chat, upload_file, and get_tools, followed by functions generated from the live server catalog. Media publishing remains a separate stable Python/CLI operation and is intentionally not added to the dynamic MCP catalog.

This README is the default advice agents should receive. Project-specific instructions belong in AGENTS.md for Codex or CLAUDE.md for Claude Code.

Rules of thumb for agents

Uploads are currently normalized to exactly 600 frames at 3 megapixels, slightly above 1080p. upload_file() detects a source beyond the input limit and automatically creates independently decodable parts without dropping source frames. Use upload_batch() directly for several logical assets or when you want an explicitly nested result.

For a novel effect, first probe frame rate and time base and render one representative preview clip targeting about 20 seconds while staying safely below the service limit—590 frames is a useful cap when the maximum is 600. Use an explicit range when supplied; otherwise the source midpoint is a deterministic fallback, while a vision-capable agent may deliberately select a high-motion or representative region. Upload that preview once and map several effect variants over it. After choosing an effect, render the complete ordered segments once. Direct full-source rendering remains reasonable for a known or trivial effect.

Each TUS resource uses sequential, resumable chunks; core TUS requires ordered offsets within that resource. For a long source, upload_file() delegates to upload_batch(), which performs media-aware splitting before network activity, reserves every output name in one transaction, and uploads those separate resources concurrently. TUS chunks are transport details; they are not independently decodable media segments. Upload completion order is never media order.

Bobbie is a custom rendering engine that runs arbitrary GLSL. You describe the pipeline but do not control bindings; Bobbie assigns them programmatically. A pipeline can combine GLSL, smaller ML models, optimized CUDA kernels for classical computer vision, and optional Bayesian priors over color and shape. CUDA-GLSL interop keeps video on the GPU. Because Bobbie controls every tensor dimension, supported pipelines should not run out of GPU memory.

Invalid requests should fail before rendering. The expected render-time failures are a video timeout and file too large. Incompressible output such as raw static can exceed 100 MB; the service deletes it. Elaborate pipelines can serve either image/video editing or advanced computer-vision work.

A separate tool extracts frames for a VLM. The current model is Qwen-VL 27B; custom VLMs are not supported. Bobbie requires grounded objects and coordinates normalized to the half-open interval [0, 1000).

Frame/VLM inspection plus Bobbie lets an agent move through any video. Use code to crop, zoom, and rotate. Use vision tools to measure camera and object motion and perform tracking, detection, and segmentation. Other catalog tools expose metadata about uploaded files, processes, and jobs.

IDEA deletes media aggressively. If an edit matters, download it for continued local editing or copy it immediately into durable storage such as S3.

Metadata

Release files for idiotproof 0.5.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for idiotproof 0.5.0
File Size Uploaded
idiotproof-0.5.0.tar.gz 74.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for idiotproof 0.5.0
File Interpreter ABI Platform
idiotproof-0.5.0-py3-none-any.whl Python 3 none any Details

Total release size: 110.0 kB

Release files / idiotproof-0.5.0.tar.gz

Download URL idiotproof-0.5.0.tar.gz
Size 74.9 kB
Tags Source
SHA-256 checksum
How to use checksums
1ba06596bd42051129f0177e9cc80e5514fbb5c9a837be7b5206c940c6b19216
BLAKE2b-256 checksum
How to use checksums
eee969bcbc377b617cdb6d01ae01589f8c4a8d6c734bb331ffdc71bc710db48b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 6, 2026.

Transparency log

Release files / idiotproof-0.5.0-py3-none-any.whl

Download URL idiotproof-0.5.0-py3-none-any.whl
Size 35.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5df8b1c69ce690e69aeeea753c119da96c87e58212da9acf8b9ed8b6edd83c88
BLAKE2b-256 checksum
How to use checksums
9c7298ae0366666812941545acb617a374ff8831527888e56fad363af96bd10a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 6, 2026.

Transparency log

Release history Release notifications | RSS feed

6.5.1

2 release files

6.5.0

2 release files

0.6.4

2 release files

0.6.3

2 release files

0.6.2

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.14

2 release files

0.5.13

2 release files

0.5.12

2 release files

0.5.11

2 release files

0.5.9

2 release files

0.5.8

2 release files

0.5.7

2 release files

0.5.6

2 release files

0.5.5

2 release files

0.5.4

2 release files

0.5.3

2 release files

0.5.1

2 release files

This release

0.5.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page