venice-cli
A Python CLI wrapping the Venice.ai API. The base is
stdlib-only; the optional venice chat and venice embed commands use the
official OpenAI SDK (Venice is OpenAI-compatible).
pip install venice-cli
Unofficial. This is an independent, community-maintained client. It is not affiliated with, endorsed by, or supported by Venice.ai. "Venice" and "Venice.ai" belong to their respective owners. For official support, see venice.ai.
Ships working venice login, venice sfx (sound-effect generation),
venice music (long-form ambience/music), venice video (text/image-to-video),
venice tts (text-to-speech), venice image (image generation),
venice upscale / venice bg-remove (image post-processing),
venice master (audio mastering), venice contact-sheet (montage grids of
generated images), venice chat (one-shot or interactive chat completions with
Venice extensions), venice embed (text embeddings), venice index /
venice search (project semantic search), venice balance (budget
tracking), and venice models (catalog browser).
Install
pip install venice-cli # base: stdlib-only, no dependencies
pip install "venice-cli[openai]" # + venice chat / venice embed
The distribution is named venice-cli, but the command is venice (and the
import package is venice). pipx install "venice-cli[openai]" works too, and
keeps the CLI out of your system site-packages.
Dependencies
The base install pulls in nothing — every command is stdlib-only except
venice chat and venice embed, which use the official OpenAI SDK against
Venice's OpenAI-compatible API, and venice mcp-serve / venice chat --mcp,
which use the MCP SDK. Those SDKs are lazy-imported, so they ship as optional
extras rather than hard requirements: if you only generate images or audio, you
don't pay for them. Without the relevant extra, that command exits 2 with a hint;
every other command works normally.
Extras are per-feature and additive, so the pattern holds as the CLI grows:
| Install | Enables |
|---|---|
venice-cli |
everything except chat/embed and the MCP commands |
venice-cli[openai] |
venice chat, venice embed |
venice-cli[mcp] |
venice mcp-serve (MCP server) and venice chat --mcp (MCP client); needs Python ≥ 3.10 |
venice-cli[all] |
every extra (openai + mcp) |
The [mcp] extra pulls in the mcp SDK, which
requires Python ≥ 3.10. The base CLI still supports 3.9 — on 3.9 the extra
resolves to nothing and only venice mcp-serve and venice chat --mcp are
unavailable.
Some commands shell out to external binaries when present: venice master and
venice contact-sheet use ffmpeg/ffprobe (and ImageMagick's montage if
available); audio playback uses mpg123, ffplay, or paplay. These are
detected at runtime — nothing breaks if they're missing.
From source (development)
Clone anywhere; no install is needed to run it:
git clone https://github.com/gobha-me/venice-cli.git
cd venice-cli
PYTHONPATH=src python3 -m venice --help
For an editable install: pip install -e ".[openai]".
Alternatively ./install.sh puts venice on your PATH without pip, by creating
two symlinks:
~/.local/bin/venice-><repo>/bin/venice~/.local/lib/venice-><repo>/src/venice
and ~/.config/venice/ (mode 0700) for the credentials file. The installer
resolves the repo path itself, so the clone can live wherever you like.
~/.local/bin should be on your PATH.
Don't mix pip and
./install.sh. Both own~/.local/bin/venice. If pip got there first,install.shrefuses to clobber the real file and exits 1. Ifinstall.shgot there first, pip silently replaces the symlink — your repo edits stop taking effect with no error. Pick one;pip uninstall venice-clior./uninstall.shto back the other out.
Shell completion
venice completion [bash|zsh] prints a tab-completion script to stdout, generated
by introspecting the CLI so it never drifts from the real commands and flags:
# bash -- current shell:
source <(venice completion bash)
# bash -- persistent (per-user):
venice completion bash > ~/.local/share/bash-completion/completions/venice
# zsh -- drop it somewhere on your $fpath, named `_venice`:
venice completion zsh > "${fpath[1]}/_venice" # then restart zsh
./install.sh writes the bash script for you (source installs only); a
pip install uses the source <(...) line above. Completion covers subcommands,
the config/secret nested actions (with their aliases), and each command's
flags.
First-time setup
venice login
You'll be prompted (hidden input) for your API key from
https://venice.ai/settings/api. The key is stored at
~/.config/venice/credentials with mode 0600.
$VENICE_API_KEY in the environment overrides the file.
Balance and budget tracking
Venice has up to four spendable buckets, drained in this order:
- DIEM allowance — daily credit derived from staked DIEM tokens
(1 DIEM staked = $1/day). Resets every 24h at
nextEpochBegins. Per-epoch use-it-or-lose-it. - Monthly credit (BUNDLED_CREDITS) — bundle granted with paid subscriptions; drains before cash.
- VCU — Venice Compute Units, per-tier inclusions.
- USD cash — one-and-done prepaid USD balance.
1 unit of any bucket == $1 of purchasing power (per-model pricing in
/models lists the same number in both usd and diem fields).
Inference-key visibility: this CLI reads USD + DIEM and the
epoch reset time. The monthly-bundle and VCU balances live behind
admin-key endpoints (/billing/balance, 401 with inference keys), so
the CLI documents them but can't show their values. If you have
monthly credit, the actual debit lands there before USD cash, so the
"After charge" line on the USD cash side may be slightly pessimistic.
venice balance # -> $32.70 USD (combined visible total)
venice balance --verbose # buckets, epoch reset, spend order
venice balance --json # incl. total_usd_equiv, spend_order, notes
venice balance --min 5 # exit 1 if total < $5 (useful in scripts)
venice sfx and venice tts print the balance line inline next to the
cost quote, with DIEM listed first to mirror the drain order:
Balance: $32.70 USD (6.56 DIEM allowance + 26.14 USD cash)
After charge: $32.69 USD
Suppress with --no-balance, or set defaults.no_balance once to suppress it
everywhere (--show-balance re-enables it for a single run). Hard-cap a single call
with --max-spend USD (refuses to queue / synthesize if the estimate exceeds
the cap).
Sound effects
# Quote only -- no charge, no audio. Shows balance + estimated remaining.
venice sfx "thunderstorm rolling in" --duration 8 --dry-run
# Generate, confirm cost, save to ./venice-sfx-<id>.mp3.
venice sfx "soft chime" --duration 2
# Auto-confirm, custom output path, no playback, hard budget cap.
venice sfx "rain on tin roof" --duration 4 --yes --max-spend 0.05 \
-o /tmp/rain.mp3 --no-play
# Background: prints queue_id to stdout, fetch later.
ID=$(venice sfx "ocean waves" --duration 10 --yes --background)
venice sfx-status "$ID" -o /tmp/ocean.mp3
SFX Models
| slug | max duration |
|---|---|
elevenlabs-sound-effects-v2 (default) |
22 s |
mmaudio-v2-text-to-audio |
30 s |
Durations longer than the model max are clamped (warning on stderr).
Text-to-speech
# Positional text, auto-confirm, sub-cent cap.
venice tts "Hello from Venice." --yes --max-spend 0.01
# Read input from a file.
venice tts --from-file speech.txt --yes -o out.mp3
# Read input from stdin (pipe-friendly).
cat speech.txt | venice tts --stdin --yes -o out.mp3
echo "quick line" | venice tts --stdin --yes
# Specific voice and WAV output.
venice tts "Sky voice in wav." --voice af_sky --format wav --yes -o sky.wav
# Different model (e.g. ElevenLabs Turbo for higher quality).
venice tts "Demo line." --model tts-elevenlabs-turbo-v2-5 --voice <id> --yes
# Speed control (0.25-4.0).
venice tts "Fast talker." --speed 1.4 --yes
# Dry-run shows estimated cost + balance without spending.
venice tts "How much will this cost?" --dry-run
TTS models and pricing (per 1M characters)
| slug | price | voices |
|---|---|---|
tts-kokoro (default) |
$3.50 | 54 |
tts-inworld-1-5-max |
$12.50 | 14 |
tts-xai-v1 |
$18.75 | 5 |
tts-chatterbox-hd |
$50.00 | 9 |
tts-orpheus |
$62.50 | 8 |
tts-elevenlabs-turbo-v2-5 |
$62.50 | 21 |
tts-qwen3-0-6b |
$87.50 | 9 |
tts-qwen3-1-7b |
$112.50 | 9 |
tts-minimax-speech-02-hd |
$125.00 | 15 |
tts-gemini-3-1-flash |
$187.50 | 30 |
To see the voice list for any TTS model:
venice models tts-kokoro | jq '.model_spec.voices'
If --voice is omitted Venice uses each model's built-in default.
Formats supported: mp3 (default), opus, aac, flac, wav, pcm.
Image generation
Sync POST /image/generate, same budget rails as sfx/tts. Default
output is PNG (lossless, good for card art and upscaling).
# Single image. Confirm cost, save to ./venice-image-<id>.png.
venice image "a fierce red dragon, trading-card art"
# Meaningful filename + auto-confirm + hard cap.
venice image "ancient stone golem" --name stone-golem --yes --max-spend 0.05
# Generate 4 variants to pick the best -> ...-1.png ... -4.png.
venice image "frost wyrm, splash art" --variants 4 --yes
# Sizing + tuning (pixel-based models take --width/--height, max 1280).
venice image "portrait card frame" --width 768 --height 1024 \
--negative-prompt "text, watermark" --seed 42 --cfg-scale 7.5 --yes
# Omit the Venice watermark (best for finished card art).
venice image "frost wyrm, splash art" --hide-watermark --yes
# Don't blur flagged fantasy/battle art.
venice image "epic battle, dramatic" --no-safe-mode --yes
# Dry-run: estimate cost + balance, list planned files, spend nothing.
venice image "how much will this cost?" --dry-run
# Batch a whole card set from a file (one prompt per line;
# optional 'name<TAB>prompt'; blank lines and '#' comments skipped).
venice image --from-file cards.tsv --yes -o ./card-art/
venice image --from-file cards.tsv --variants 2 --dry-run
# Shared look across a whole set: a style prefix prepended to every prompt
# plus one negative prompt applied to the entire batch. Output filenames
# stay based on each card's own prompt, not the prefix.
venice image --from-file cards.tsv -o ./card-art/ --yes \
--style-prefix "dark fantasy oil painting, dramatic cinematic lighting" \
--negative-prompt "text, watermark, signature, blurry, lowres"
Shared style templating
For a consistent set (e.g. a whole card deck), keep the long style + negative strings in one place with a preset instead of retyping them:
venice image --from-file cards.tsv -o ./card-art/ --yes \
--preset dark-fantasy --preset-file ./presets.json
--preset-file defaults to ~/.config/venice/image_presets.json; point it at a
file in your project to version presets alongside the assets. Format:
{
"dark-fantasy": {
"style_prefix": "dark fantasy oil painting, dramatic cinematic lighting",
"negative_prompt": "text, watermark, signature, blurry, lowres"
}
}
Precedence: an explicit --style-prefix / --negative-prompt on the command
line overrides the preset; the preset fills whatever you leave off. A single
--negative-prompt (or the preset's) applies to every image in a --from-file
batch.
cards.tsv example (tab between name and prompt):
fire-dragon A fierce red dragon breathing flame, trading-card art
stone-golem An ancient moss-covered stone golem, trading-card art
An unnamed prompt gets a slug from its first few words
Choose a model with --model (default venice-sd35); see
venice models --type image --detail for ids and per-image pricing.
Formats: png (default), webp, jpeg. Aspect-ratio/resolution-tier
models take --aspect-ratio/--resolution instead of --width/--height.
--hide-watermark drops the Venice watermark (Venice may keep it for some
content); --no-safe-mode stops adult-classified art from being blurred. To
drop the watermark by default, set defaults.image.hide_watermark in config
(venice config set defaults.image.hide_watermark true); --no-hide-watermark
forces it back on for a single call.
Upscale images
venice image caps output at 1280px, so take art larger by upscaling it
(1-4x, default 2x) via /image/upscale:
# 2x upscale -> ./env-upscaled.png (960x540 -> 1920x1080).
venice upscale env.png --scale 2 --yes
# Enhance-only pass (scale 1 requires --enhance) with a style hint.
venice upscale portrait.png --scale 1 --enhance --enhance-prompt gold --yes
# ...or set `defaults.upscale.enhance true` once and --scale 1 works bare.
# Custom output, tune how much the enhancer may change the image.
venice upscale card.png --scale 4 --enhance --enhance-creativity 0.3 \
-o card-4k.png --yes
# Dry-run: show the planned output + balance, spend nothing.
venice upscale env.png --dry-run
Input is a PNG/JPEG file under 25 MB. Pricing is dynamic (Venice bills
$0.001-$10.00 per call by input size and scale), so there's no reliable
pre-charge estimate; the balance is shown and you confirm (or --yes).
Remove backgrounds
Venice's generate call ignores background: transparent, so make an asset
opaque then strip its background via /image/background-remove for a
transparent PNG (e.g. rank insignia, icons):
# Local file -> ./insignia-nobg.png (transparent).
venice bg-remove insignia.png --yes
# From a URL instead of a local file.
venice bg-remove --image-url https://example.com/logo.png -o logo-nobg.png --yes
# Dry-run: show the planned output + balance, spend nothing.
venice bg-remove insignia.png --dry-run
Provide exactly one source: a positional file (base64-encoded under 25 MB) or
--image-url. Pricing is dynamic like upscale; balance is shown and you
confirm before the charge.
Edit images
Iterate on already-generated art without regenerating it — recolor, restyle, or
inpaint a card — via /image/edit. Add one or two --layer images (masks or
overlays) to composite instead, which routes to /image/multi-edit (up to 3
images total, base first):
# Prompt-only edit -> ./card-edit.png.
venice image-edit card.png -p "change the sky to a sunrise" --yes
# From a URL, request a 16:9 JPEG at 2K.
venice image-edit --image-url https://example.com/card.png \
-p "make it snow" --aspect-ratio 16:9 --output-format jpeg \
--resolution 2K -o card-winter.jpg --yes
# Mask/overlay composite -> /image/multi-edit (base first, then layers).
venice image-edit base.png -p "apply this mask" --layer mask.png --yes
# Dry-run: show the planned output + balance, spend nothing.
venice image-edit card.png -p "brighter" --dry-run
Provide exactly one base source: a positional file (base64-encoded under 25 MB)
or --image-url. --prompt/-p is required. Optional --model, --aspect-ratio,
--resolution, --output-format, and --safe-mode/--no-safe-mode (config-backable
via defaults.image_edit.safe_mode) map straight to the API;
omit them to take the model defaults (firered-image-edit, PNG for 1K). Pricing
is dynamic like upscale; balance is shown and you confirm before the charge.
Master audio
Venice's audio queue returns a model-default container (sfx = mp3) and its
/audio/speech exposes no sample-rate/bit-depth control, so loudness
normalization, true-peak limiting, and seamless looping are done locally.
venice master shells out to ffmpeg (and ffprobe for --loop) — no
API call, no spend — to produce a WAV master (default 48kHz/24-bit) with
2-pass loudnorm:
# 48k/24-bit WAV master, LUFS -16 / true-peak -1 dBTP -> ./track.mastered.wav
venice master track.mp3
# Seamless-loop ambience (crossfade the tail into the head).
venice master ambience.mp3 --loop --loop-crossfade 3 -o ambience-loop.wav
# Tune targets / format.
venice master pad.wav --lufs -14 --true-peak -1.5 --bit-depth 16 --sample-rate 44100
# Show the ffmpeg commands without running them (works without ffmpeg installed).
venice master track.mp3 --dry-run
The same flags are available on venice music / venice sfx via --master,
which masters the generated file right after it's saved (writing a sibling
*.mastered.wav). Requires ffmpeg; if it's missing the command errors before
spending rather than after generating:
venice music "tense dungeon drone" --duration 60 --yes --master --loop
venice sfx "campfire crackle" --duration 8 --yes --master
Needs ffmpeg on PATH (sudo apt install ffmpeg). ffprobe (bundled with ffmpeg)
is required only for --loop. Both toggles are config-backable
(defaults.sfx.master, defaults.music.master, defaults.{master,sfx,music}.loop)
and both have --no-master/--no-loop to override a config default for one run.
Video
Text-to-video (and image-to-video, see below) on the same async queue as
sfx/music (/video/quote → /video/queue → /video/retrieve →
/video/complete), writing an mp4 to venice-video-<id>.mp4. Generation runs
minutes, not seconds, so it polls less often and waits longer by default
(--poll-interval 5s, --max-wait 900s). --model defaults to the catalog's
default-trait video model; available durations, resolutions, aspect ratios,
and the media-input modes below all vary by model.
# Quote only -- no spend.
venice video "a koi pond at dawn, slow push-in" --dry-run
# Generate a 5s clip (default), confirm the spend, save to ./out.mp4.
venice video "a koi pond at dawn" --duration 5s --resolution 720p -o out.mp4 --yes
# Pick a model / aspect ratio; drop the generated audio track.
# (`defaults.video.no_audio` makes that the default; `--with-audio` re-enables
# it for one run. Note `--audio` is different -- it supplies an INPUT audio file.)
venice video "neon city flythrough" --model seedance-2-0-text-to-video \
--aspect-ratio 16:9 --no-audio --yes
# Queue now, fetch later.
ID=$(venice video "storm clouds timelapse" --duration 10s --background)
venice video-status "$ID"
Media inputs (image-to-video & references)
For models that support them, the generation can be conditioned on media.
Every media flag accepts a local file path or an http(s)/data: URL —
local files are read, size-checked, and encoded to a data: URL for you.
# Image-to-video: animate a still (optionally with an end frame).
venice video "slow zoom out from the figure" --image hero.png --yes
venice video "morph A into B" --image a.png --end-image b.png --yes
# Reference images for character/style consistency (repeatable, up to 9).
venice video "the same knight, new scene" \
--reference-image knight1.png --reference-image knight2.png --yes
# Video-to-video / upscale, and reference videos (repeatable, up to 3). The
# aggregate reference-video duration feeds the *quote* so R2V pricing is right.
venice video "restyle this clip" --video source.mp4 \
--reference-video ref.mp4 --reference-video-duration 5 --yes
# Advanced @Element composition (Kling O3): pass each element as a JSON object.
# Local paths inside the JSON are encoded just like the flags above.
venice video "@Element1 greets @Element2 at @Image1" \
--element '{"frontal_image_url":"alice.png"}' \
--element '{"frontal_image_url":"bob.png"}' \
--scene-image plaza.png --yes
Full media flags: --image, --end-image, --video, --audio (background
music, distinct from --no-audio), --reference-image (≤9),
--reference-video (≤3), --reference-audio (≤3), --scene-image (≤4),
--reference-video-duration, and --element (JSON, ≤4). Image/reference
inputs condition generation and are sent only on /video/queue; --video and
--reference-video-duration also reach /video/quote because they change the
price. Per-model support varies — the API rejects an unsupported combination.
Some (VPS-backed) models return a presigned download_url at queue time and
stream nothing back from /video/retrieve; the CLI fetches the mp4 from that
URL transparently. When queued with --background, the URL is printed alongside
the queue id — pass it back via venice video-status <id> --download-url <url>.
Chat
POST /chat/completions via the OpenAI SDK (see
Optional dependency) — one-shot, or an interactive
multi-turn REPL (see Interactive mode). Streams the reply by
default; --model is validated against /models?type=text (a free GET) before
the paid call, and defaults to the catalog's default-trait text model.
# Simplest: message as an argument, streamed to stdout.
venice chat "Explain DIEM staking in one sentence."
# System prompt, explicit model, no streaming.
venice chat "Rewrite this as a haiku." --system "You are a poet." \
--model llama-3.3-70b --no-stream
# Read the message from stdin (either form).
echo "Summarize this changelog." | venice chat -
git log --oneline -20 | venice chat - --system "Group these into release notes."
# Raw response object for scripting (forces --no-stream).
venice chat "ping" --json | jq '.choices[0].message.content'
Interactive mode
With -i/--interactive — or simply no message on a terminal — venice chat
drops into a REPL that holds the conversation in memory across turns. All the
Venice extensions and --tools (each turn becomes an
agent turn) carry over. Each REPL is a session that
auto-saves after every turn (see Sessions, below), so you can pick
it back up later — settings and all.
# Start a conversation (or just run `venice chat` on a TTY).
venice chat -i --system "You are a terse assistant."
# Resume the most recent chat session, or a specific one by id.
venice chat --continue
venice chat --resume 20260722T220353-9ab8e7
# Resume an old hand-saved transcript file (still works).
venice chat --resume session.json
# Don't persist this one.
venice chat -i --ephemeral
In-REPL slash-commands: /system [text] (show/set the system prompt),
/persona [name] (load a saved system prompt from
~/.config/venice/personas/; with no name, list the available ones),
/model [name] (switch model; with no name, show the current one and list the
catalog), /models (list the available models, marking the current and the
default), /auto and /manual (toggle auto-accepting paid/side-effecting tool
calls for following turns), /compact [N] (summarize older history into one
message, keeping the last N turns verbatim),
/cost (this session's estimated spend so far; --session-max-spend adds a
cap), /usage (a token + cost breakdown for the session, keeping the
cache-read/cache-write/uncached input split distinct so cache-heavy sessions
cost out correctly), /reset (clear history, keep the system prompt),
/save [file] (write the transcript JSON; defaults to the --resume file),
/paste (compose a multi-line message a line at a time, ending with /end —
/cancel aborts), /edit [text] (compose your next message in $EDITOR, like
git commit, optionally pre-seeded with text),
/help, and /exit (or /quit, or Ctrl-D). Ctrl-C aborts the current turn
without ending the session (while a tool-loop turn is running it first pauses to
steer — see Mid-run steering). Tab completes slash-commands (and model ids after
/model , persona names after /persona ). At a per-tool confirmation prompt, a accepts that call and
auto-accepts the rest of the run. --max-tool-calls 0 runs until the model
stops on its own (instead of capping at the default and asking to continue).
Long sessions can also compact automatically. --auto-compact (or
defaults.chat.auto_compact / defaults.code.auto_compact) summarizes the
older prefix into one synthetic message once the prompt crosses
--compact-threshold tokens (default 100 000), keeping the system prompt and
the last --compact-keep-turns turns (default 10) verbatim. The trigger uses
the server-reported usage.prompt_tokens when a turn provides it (falling
back to a chars-per-token estimate otherwise), so it fires on the real prompt
size rather than a guess. Compaction is best-effort — a failed summarization
call leaves the history untouched — and never orphans a tool result from its
assistant turn. It's off by default because it costs a summarization call.
Sessions
Every venice chat / venice code REPL is a session that auto-saves after
each turn to ~/.config/venice/sessions/<id>.json (mode 0600;
$VENICE_SESSIONS_DIR overrides the location). Unlike a bare /save transcript,
a session travels with its settings — model, system prompt, generation
parameters, max-tool-calls, the venice code sandbox root, and the running
token/cost usage — so resuming restores the whole context, not just the messages.
The API key is never written to a session.
One-shot venice code "task" runs are sessions too (they persist unless
--ephemeral), so an unattended --auto run is resumable, inspectable, and —
new in this release — steerable while it runs (see below).
venice sessions ls # list saved sessions (newest first; flags pending steers)
venice sessions ls --json # same, as JSON (adds each session's pending-steer count)
venice sessions show <id> # settings + message summary for one session
venice sessions send <id> "…" # queue a mid-run steering message (see below)
venice sessions rm <id> # delete a session (and its steering mailbox)
venice chat --continue # resume the most recent chat session
venice chat --resume <id> # resume a specific session by id (restores settings)
venice code --continue # (code's most recent session re-sandboxes to its root)
Resume precedence is explicit flag > saved session > config default: passing
e.g. --model on resume overrides the session's saved model, but omitting it
keeps what the session used. --resume still accepts a plain transcript file
for back-compat (it's imported into a fresh session, leaving the file untouched).
Pass --ephemeral (alias --no-save) to run without persisting a session (which
also makes the run unsteerable). /save [file] remains an explicit, separate
transcript export.
Mid-run steering (venice sessions send)
A running agent — especially venice code --auto — used to have only two
controls: let it finish, or kill it (losing uncommitted work and metered spend).
venice sessions send adds a third: steer it without stopping it.
venice sessions send latest "actually, prioritize the #3 regression first"
venice sessions send 20260724T101530-ab12cd "skip the CSS, focus on the API"
echo "long note…" | venice sessions send latest - # read the message from stdin
The message is dropped into the session's file mailbox
(~/.config/venice/sessions/<id>/mailbox/, one atomic 0600 file per message).
At its next checkpoint — the boundary between tool calls, before the next model
turn — the agent drains the mailbox and consumes each message as a tagged user
turn, exactly as if you had typed it interactively. It's additive input, not a
reset: your --max-tool-calls / --session-max-spend budgets are unchanged.
latest targets the most recent code session (sessions ls --json shows ids +
pending counts for scripting). Targeting is by recency, not liveness — there's
no process tracking — so a message sent to a session that has already finished
simply waits in its mailbox and is drained the next time you --resume it. The
mailbox is a local, owner-only directory (0700), not a network channel: a steer
carries the same trust as the original task. This is stdlib-only — no daemon, no
sockets, no extra dependency.
When you're watching an attached run in a terminal (venice code --interactive,
or a foreground --auto/--manual run), you don't need a second shell — just press
Ctrl+C:
- First Ctrl+C lets the current step finish, then pauses at the next checkpoint and
asks:
[paused] message to the agent (empty = resume, Ctrl+C = abort):. Type a line and it's fed in as a steer exactly likesessions send; press Enter (or Ctrl+D) on an empty line to resume unchanged. - Second Ctrl+C — either at that prompt, or before the checkpoint is reached — aborts, as it always has (the one-shot run exits 130 with its partial transcript saved; the REPL rolls the current turn back and keeps the session).
This is tty-only (a piped or --json run keeps the plain, non-interactive behavior)
and needs no flag or mailbox — it works even for a --ephemeral attached run. Under the
hood it's the same checkpoint as sessions send, so an in-flight tool call is never cut
off mid-write.
Personas (local system-prompt files)
Keep a library of your own reusable system prompts as plain .md/.txt files
under ~/.config/venice/personas/ — private, version-controllable, offline, and
complementary to Venice's server-side --character slugs. Drop a file in the
directory, then load it by bare name:
mkdir -p ~/.config/venice/personas
printf 'You are a terse pirate. Answer in one sentence.\n' \
> ~/.config/venice/personas/pirate.md
venice chat --persona pirate # load at launch
# ...or mid-session in the REPL:
# /persona -> list personas (name + first line)
# /persona pirate -> load personas/pirate.md as the system prompt
/persona <name> replaces the system prompt but keeps the conversation (use
/reset for a clean slate), exactly like /system. At launch, --persona
(or the defaults.chat.persona config default) seeds the system prompt; an
explicit --system / defaults.chat.system takes precedence. Names are
bare only — a name with a path separator or .. is refused, and the lister
enumerates only the personas/ directory, so a persona can never read the
neighbouring credentials file. venice chat only for now.
Venice extensions
Venice augments the OpenAI schema with a venice_parameters block; these flags
map onto it:
# Live web search + inline source citations (printed to stderr).
venice chat "What shipped in the latest Venice API update?" \
--web-search on --web-citations
# Scrape URLs in the message via Firecrawl.
venice chat "Summarize https://venice.ai/blog" --web-scraping
# Talk to a public Venice character by its Public ID slug.
venice chat "Introduce yourself." --character venice
# Reasoning models: drop <think> blocks, or disable thinking entirely.
venice chat "Tricky logic puzzle..." --strip-thinking
venice chat "Just answer fast." --no-thinking
# Omit Venice's supplied system prompt (uncensored/raw behavior).
venice chat "..." --no-venice-system-prompt
# xAI native web+X search on grok models (extra ~$0.01/search).
venice chat "Latest posts about Venice?" --model grok-4-20 --x-search
| flag | effect |
|---|---|
--web-search {auto,on,off} |
Venice web search (default off) |
--web-citations |
cite web sources (with --web-search) |
--web-scraping |
Firecrawl-scrape URLs in the message |
--character SLUG |
use a public Venice character |
--no-venice-system-prompt |
omit Venice's supplied system prompt |
--strip-thinking |
strip <think> blocks (reasoning models) |
--no-thinking |
disable thinking (reasoning models) |
--x-search |
xAI web+X search (grok; extra ~$0.01/search) |
Chat pricing is dynamic (per token, model-dependent), so there's no pre-call
quote; pass --json or watch the usage: line on stderr to see token counts.
Agent / tool calling
With --tools (alias --agent), venice chat becomes a self-contained agent:
the model can call venice's own endpoints as in-process function tools and the
completion runs in a loop (model → tool call → tool result → repeat) until it
produces a final answer. These run in-process on the [openai] extra alone
(no mcp SDK, no subprocess):
venice_image, venice_tts, venice_sfx, venice_music, venice_upscale,
venice_bg_remove, and venice_chat (a sub-completion / subagent primitive) —
seven of the capabilities venice mcp-serve exposes (which adds venice_video
and venice_image_edit) — plus project_search,
a read-only semantic search over the project's local
venice index for locating code by meaning before acting on it (a snapshot of
the last index build — pair it with reindex, a paid tool that rebuilds the index
so recall reflects edits made this session), and
venice_models, a read-only lookup that lists model ids for a given catalog
type (its single type arg — text/code/image/video/music/tts/embedding/upscale,
or all) so the model can pick a valid model for the other tools instead of
guessing, and venice_model_details (single model arg) which returns one
model's pricing (cost), capabilities (text models — supportsVision etc.),
constraints (image/media — aspect ratios, resolutions, qualities, prompt-length
limit), and the full model_spec, so the agent can budget input and confirm a
model fits, and venice_vision, which sends a local image (input_path, as a
base64 data-URL) or an image_url to a vision-capable text model and returns
what it sees — the agent's eyes, so it can inspect its own generations
(watermarks, character consistency, glitches) instead of working blind; an
optional prompt directs the question, and when model is omitted a
supportsVision model is auto-picked from the catalog. Not spend-gated.
(venice code gets all three too.)
# One command, multiple steps: the model generates an image, then critiques it.
venice chat --tools "Generate a fire-elemental trading card, then critique it."
# Text-only agentic reasoning via the venice_chat subagent tool (no paid media).
venice chat --tools "Use venice_chat to draft a haiku, then improve it."
# Restrict the toolset and cap the number of tool calls.
venice chat --tools --tool venice_image --max-tool-calls 3 "Draw three logo ideas."
Details and safety:
- Capability guard. Tools are offered only if the chosen model advertises
supportsFunctionCalling; on a non-tool model the command prints a note and degrades to a plain one-shot chat. Without--tools,venice chatis unchanged. - Spend gating (paid tools) reuses the MCP rails: each paid call auto-approves
under a per-call cap (
--max-spend, default$0.10/$VENICE_MCP_MAX_SPEND). An over-cap call prompts[y/N]on a TTY; non-interactively (or if you decline) the block is handed back to the model, which adapts.--yesauto-approves every paid call (this bypasses the per-call cap —--max-tool-callsstill bounds the count). The model itself can never raise its spending authority. - Session spend cap (
--session-max-spend, #66) meters the chat completions themselves — not just paid tools. Each turn's server-reportedusageis priced against the session model's per-1M-token catalog rate; once the running total reaches the cap the loop stops starting new turns and forces a final answer (chat has no pre-call quote, so it bounds further spend, not a turn already in flight). Config-backable viadefaults.chat.session_max_spend/defaults.code.session_max_spend. Distinct from--max-spend(the per-call tool cap). A model with unknown pricing is counted (tokens) but not charged. --output DIRsets where generated files are written (default: cwd).- Non-streamed in v1. The tool path buffers each turn, so
--streamis ignored when--toolsis on;--jsonprints the final completion object.
| flag | effect |
|---|---|
--tools / --agent |
enable the in-process tool-calling loop |
--tool NAME |
restrict to this tool (repeatable; default: all of them) |
--max-tool-calls N |
cap tool invocations before forcing an answer (default 8) |
--max-spend USD |
per-call auto-approve cap for paid tools |
--session-max-spend USD |
cap total chat-completion spend for the session |
--yes / -y |
auto-approve every paid tool call and side-effecting MCP tool |
--output DIR / -o |
directory for generated files |
--shell / --exec |
add a gated shell tool (/bin/sh -c in the cwd); implies --tools |
--shell-allow CMD |
allow only these commands for --shell (repeatable; adds to config shell.allow) |
--shell-deny PATTERN |
refuse commands matching these globs (repeatable; adds to config shell.deny) |
--shell-unrestricted |
acknowledge an empty allowlist under --yes (required for that combination) |
--browser |
add web_fetch + browser_capture tools (fetch a URL / headless-render a page); implies --tools |
--browser-allow HOST |
allow only these hosts for the browser tools (repeatable; adds to config browser.allow) |
--browser-deny PATTERN |
refuse URLs whose host/URL matches these globs (repeatable; adds to config browser.deny) |
--memory |
add persistent memory + task tools (durable notes + a checklist); implies --tools (see Memory & tasks) |
--mcp NAME |
attach a registered external MCP server's tools (repeatable) |
--no-mcp |
attach no MCP servers (overrides a configured default) |
Shell exec tool (--shell)
--shell (alias --exec) adds a gated shell tool so the agent can run any CLI
(gh, git, curl, build/test commands) — the same /bin/sh -c rail
venice code uses, with cwd set to the current directory,
a timeout, size-capped output, and the Venice API keys scrubbed from the child
environment. Every command is shown and confirmed before it runs ([y/N]), unless
--yes auto-approves. --shell implies --tools.
venice chat --shell --shell-allow gh --shell-allow git "Open my oldest assigned issue."
Scope it with an allow/deny policy — CLI flags add to a shared top-level shell
section in config (shell.allow / shell.deny), the single source of truth
for both venice chat --shell and venice code's run tool:
venice config set shell.deny '["rm *", "sudo *", "* --force*"]'
venice config set shell.allow '["git", "gh", "ls", "cat"]'
- Deny globs are matched on the whole command line and on each token, are always
enforced, and win over allow. Use
sudoto block by name,*rm -rf*for a substring. - A non-empty allowlist additionally requires a single simple command — no shell
operators, pipes, redirects, substitutions, or variables (
; | & < > ( ) \$) — and the leading command's basename must be allowlisted (globs likegit*are fine). This stops an allowlistedgh && rm -rf ~` from slipping through. - An empty allowlist is unrestricted (only the confirm gate + deny apply). Combining
that with
--yes(auto-approved arbitrary shell) is refused unless you pass--shell-unrestrictedto acknowledge it. - Not exposed over
venice mcp-serve— a shared/remote server running arbitrary shell is a much larger blast radius and is deliberately out of scope.
Web & browser tools (--browser)
--browser adds two tools so the agent can look at the web — useful for verifying a page
it (or you) just built. Implies --tools. Both are read-only and not spend-gated;
the URL policy below is the guard.
web_fetch— a stdliburllibGET that returns the page as text (tags stripped) or raw HTML (mode=html). Zero extra dependencies; good for non-JS "read this page."browser_capture— headless-renders the page with a Chromium-family browser (chromium/chrome/brave) and returns the post-JS DOM (mode=dom/text) and/or a screenshot PNG (mode=screenshot/both, written to--outputand returned as a path, never inline). Passassert_contains="…"to deterministically check the rendered DOM contains a string — the robust "did the JS actually appear" check (beats eyeballing a screenshot). If only Firefox is installed it degrades to screenshot-only (its headless CLI can't dump the DOM); with no browser at all it reports "no headless browser available," like thegitrail.
venice chat --browser "Fetch https://example.com and summarize it."
venice chat --browser --browser-allow 'localhost' \
"Open http://localhost:8123/ and confirm the game cards rendered (assert_contains)."
URL safety. The scheme must be http/https and the cloud-metadata endpoint
(169.254.169.254) is always blocked — these hard stops are not configurable. file://
and everything non-http is refused. On top of that, an operator allow/deny policy (CLI
flags add to a shared top-level browser section in config) scopes which hosts
are reachable:
venice config set browser.deny '["*.internal", "10.*"]'
venice config set browser.allow '["localhost", "*.example.com"]'
- Deny globs match the URL host and the full URL, are always enforced, and win over
allow. A non-empty allow list restricts to matching hosts; empty = any host (still
subject to the hard stops + deny).
localhostis reachable by default (the common case is verifying a locally-served page). - The headless browser runs untrusted page JS under an allowlisted environment (the Venice API key and other ambient tokens are dropped) and a throwaway profile dir.
- DOM/text output and downloads are size-capped; tune with
defaults.browser.*(wait_ms,timeout,max_bytes). - Not exposed over
venice mcp-servein this release (chat/code only), likeproject_search/reindex.
Web search (--web-search)
--browser can only fetch a URL the agent already knows. --web-search closes the
other half of the loop — discovery. It adds one tool, venice_web_search, so the coding
agent can look something up on the web when a fix needs documentation it doesn't have a link
for (an API's docs, a library's usage, an error message).
venice_web_search(query) makes one Venice web-search completion (server-side
enable_web_search + enable_web_citations, the same feature behind venice chat --web-search) and returns a short answer plus the cited URLs. To then read a cited
page in full, follow up with web_fetch — so pair --web-search with --browser, and every
fetched URL stays under the browser.* allow/deny policy above. Search discovers; the
browser policy still governs what gets read.
# discover + read: search finds the doc, the browser fetches it (under browser.* policy)
venice code --web-search --browser --auto \
"The stripe SDK call is failing with an idempotency error — look up the fix and apply it."
- Model. Web search needs a model advertising
supportsWebSearch. By default the coding--modelis used when it qualifies, else the first capable model in the catalog; override with--web-search-model MODEL(ordefaults.code.web_search_model). No model id is hard-coded — it's resolved against the live/modelscatalog. - Billed, bounded. It rides the normal completion path (same key, same billing) rather
than a scraper, so there's no new dependency or secret. It isn't per-call spend-gated; its
cost is bounded by the agent's tool-call budget, and each result carries a best-effort
cost_estimate_usd. - Who gets it. The coding agent (and a
--planner) can use it directly; with--scoutthe read-only scout becomes a "docs scout" (read the tree and search the web). Spawn workers never get it — a worker acting on instructions injected via a search result is the blast-radius case, so web tools are denied to workers by default. - Config:
defaults.code.web_search(bool) /defaults.code.web_search_model. Not exposed overvenice mcp-serve.
Memory & tasks (--memory)
--memory gives the agent a durable place to keep notes and a checklist, so multi-step
and cross-session work survives beyond one transcript. Implies --tools. Seven free,
local, offline tools (no API calls, no spend gate):
memory_write/memory_read/memory_search/memory_list— named notes the agent recalls later. Two tiers:scope="project"(default) rides the repo at<root>/.venice/memory/so it's shared by anything working in that tree;scope="global"lives user-global (travels with the agent across projects).memory_searchis a plain substring match over names/descriptions/bodies (zero-dep, always works);memory_listreturns metadata only (names/types/descriptions/timestamps — a cheap index to decide what to read).task_add/task_update/task_list— a lightweight project-only checklist (pending→in_progress→done) the agent maintains across turns and--resume.
venice code --memory --auto "Refactor the parser; track your steps as tasks."
venice chat --memory "Remember that this project uses tabs, not spaces (scope=project)."
Inspect or prune what the agent stored with the venice memory command:
venice memory ls # both tiers (metadata only)
venice memory ls --scope global # just the global tier
venice memory show <name> # one note, including its body
venice memory rm <name> # delete a note (default: project tier)
venice memory tasks # the project checklist (--status filters)
- Locations: project notes/tasks at
<root>/.venice/memory/(git-ignored by default, like the semantic index); global notes at~/.config/venice/memory/($VENICE_MEMORY_DIRoverrides). All store files are0600. - Hygiene (CLAUDE.md): a note name is refused if it's secret-shaped
(
credentials,id_rsa*,*.key,.env,*secrets*, …), so the store can't be used to label or stash a credential. - Not exposed over
venice mcp-serve(chat/code only), like the browser rails.
External MCP tools (--mcp)
--mcp NAME attaches the tools of an external MCP
server (filesystem, git, shell, ...) alongside the built-in venice tools, so one
agent can drive both. Register servers first with
venice config add (stdio or http/sse), then name them:
venice config add fs --command npx --arg -y --arg @modelcontextprotocol/server-filesystem --arg /work
venice chat --mcp fs "Summarize the TODOs across the source files."
- Needs the
[mcp]extra (pip install "venice-cli[mcp]", Python ≥ 3.10).--mcpimplies the agent loop (no separate--toolsneeded); it still requires a function-calling model and degrades to plain chat otherwise. - Remote tools are advertised as
server__tool(namespaced to avoid collisions). - Side-effecting tools are gated. A remote tool that isn't annotated read-only
prompts for confirmation on a TTY (or feeds the request back to the model
non-interactively) before it runs; read-only tools run freely.
--yesbypasses the gate. This rides the same confirm rail as paid built-in tools. - Multiple
--mcpflags attach multiple servers;--no-mcpoverrides adefaults.chat.mcpconfig default. Attach timeouts:$VENICE_MCP_CONNECT_TIMEOUT,$VENICE_MCP_CALL_TIMEOUT. - Works the same in interactive mode — servers stay attached for the whole session and are torn down on exit.
Embeddings
Turn text into embedding vectors with a Venice embedding model (/embeddings,
via the OpenAI SDK). The model defaults to the catalog's default-trait
embedding model; pass --model to pick another (see
venice models --type embedding).
# Single input -> one JSON array on stdout.
venice embed "the quick brown fox"
# Read the input from stdin.
echo "summarize me" | venice embed -
# Batch: one input per non-empty line -> one vector per line (index order).
venice embed --from-file corpus.txt
# Pipe vectors to jq (newline-delimited JSON, one array per line).
venice embed "hello" | jq 'length'
# Truncate dimensions (if the model supports it) and pick a model.
venice embed "hello" --model text-embedding-qwen3-8b --dimensions 256
# Full raw response object (model, data, usage) instead of bare vectors.
venice embed "hello" --json | jq '.usage'
By default each embedding prints as a JSON array, one per line;
--encoding-format base64 requests base64-packed vectors instead of floats.
Local / alternate backend
Because venice embed rides on the OpenAI SDK, it can point at any
OpenAI-compatible embeddings endpoint — including a local one (llama.cpp,
Ollama, text-embeddings-inference). Pass --embed-base-url (with
--embed-model, since the alternate server has its own catalog) to swap
backends; this skips the Venice catalog and needs no Venice key. Venice stays
the default when the flag is absent.
# Embed against a local server -- no Venice key required.
venice embed --embed-base-url http://localhost:1234/v1 \
--embed-model my-local-model "the quick brown fox"
The URL can also come from $VENICE_EMBED_BASE_URL, and a key (if the backend
needs one) from $VENICE_EMBED_API_KEY or the named-secret store (see
venice secret) — never config.json. Persist it once
with venice login --embed (or venice secret set embed) to retire the
export VENICE_EMBED_API_KEY=… in your shell profile; the env var still wins when
set. The base-url/CA-bundle flags are config-backable per-flag via defaults.embed.*
(see venice config).
Self-signed backends. A local embedder fronted by Traefik/Caddy often serves
a private or self-signed TLS cert, which the OpenAI SDK rejects
(embed: connection error, exit 8). Two opt-in escape hatches — applied only
to --embed-base-url, never the Venice endpoint:
# Trust a private CA (verification stays ON -- preferred):
venice embed --embed-base-url https://embed.local/v1 --embed-model bge-m3 \
--embed-ca-bundle /etc/ssl/my-ca.pem "hi" # or $VENICE_EMBED_CA_BUNDLE
# Disable verification entirely (self-signed, no CA handy -- prints a warning):
venice embed --embed-base-url https://embed.local/v1 --embed-model bge-m3 \
--embed-insecure "hi"
--embed-ca-bundle is config-backable (defaults.embed.embed_ca_bundle) and
reads $VENICE_EMBED_CA_BUNDLE. --embed-insecure is CLI-only by design —
turning verification off should always be an explicit, visible choice, never
something a stale env var or config file switches on. The two are mutually
exclusive, and passing either without --embed-base-url is an error (exit 2).
Semantic search
venice index builds a local semantic index of a project tree, and venice search finds the chunks most relevant to a natural-language query by meaning
rather than by keyword. Both use the same embedding machinery as venice embed
(the [openai] extra, Venice or a local backend); the vector store and the
cosine search are pure-stdlib, so no extra dependency is needed.
# Index the current tree. Venice has no default embedding model, so pass one
# (or set defaults.index.model). Vectors land in ./.venice/index/.
venice index . --model text-embedding-bge-m3
# Search it from anywhere in the tree (walks up to find .venice/index).
venice search "where is the retry/backoff logic"
# Top-3 results as JSON (path, line range, score, preview).
venice search "jwt refresh handling" -k 3 --json
Text output is one hit per line as SCORE path:start-end, followed by a short
preview of the matched lines:
0.8137 src/venice/client.py:88-120
def post_for_bytes_or_json(self, path, body, ...):
Incremental. Re-running venice index re-embeds only the files whose
contents changed (keyed on a SHA-256 of each file); unchanged files keep their
vectors and deleted files are dropped. --rebuild forces a full re-index — also
required if you switch model/dimensions/backend, since vectors from different
embedding spaces are not comparable.
What gets indexed. UTF-8 text files under the tree, chunked into overlapping
line windows (--chunk-lines / --chunk-overlap). The walker skips binaries,
oversized files, .git/node_modules/virtualenvs and similar, and honors a
simple top-level .gitignore plus any --exclude GLOB. Credential- and
secret-shaped files (.env, credentials, *.pem, *.key, id_rsa*) are
never indexed, and symlinks pointing outside the tree are ignored.
Local backend. As with venice embed, --embed-base-url (+ --embed-model,
or $VENICE_EMBED_BASE_URL / $VENICE_EMBED_API_KEY) points indexing at a local
OpenAI-compatible server — cheap for embedding a whole tree, and needs no Venice
key:
venice index . --embed-base-url http://localhost:1234/v1 --embed-model bge-m3
venice search "parse the queue response" # uses the index's own backend/model
Self-signed backends. Both commands accept the same TLS escape hatches as
venice embed — applied only to a local backend, never the Venice endpoint:
venice index . --embed-base-url https://embed.local/v1 --embed-model bge-m3 \
--embed-ca-bundle /etc/ssl/my-ca.pem # trust a private CA (or $VENICE_EMBED_CA_BUNDLE)
venice search "parse the queue response" --embed-ca-bundle /etc/ssl/my-ca.pem
venice index . --embed-base-url https://embed.local/v1 --embed-model bge-m3 \
--embed-insecure # disable verification (warns; CLI-only)
--embed-ca-bundle reads $VENICE_EMBED_CA_BUNDLE and is config-backable
(defaults.index.embed_ca_bundle / defaults.search.embed_ca_bundle);
--embed-insecure is CLI-only, mutually exclusive with it, and errors (exit 2) if
the flags don't apply (no --embed-base-url for index, or a Venice-built index
for search). For search the CA bundle is supplied fresh at query time — it is
never baked into the index — and the project_search agent tool also honours
$VENICE_EMBED_CA_BUNDLE, so a venice chat/venice code session can search an
index built against a self-signed embedder.
The index is machine-generated: venice index drops a self-ignoring
.venice/.gitignore, so it won't be committed even if your repo doesn't already
ignore .venice/. Config-backable per-flag via defaults.index.* /
defaults.search.*. venice search is also exposed to the chat agent as the
project_search tool (see Agent / tool calling), so a venice chat --tools
session can locate code by meaning before acting on it.
project_searchis a snapshot;grepis live. The.veniceindex is a point-in-time build, soproject_searchreturns pre-edit content for files the agent changed this session, while thegreptool always walks the working tree. After edits, the agent can call thereindextool to rebuild the index (it re-embeds only the files whose contents changed, reusing the index's existing embedding backend).reindexis paid (it calls the embeddings API) and routes through the same confirm gate as the other paid tools — approve it at they/a/Nprompt (or run non-interactively with--yes). If no index exists yet it tells you to runvenice indexfirst.
Coding agent (venice code)
venice code is a self-contained coding agent ("vcoder") built on the tool loop.
Point it at a project and give it a task: it proposes a plan, waits for your
acceptance, then reads, edits, and runs commands using built-in, path-sandboxed
tools, powered by a function-calling Venice model. Needs the [openai] extra and a
tool-calling model (unlike venice chat --tools, it errors out rather than degrading
if the model can't call tools). The coding engine itself is pure stdlib — no new
dependency.
# Human at a terminal: see the plan, then choose auto/step at the prompt.
venice code "add retry with backoff to the HTTP client and a test" -m mistral-31-24b
# Autonomous, unattended (a script/cron): accept + run to completion, JSON out.
venice code --auto --json "bump the version and update CHANGELOG" > result.json
# Two-step (for a script or a controlling LLM to approve out of band):
venice code --plan-only --json "refactor the parser" > plan.json # prints plan, exits
venice code --auto "refactor the parser" # then execute
# An interactive coding session (tools on; changes confirm per step).
venice code -i
Plan → acceptance → run. The command always plans first (one no-tools turn that emits a numbered plan + acceptance criteria), then crosses an acceptance boundary three possible ways, then executes and finally self-checks the criteria:
| How it's launched | How the plan is accepted | Run mode |
|---|---|---|
| Human, terminal | Interactive prompt: [a]uto / [s]tep / [e]dit / [N]o |
chosen at the prompt |
| Flag-driven | --auto (accept + autonomous) or --manual (accept + step) |
from the flag |
| Out of band | --plan-only prints the plan and exits 0; the caller re-invokes to run |
deferred |
Non-interactive with neither --auto nor --plan-only aborts (exit 2) before any
model call — side-effecting work never runs unattended without an explicit opt-in.
After execution a final turn reports each criterion MET/NOT MET and ends with an
ACCEPTANCE: PASS/FAIL verdict; with --json the verdict lands in the envelope
(acceptance.verdict = pass/fail/unknown). The verdict parse is
case/format-tolerant and re-prompts once for the verdict line if the first reply
lacks it, so a correct run whose model phrased its verdict loosely still exits 0. The
exit code reflects it: 0 = all met (or check skipped), 1 = not met, 10 = the model
never emitted a parseable verdict even after the re-prompt (the work may still be
complete — a loud stderr warning is printed).
Tools (path-sandboxed to the project root; mutating tools confirm unless --auto):
| Tool | Does | Confirms? |
|---|---|---|
read_file / list_dir / grep |
read a file, list a dir, regex-search the tree | no |
git |
read-only git (status/diff/log/show/…) |
no |
project_search |
semantic search over the .venice index (if built) — a snapshot of the last build; use grep for live matches |
no |
reindex |
rebuild the .venice index so project_search reflects this session's edits (re-embeds only changed files); present only when an index exists |
yes (paid) |
venice_vision |
describe/inspect a local image or image URL via a vision-capable model | no |
write_file |
create/overwrite a file (atomic) | yes |
edit_file |
replace an exact, unique string in a file | yes |
apply_patch |
apply a batch of edits grouped per file, atomically per file (use occurrence=N for non-unique strings) |
yes |
run |
run a shell command (/bin/sh -c) at the active root |
yes |
attach_root |
register another directory as a project root (for work spanning repos) and, by default, switch the active root into it so relative paths and run/git follow — writes outside the writable roots fail loudly |
no |
venice_image / venice_image_edit / venice_sfx / venice_music / venice_tts / venice_upscale / venice_bg_remove / venice_video |
generate/edit images, audio & video into the project — opt-in with --assets |
yes |
web_fetch / browser_capture |
fetch a URL (text/HTML) or headless-render a page (post-JS DOM + screenshot) to verify its own work — opt-in with --browser (see Web & browser tools) |
no |
memory_write / memory_read / memory_search / memory_list |
durable notes the agent recalls across turns/sessions (two tiers: project + global) — opt-in with --memory (see Memory & tasks) |
no |
task_add / task_update / task_list |
a project-only checklist the agent tracks (pending/in_progress/done) — opt-in with --memory |
no |
venice_scout |
delegate a read-only investigation to a disposable subagent with a fresh context; returns a structured report so exploration doesn't pollute the main context — opt-in with --scout (see Scout subagent) |
no |
venice_web_search |
search the web to discover documentation you don't have a URL for; returns a short answer + cited URLs (billed, but bounded by the tool-call budget) — opt-in with --web-search (see Web search) |
no |
Safety. Every filesystem path is resolved and confined to the writable roots —
the startup root (default: cwd, or --root / $VENICE_CODE_ROOT) plus any added with
--allow-root / config roots.allow / the attach_root tool. A write that lands
outside the writable set fails loudly (naming the roots) rather than silently
redirecting — so a session that spans repos can't leak files into the wrong one; deny
roots (--deny-root / roots.deny) are readable but never writable (deny wins). This
is a guardrail, not a sandbox: a shell command can still write anywhere. A path that
also names a secret-shaped file (credentials, .env, *.pem, *.key, …) or lives
under .git/.venice is refused — the same denylist venice index uses. run executes
with the working directory forced to the active root, a timeout (--exec-timeout),
size-capped output, and the Venice API keys scrubbed from the child environment. Note
that a shell command can still touch paths outside the root (cat ../x); run's
boundary is the confirm gate (the exact command is shown before it runs) plus the
forced cwd, timeout, and env-scrub — which is why it always confirms. git mutations
(add/commit) go through the gated run tool. run also honors the shared
shell allow/deny policy (config shell.* or
--shell-allow/--shell-deny) — a denied command is refused; an empty policy leaves
it unrestricted (unchanged behavior).
| flag | effect |
|---|---|
--auto, -y |
accept the plan and run autonomously (auto-approve every tool call); required to run with no terminal |
--manual |
accept and run with per-step confirmation (default on a terminal) |
--plan-only |
print the plan and exit without executing |
--no-plan |
skip the planning turn and execute directly |
--no-verify |
skip the post-run acceptance-criteria check |
--root DIR |
project directory to sandbox to (default: cwd) |
--allow-root DIR / --deny-root DIR |
extra directories the file tools may read+write / roots excluded from writes (repeatable; adds to config roots.*). The agent can also add roots at runtime with attach_root |
--max-tool-calls N |
cap tool invocations before forcing a final answer (default 25) |
--exec-timeout SECS |
timeout for run/git (default 120) |
--shell-allow CMD / --shell-deny PATTERN |
scope the run tool with the shared allow/deny policy (repeatable; adds to config shell.*) |
--browser / --browser-allow HOST / --browser-deny PATTERN |
expose web_fetch + browser_capture so the agent can verify a rendered page; scope hosts with the browser.* allow/deny policy (see Web & browser tools) |
--assets |
also expose the in-process asset-generation tools (image / image-edit / sfx / music / tts / upscale / bg-remove / video) so the agent can create images, audio & video in the project; paid — each confirms per call unless --auto |
--scout |
expose venice_scout: delegate a read-only investigation to a disposable subagent with a fresh context; keeps exploration out of the main context (see Scout subagent) |
--spawn |
expose venice_spawn: delegate a bounded write/paid task to a disposable worker subagent with a fresh context and a role-scoped subset of your tools; edit churn stays quarantined and it returns a structured report to merge (see Worker subagent) |
--spawn-max-spend USD |
per-worker USD cap on an asset worker's cumulative estimated media spend (default $2.00; <= 0 disables); config defaults.code.spawn_max_spend |
--subagent-max-tokens N |
per-subagent cap on the cumulative prompt+completion tokens a venice_scout or venice_spawn subagent spends across its turns (default off; <= 0 disables); once crossed the subagent is asked to wrap up and its report carries the token count. A cumulative-usage ceiling, not a context-window size limit, and distinct from --max-tokens (per-turn output); config defaults.code.subagent_max_tokens |
--planner |
planner harness: implies --scout --spawn --memory, mandates the decompose → dispatch → track → merge protocol, and adds venice_merge — a consolidated rollup of every dispatch (see Planner harness) |
--parallel |
dispatch independent venice_scout/venice_spawn subagents concurrently (bounded pool) instead of one at a time, so a planner's independent units overlap in wall-clock; opt-in, serial otherwise; config defaults.code.parallel (see Parallel dispatch) |
--web-search / --web-search-model MODEL |
expose venice_web_search so the agent can discover documentation on the web (answer + cited URLs); pairs with --browser to then read a cited page under the browser.* policy (see Web search) |
--auto-compact |
summarize older history once the prompt crosses --compact-threshold tokens (default 100 000), keeping the last --compact-keep-turns turns (default 10); long runs stay in-context |
-i, --json, --model, --system |
interactive REPL · JSON envelope · model · extra system instructions |
--persona NAME |
load ~/.config/venice/personas/NAME.md as the system prompt at launch (/persona in the REPL) |
With --assets, generated files land in $VENICE_MCP_OUTPUT_DIR or, by default, under
the project root, and paid calls are capped per call by $VENICE_MCP_MAX_SPEND (default
$0.10) — except that --auto auto-approves every call and so bypasses that cap;
--auto --assets can incur up to --max-tool-calls paid generations, so use a cheap
model and a sane --max-tool-calls when running unattended.
Per-flag config defaults live under defaults.code.* (e.g. model, root, auto,
assets, scout, spawn, spawn_max_spend, subagent_max_tokens, planner,
max_tool_calls).
Scout subagent (--scout)
--scout adds one tool, venice_scout, so the coding agent can delegate an
investigation to a disposable subagent instead of exploring in its own context. The
scout starts from a fresh context, gets only read-only tools
(read_file/list_dir/grep/read-only git, plus project_search when a .venice
index exists), investigates, and returns a single structured report — findings,
confidence, dead-ends, what it did not check, and which claims it verified live vs.
inferred. Those sections are also parsed into a fields map on the returned report, so
the planner can consume the handoff programmatically rather than re-reading prose. The
caller sees only that report, not the dozens of tool calls behind it, so a
big "where/how does X work?" question doesn't fill the main agent's context with
exploration noise. This is a context firewall, not a role-specialized worker.
A scout cannot edit files or run commands (read-only by construction — it can never
be handed a write/exec tool, nor spawn another scout), and each run is bounded by a
tool-call budget (max_tool_calls in the call, default 6, hard max 15) so it can't run
away. Use it before an edit to scope the change:
venice code --scout --auto "Add rate-limiting to the API client. First scout how the \
existing client is structured and where requests are made, then implement it."
Worker subagent (--spawn)
Where the scout is a read-only context firewall, --spawn adds venice_spawn — the
same firewall for doers. The coding agent can delegate a bounded implementation task
to a disposable worker that runs in a fresh context, does the work, and returns a
single structured report — outcome, changes (files + commands), what it verified live
vs. assumed, follow-ups, and blockers. As with the scout, those sections are also parsed
into a fields map on the returned report, and the caller sees only the report, not the
edit churn behind it — so the planner's context stays clean while the worker does the
churny part.
A worker gets a role-scoped subset of your already-built tools (never more than the session was granted):
role="code"(default, spend-free) →read_file/list_dir/grep/write_file/edit_file/apply_patch/run/git(plusproject_searchwhen a.veniceindex exists) — thefs+exec+vcs+searchcategories.role="asset"→ the media generators and their support tools (image/audio/videopluscatalog/vision/jobs) — needs--assetson the parent, otherwise the grant is empty and the call errors.
Containment is structural, not a per-call prompt: a worker's writes flow through the
same writable roots (fail loud outside them) and the
same run allow/deny policy as the parent; it gets no attach_root (can't widen its
roots) and no scout/spawn tools (nesting is capped at one level — a planner
scouts/spawns, a worker does neither). Each run is bounded by a tool-call budget
(max_tool_calls in the call, default 12, hard max 40). An asset worker's cumulative
estimated media spend is capped in dollars by --spawn-max-spend (default $2.00,
<= 0 disables; config defaults.code.spawn_max_spend); the report then also carries
spent_usd/spend_cap_usd.
Independently, --subagent-max-tokens N caps the cumulative prompt+completion tokens
a subagent spends across its turns (default off) — this applies to both the scout and
the worker, since token burn is universal to both, whereas the dollar cap above is
spawn-only (only a worker holds paid tools). Once the ceiling is crossed the subagent is
asked for a final answer and wraps up (the crossing turn completes, so the count can
slightly exceed the cap — like the spend cap); every report carries tokens/token_cap,
and under --planner the merge rollup sums totals.tokens and warns when a subagent hit
its cap. It's a cumulative-usage ceiling, not a context-window size limit, and is
distinct from --max-tokens (per-turn output). Config: defaults.code.subagent_max_tokens.
venice code --spawn --auto "Add a /health endpoint and a test for it. Spawn a code worker \
to implement and test it, then review its report before finishing."
Planner harness (--planner)
--planner turns the pieces above into one coherent workflow. It implies
--scout --spawn --memory and adds a protocol to the system prompt: decompose the
task into small self-contained units (task_add each one first), dispatch them
serially (task_update in progress → optional venice_scout → venice_spawn with the
unit's task_id → task_update done), then merge.
Merge is first-class, not prose: the harness records every launched scout/spawn dispatch
— its parsed report fields, task_id link, tool calls, spend, and whether it errored
or was truncated — and exposes venice_merge, a free tool that rolls all of it up
together with the task checklist and structural warnings
(a task not done or never dispatched, a dispatch that errored or hit its tool-call cap,
a task_id that matches no task). The planner is told to resolve those warnings and end
with a MERGE SUMMARY: section. With --json, the envelope carries the same rollup
under planner even if the model skipped the merge call. Workers can never hold
venice_merge (merging is the planner's job). Dispatch is serial by default; add
--parallel to let independent units overlap.
venice code --planner --auto --json "Split the CSV importer into reader/validator/writer \
modules with tests. Decompose into units, dispatch a worker per unit, and merge."
Parallel dispatch (--parallel)
By default a planner dispatches one subagent at a time, so three independent units run
back-to-back — the wall-clock is the sum of their nested loops. --parallel lets the
model emit several venice_scout/venice_spawn calls in a single turn and runs them
concurrently on a bounded pool (up to 4 at once), so independent units overlap and the
wall-clock drops to roughly the slowest single unit. It is opt-in (serial
otherwise) and only affects the two subagent tools — every other tool still runs serially,
and results are stitched back in the model's original order so the transcript is
deterministic.
The prompt overlay tells the planner to dispatch units together only when they are truly
independent (no unit needs another's output and no two touch the same files); dependent
units stay serial. Best paired with --planner (it is inert without a subagent rail).
Config: defaults.code.parallel.
venice code --planner --parallel --auto "Add unit tests for the parser, the formatter, \
and the CLI — three independent modules. Decompose, dispatch the independent units \
together, and merge."
MCP server
venice mcp-serve runs an MCP server over
stdio, exposing venice's generators as tools that an MCP host (Claude Code, or
any MCP client) can call directly instead of shelling out to the CLI. It needs
the [mcp] extra (Python ≥ 3.10):
pip install "venice-cli[mcp]"
# Register it with Claude Code:
claude mcp add venice -- venice mcp-serve
The server exposes nine tools:
| Tool | Does | Paid? |
|---|---|---|
venice_image |
generate image(s) → file path(s) | yes (estimated) |
venice_tts |
synthesize speech → audio file | yes (estimated) |
venice_sfx |
sound effect (async queue) → audio file | yes (quoted) |
venice_music |
long-form music/ambience (async queue) → audio file | yes (quoted) |
venice_video |
text/image-to-video (async queue, long-running) → video file | yes (quoted) |
venice_upscale |
upscale/enhance a local image → image file | yes (dynamic) |
venice_bg_remove |
remove a background → transparent PNG | yes (dynamic) |
venice_image_edit |
edit/inpaint an image (+ optional mask layers) → image file | yes (dynamic) |
venice_chat |
one-shot chat completion → reply text | no |
Spend gating. MCP is non-interactive, so instead of a [y/N] prompt the
paid tools gate on cost. A tool call whose estimated cost is at or under the
auto-approve cap (VENICE_MCP_MAX_SPEND, default $0.10) runs immediately.
If the estimate is over the cap — or can't be known up front, as with the
dynamically-priced venice_upscale / venice_bg_remove / venice_image_edit — the tool returns
{"status": "confirmation_required", ...} with the estimate and cap, and the
host must re-call with confirm=true (or a higher max_spend). Nothing is
spent and no file is written on a gated call. venice_chat is cheap and not
gated.
Output. Tools write their result to a file and return its path (never
inline base64). Files land in VENICE_MCP_OUTPUT_DIR (default: the current
working directory), or a per-call output_dir. The API key is read the usual
way ($VENICE_API_KEY or the credentials file) and is never echoed.
Only stdout carries the JSON-RPC protocol; the server's own diagnostics go to
stderr. Video generation and image editing are exposed over MCP too: the
venice_video and venice_image_edit tools cover the same capabilities as the
venice video and venice image-edit CLI commands.
The reverse direction — venice as an MCP client, calling other servers'
tools inside venice chat — is venice chat --mcp.
Config
venice config manages a persistent, non-secret config file at
~/.config/venice/config.json (created mode 0600). It holds two things: an
MCP server registry (attached by venice chat --mcp)
and default flag values so you stop repeating --model / -o / --yes /
--max-spend.
# MCP server registry (like `claude mcp add`)
venice config add venice --command venice --arg mcp-serve # stdio server
venice config add remote --url https://host/mcp --header 'Authorization: Bearer T'
venice config list
venice config show [NAME]
venice config remove venice
# Default flag values (dotted keys)
venice config set defaults.chat.model llama-3.3-70b
venice config set defaults.max_spend 0.50
venice config get defaults.chat.model
venice config unset defaults.chat.model
The file looks like:
{
"version": 1,
"mcpServers": {
"venice": { "command": "venice", "args": ["mcp-serve"] }
},
"defaults": {
"output_dir": "~/venice-out",
"max_spend": 0.50,
"chat": { "model": "llama-3.3-70b", "web_search": "auto" }
}
}
Global keys under defaults (output_dir, max_spend, yes, no_balance)
apply to any command that has the flag; a per-command section (e.g.
defaults.chat) overrides them. no_balance covers every spend-incurring
command at once -- --show-balance forces the display back on for a single run.
Precedence for any flag is: explicit CLI flag > environment variable > config
file > built-in default — so a config default never shadows
something you pass on the command line or set in the environment.
Per-command sections cover the persistent preferences of most commands — the knob is "if it expresses a preference (model, format, voice, sizing, style, safety), it should be settable in config." Currently config-backable:
defaults.image.*—width,height,aspect_ratio,resolution,style_prefix,preset,preset_file,negative_prompt,cfg_scale,steps,style_preset,hide_watermark,safe_mode(tri-state--safe-mode/--no-safe-mode; setfalseto skip Venice's safety blur)defaults.image_edit.*—model,aspect_ratio,resolution,output_format,safe_mode(tri-state--safe-mode/--no-safe-mode)defaults.tts.*—voice,speed,playdefaults.sfx.*—play,master,loop,no_cleanup(looponly takes effect whenmasteris also on -- it is a mastering-chain knob)defaults.music.*—duration,speed,play,instrumental,master,loop,no_cleanup(sameloopcaveat assfx)defaults.master.*—loop. This is the standalonevenice mastercommand's own section; thevenice sfx/venice music--mastertoggle is the keydefaults.sfx.master/defaults.music.master, not this tabledefaults.video.*—model,resolution,aspect_ratio,negative_prompt,no_audio,no_cleanupdefaults.upscale.*—enhance,enhance_creativity,enhance_prompt,replicationdefaults.chat.*,defaults.code.*,defaults.embed.*,defaults.index.*,defaults.search.*— see each command's section above
Per-invocation flags (--dry-run, --json, --resume, --seed, inputs and
positionals) stay CLI-only by design.
These per-command defaults also apply when a generator runs as an agent tool
inside venice chat --tools and venice code — e.g. defaults.image.safe_mode
is honored when the model calls venice_image, not just on the venice image
CLI. defaults.image_edit.safe_mode, defaults.music.instrumental,
defaults.video.no_audio and defaults.upscale.enhance reach the tools the same
way. An explicit argument the model puts in the tool call still wins over config.
venice mcp-serve threads the same defaults into its wrappers.
The API key is never stored here — it stays in
~/.config/venice/credentials. Unknown keys are preserved on write, so the
schema is forward-compatible.
Browse the model catalog
venice models # count by type
venice models --type music # list ids, one per line
venice models --type music --detail # ids + name + pricing + capabilities
venice models elevenlabs-sound-effects-v2 # full JSON for one model
venice models --type all --json # everything, raw
At time of writing the catalog spans ~258 models across text (80), code (30), image (26), video (92), music+sfx (10), tts (10), embedding (9), and upscale (1). The video models include Sora 2, Veo 3.1, Kling, Runway Gen4, LTX-2, Wan 2.7, Seedance 2.0, and more.
Exit codes
| exit | meaning |
|---|---|
| 0 | success |
| 1 | user declined / aborted / insufficient balance (402) |
| 2 | bad input, no API key, missing prompt, stub command |
| 3 | content policy block (422) |
| 4 | rate limit (429) |
| 5 | Venice 5xx |
| 6 | job not found / expired (404) |
| 7 | poll timeout |
| 8 | network / connection error |
| 9 | disk write error |
| 10 | acceptance verdict unparseable / ambiguous (venice code) |
| 130 | Ctrl-C |
Audio playback caveat
Auto-play depends on whatever player your system has. If all you have
is paplay, note that it plays WAV natively but handles MP3 (Venice's
default output) via PulseAudio's GStreamer plumbing, which can fail
silently. If playback fails the file is still saved; the CLI just
won't auto-play it. For reliable MP3 playback in-CLI, install one of:
sudo apt install mpg123 # or: ffmpeg
The player list (paplay -> aplay -> ffplay -> mpg123 -> play
-> afplay) auto-picks the first available.
Environment overrides
| var | meaning |
|---|---|
VENICE_API_KEY |
overrides the file-based key (no disk read) |
VENICE_BASE_URL |
override the API base URL (testing, proxy) |
VENICE_EMBED_BASE_URL |
embed alternate OpenAI-compatible endpoint (local backend) |
VENICE_EMBED_API_KEY |
key for VENICE_EMBED_BASE_URL (if the backend needs one) |
VENICE_EMBED_CA_BUNDLE |
CA bundle to trust for a self-signed embedding backend (embed, index, search, and the project_search agent tool) |
VENICE_MCP_MAX_SPEND |
mcp-serve auto-approve cap in USD (default 0.10) |
VENICE_MCP_OUTPUT_DIR |
where mcp-serve tools write files (default: cwd) |
Commands at a glance
| command | does what |
|---|---|
venice login [--embed] |
store the Venice API key (interactive, hidden, mode 0600); --embed stores the embed-backend key in the secret store instead |
venice secret set/ls/rm NAME |
manage named secrets in ~/.config/venice/secrets.json (0600); values never printed, ls shows lengths only |
venice balance [--verbose|--json|--min N] |
current USD + DIEM balance |
venice models [--type T] [--detail] [SLUG] |
browse the catalog |
venice sfx PROMPT [--duration N] [--max-spend USD] [...] |
generate a sound effect |
venice sfx-status QUEUE_ID |
fetch a backgrounded SFX job |
venice tts TEXT [--voice V] [--format F] [--speed N] [...] |
synthesize speech (sync) |
venice image PROMPT [--variants N] [--name NAME] [--max-spend USD] [...] |
generate image(s) (sync) |
venice image --from-file PATH [...] |
batch-generate a card set |
venice music PROMPT [--duration N] [--master] [--loop] [...] |
generate long-form ambience/music |
venice video PROMPT [--duration 5s] [--resolution R] [--aspect-ratio A] [--image F] [--reference-image F ...] [--element JSON] [...] |
generate a video (async queue, mp4); text- or image-to-video with reference inputs |
venice video-status QUEUE_ID [--download-url URL] |
fetch a backgrounded video job |
venice master INPUT [--loop] [--lufs N] [--bit-depth N] [...] |
master audio to WAV (48k/24-bit, LUFS/true-peak) |
venice contact-sheet DIR_OR_GLOB [--cols N] [--cell WxH] [--label] [...] |
tile images into one contact sheet (no API call) |
venice chat MESSAGE [--system S] [--model M] [--web-search on] [...] |
one-shot chat completion (OpenAI SDK) |
venice chat [-i] [--continue|--resume ID|FILE] [--ephemeral] |
interactive multi-turn REPL (auto-saved sessions, /-commands, transcripts) |
venice sessions ls|show|rm [ID] |
list/inspect/remove auto-saved chat & code sessions (~/.config/venice/sessions/, 0600) |
venice embed [TEXT] [--from-file PATH] [--model M] [--dimensions N] [--json] [--embed-base-url URL --embed-model M [--embed-ca-bundle PATH | --embed-insecure]] |
text embeddings (OpenAI SDK; alt/local backend) |
venice index [PATH] [--model M] [--embed-base-url URL --embed-model M [--embed-ca-bundle PATH | --embed-insecure]] [...] / venice search QUERY [-k N] [--json] [--embed-ca-bundle PATH | --embed-insecure] |
build / query a local semantic index of a project tree |
venice code [TASK] [--auto|--manual] [--plan-only] [-i] [--root DIR] [--continue|--resume ID|FILE] [--ephemeral] [--json] [...] |
coding agent: plan → accept → edit/run a project (needs [openai] + tool-calling model) |
venice mcp-serve |
run an MCP server (stdio) exposing venice tools (needs [mcp]) |
venice config add|list|remove|show |
manage the MCP server registry |
venice config get|set|unset KEY [VALUE] |
manage default flag values |
venice completion bash|zsh |
print a shell tab-completion script (generated from the parser) |
Tests
make test # everything, including the drive suite
make drive # the drive suite + its fake-API fixture
Stdlib unittest only. Most tests mock urlopen (and, for chat, the OpenAI
client) and patch HOME to a tmpdir -- no live API calls, no real disk writes
outside the tmpdir. The chat and embed tests need the OpenAI SDK importable
(pip install -e ".[openai]").
On top of that, tests/test_drive_cli.py drives the real CLI over a pty
with pexpect, pointed at a local fake API via $VENICE_BASE_URL. It asserts
on actual terminal output, prompts, and exit codes -- including two interleaved
multi-step dialogues (the chat REPL and the venice code plan gate), because
that's where interactive breakage hides. It needs the test extra
(pip install -e ".[all,test]"); without it the pty cases skip cleanly under
make test, while make drive tells you the dep is missing rather than
reporting a green run of nothing.
Uninstall
If you installed with pip:
pip uninstall venice-cli
If you installed from source with ./install.sh:
./uninstall.sh
Either way the credentials file at ~/.config/venice/credentials is left
alone -- delete it manually if you want. uninstall.sh removes only the two
symlinks, and only if they point at that repo.
Security note
The API key is stored plaintext on disk at
~/.config/venice/credentials (mode 0600, inside a 0700 directory).
There is no OS keychain integration -- file permissions are the only
protection, so anything that can read your home directory can read the
key. The venice config file (~/.config/venice/config.json) is written
mode 0600 for the same reason -- an MCP env/headers entry can carry a
bearer token -- but the API key itself is never written there.
Secrets (venice secret)
Named secrets other than the main Venice key -- currently the embed-backend key,
later MCP/cluster tokens -- live in a structured ~/.config/venice/secrets.json
(mode 0600, same plaintext-file model as credentials), so they don't have to sit in
your shell profile or plaintext config.json:
venice secret set embed # hidden prompt; stores the 'embed' secret
venice secret ls # names + lengths only, never values
venice secret rm embed
venice login --embed # convenience alias for `secret set embed`
A secret is only ever read from a hidden prompt (never on argv) and is never
printed back -- there is deliberately no command that outputs a value; ls shows
character counts. For a name with a canonical env var (embed →
$VENICE_EMBED_API_KEY), the env var still overrides the stored value, so CI can keep
injecting it from that system's secret store. The same plaintext caveat applies:
file permissions are the only protection.
- In CI or any shared environment, prefer
$VENICE_API_KEY(it overrides the file) sourced from that system's secret store, and don't runvenice loginthere. - The key is never logged or printed by this tool, but it is visible to anything reading your process environment or scrollback. Be aware of what's on screen when sharing a terminal.
- If a key is exposed, revoke and rotate it at https://venice.ai/settings/api.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file venice_cli-0.71.0.tar.gz.
File metadata
- Download URL: venice_cli-0.71.0.tar.gz
- Upload date:
- Size: 438.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9992ed52b4ee06a5086aac3e23fcfecb8d475f3c76d489468684eb658879e210
|
|
| MD5 |
dde0212b6e81057a7ff39b096c7e1d3b
|
|
| BLAKE2b-256 |
30f786577e855e76365bc14afd82e69316e1b699ce52e33e10e43eb026f7c94c
|
Provenance
The following attestation bundles were made for venice_cli-0.71.0.tar.gz:
Publisher:
publish.yml on gobha-me/venice-cli
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
venice_cli-0.71.0.tar.gz -
Subject digest:
9992ed52b4ee06a5086aac3e23fcfecb8d475f3c76d489468684eb658879e210 - Sigstore transparency entry: 2254666871
- Sigstore integration time:
-
Permalink:
gobha-me/venice-cli@d2c109ebd382f6b03a6d24a2334744adf4e46901 -
Branch / Tag:
refs/tags/v0.71.0 - Owner: https://github.com/gobha-me
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@d2c109ebd382f6b03a6d24a2334744adf4e46901 -
Trigger Event:
release
-
Statement type:
File details
Details for the file venice_cli-0.71.0-py3-none-any.whl.
File metadata
- Download URL: venice_cli-0.71.0-py3-none-any.whl
- Upload date:
- Size: 280.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/6.1.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
da5b4167f3493e7ae4c666870312da1e9c4ecec1619ca05aac7d640ea84b2ae7
|
|
| MD5 |
a53d9e380a0e4538bc1b2d2c3a797b3d
|
|
| BLAKE2b-256 |
6fa201539bc6a7da133b620b91ac4f5a9fc80b3478e16a372222abe88a7eaf39
|
Provenance
The following attestation bundles were made for venice_cli-0.71.0-py3-none-any.whl:
Publisher:
publish.yml on gobha-me/venice-cli
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
venice_cli-0.71.0-py3-none-any.whl -
Subject digest:
da5b4167f3493e7ae4c666870312da1e9c4ecec1619ca05aac7d640ea84b2ae7 - Sigstore transparency entry: 2254667012
- Sigstore integration time:
-
Permalink:
gobha-me/venice-cli@d2c109ebd382f6b03a6d24a2334744adf4e46901 -
Branch / Tag:
refs/tags/v0.71.0 - Owner: https://github.com/gobha-me
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@d2c109ebd382f6b03a6d24a2334744adf4e46901 -
Trigger Event:
release
-
Statement type: