glm-launch
A Python CLI tool that wraps Claude Code with GLM settings. Instead of running a local proxy, it configures environment variables, then exec's the claude binary directly. (codex is not supported — Z.AI has no OpenAI Responses API endpoint.)
It works with Z.AI and their GLM series of models. You'll need a Z.AI API key — grab one with a Z.AI Coding Plan subscription. Using that referral link gives you 10% off and gets me 10% off too. Prefer not to? Here's a non-affiliate link.
Requires Python 3.13+.
Usage
# 1. Set your Z.AI auth token
export GLM_AUTH_TOKEN="your-zai-api-key"
# 2. Launch Claude Code routed through Z.AI (defaults to glm-5.3)
uv run glm-launch # bare command defaults to `claude`
uv run glm-launch claude # same thing, explicit
# Pick a different model
uv run glm-launch --model glm-4.7 # bare options also go to `claude`
uv run glm-launch claude --model glm-4.7 # balanced cost/performance
uv run glm-launch claude --model glm-5-turbo # fast
uv run glm-launch claude --model glm-4.5-air # cheap
# Bootstrap your current shell so a plain `claude` uses Z.AI
eval "$(uv run glm-launch shell)"
claude
# See available models (built-in list, or --remote for the live API list)
uv run glm-launch models
uv run glm-launch models --remote
# Sanity-check connectivity / latency
uv run glm-launch bench
Examples use the installed
glm-launchentrypoint. Beforeuv syncyou can run the script directly withuv run src/main.py …— the two are interchangeable.
Installation
uv sync
This installs a glm-launch entrypoint. Run commands via uv run glm-launch <command>, or uv tool install . to get glm-launch on your PATH directly. You can also run the script without installing via uv run src/main.py <command>.
Run without cloning (uvx)
glm-launch is on PyPI, so you can run it directly with uvx (uv tool run) — no clone or manual install needed.
# From PyPI
uvx glm-launch launch claude
# Or straight from GitHub
uvx --from git+https://github.com/jefftriplett/glm-launch glm-launch launch claude
# Pin to a tag/branch/commit
uvx --from git+https://github.com/jefftriplett/glm-launch@main glm-launch models
Commands
launch claude
Launch Claude Code with GLM environment settings. Sets Anthropic env vars to route requests through Z.AI's Anthropic-compatible endpoint, then exec's the claude binary.
The
launchprefix is optional:glm-launch claudeis equivalent toglm-launch launch claude, and a bareglm-launchdefaults toclaude. Claude options can also be passed directly, soglm-launch --model glm-4.7is equivalent toglm-launch claude --model glm-4.7.
uv run glm-launch launch claude
Options:
| Flag | Env var | Default | Description |
|---|---|---|---|
--model / -m |
— | glm-5.3 |
Model name passed to claude --model; glm-5.3 serves 1M context natively, older models need the [1m] suffix for the 1M tier |
--base-url |
GLM_BASE_URL |
https://api.z.ai/api/anthropic |
API endpoint |
--api-key |
GLM_API_KEY |
"" |
API key |
--auth-token |
GLM_AUTH_TOKEN |
(required) | Z.AI auth token |
--api-timeout-ms |
API_TIMEOUT_MS |
3000000 |
Request timeout in milliseconds (positive integer) |
--default-haiku-model |
ANTHROPIC_DEFAULT_HAIKU_MODEL |
glm-4.5-air |
Model for Haiku-tier requests |
--default-sonnet-model |
ANTHROPIC_DEFAULT_SONNET_MODEL |
glm-5.3 |
Model for Sonnet-tier requests |
--default-opus-model |
ANTHROPIC_DEFAULT_OPUS_MODEL |
glm-5.3 |
Model for Opus-tier requests |
--default-fable-model |
ANTHROPIC_DEFAULT_FABLE_MODEL |
glm-5.3 |
Model for Fable-tier requests |
--subagent-model |
CLAUDE_CODE_SUBAGENT_MODEL |
glm-4.5-air |
Model used for spawned subagents |
--effort-level |
CLAUDE_CODE_EFFORT_LEVEL |
max |
Effort level for the agent loop: low, medium, high, xhigh, max, or ultracode (see Effort levels) |
--attribution-header |
CLAUDE_CODE_ATTRIBUTION_HEADER |
0 |
Attribution header toggle (0 or 1; 0 disables it) |
--auto-compact-window |
CLAUDE_CODE_AUTO_COMPACT_WINDOW |
auto |
Auto-compact context window (auto, empty, or a positive integer) |
--max-context-tokens |
CLAUDE_CODE_MAX_CONTEXT_TOKENS |
auto |
Maximum context budget (auto, empty, or a positive integer) |
--dry-run |
— | false |
Print the resolved command and masked GLM environment without launching |
The following env vars are set before exec'ing claude:
ANTHROPIC_BASE_URL— from--base-url/GLM_BASE_URLANTHROPIC_API_KEY— from--api-key/GLM_API_KEYANTHROPIC_AUTH_TOKEN— from--auth-token/GLM_AUTH_TOKENAPI_TIMEOUT_MS— from--api-timeout-ms/API_TIMEOUT_MSANTHROPIC_DEFAULT_HAIKU_MODEL— from--default-haiku-modelANTHROPIC_DEFAULT_SONNET_MODEL— from--default-sonnet-modelANTHROPIC_DEFAULT_OPUS_MODEL— from--default-opus-modelANTHROPIC_DEFAULT_FABLE_MODEL— from--default-fable-modelCLAUDE_CODE_SUBAGENT_MODEL— from--subagent-modelCLAUDE_CODE_EFFORT_LEVEL— from--effort-levelCLAUDE_CODE_ATTRIBUTION_HEADER— from--attribution-headerCLAUDE_CODE_AUTO_COMPACT_WINDOW— from--auto-compact-window(only when non-empty)CLAUDE_CODE_MAX_CONTEXT_TOKENS— from--max-context-tokens(only when non-empty)
[!NOTE] With the default
auto, the context settings are sized to the selected--modelautomatically:glm-5.3and the[1m]IDs get 1M tokens, most other models 200K, andglm-4.5/glm-4.5-air128K (unknown models fall back to 200K).glm-5.3serves the 1M window natively; for other models the[1m]suffix is what enables Z.AI's 1M context tier — plainglm-5.2andglm-5.3-flashserve the standard 200K window. Pass an explicit number to override, or an empty string to leave the env vars unset. Runglm-launch modelsto see each model's window.
Effort levels
Z.AI collapses Claude Code's effort ladder into the model's thinking tiers (source):
| Claude Code effort | GLM-5.3 effort | GLM-5.2 effort |
|---|---|---|
low |
low (light) |
high |
medium, high |
high (enhanced) |
high |
xhigh, max, ultracode |
max (deep) |
max |
GLM-5.3 no longer supports disabling thinking entirely — low is the
lightest setting. Z.AI recommends max for coding, which is the default
here. You can also switch mid-session with the /effort command in
Claude Code.
[!TIP] The GLM coding models are text-only — pasting images into Claude Code won't work through Z.AI. Coding Plan subscribers get image understanding via Z.AI's Vision MCP server (backed by
glm-4.6v) instead; see #3. Also note that Team Plan API keys are separate from regular Z.AI keys — only a Team key draws Team quota, so a mismatched key can look like an auth failure.
Examples:
# Use defaults (glm-5.3 with 1M context, Z.AI endpoint)
uv run glm-launch launch claude
# The flagship (the default) — 1M context is standard on glm-5.3
uv run glm-launch launch claude --model glm-5.3
# Native multimodal (video/image/text/file) at a much lower cost,
# with 3x the coding-plan quota of glm-5.3
uv run glm-launch launch claude --model glm-5.3-flash
# Previous flagship with the 1M context tier (the coding plan
# auto-routes glm-5.2/glm-5.1 requests to glm-5.3)
uv run glm-launch launch claude --model "glm-5.2[1m]"
# Previous flagship on the standard 200K window (cheaper)
uv run glm-launch launch claude --model glm-5.2
# Balanced cost/performance coding model
uv run glm-launch launch claude --model glm-4.7
# Fast, speed-optimized GLM-5 variant
uv run glm-launch launch claude --model glm-5-turbo
# Lightweight, low-cost model for cheaper runs
uv run glm-launch launch claude --model glm-4.5-air
# Tune the model tiers independently (e.g. cheap subagents, flagship main)
uv run glm-launch launch claude \
--model glm-5.3 \
--subagent-model glm-4.5-air \
--default-haiku-model glm-4.5-air
# Pass extra args through to claude
uv run glm-launch launch claude -- --verbose
# Inspect the command/env without launching claude
uv run glm-launch launch claude --dry-run
# Override via env vars
GLM_AUTH_TOKEN="my-token" uv run glm-launch launch claude
--dry-run does not require the claude binary to be installed.
Run uv run glm-launch models to see all valid model names (or --remote for the live list).
If claude is not on your PATH, the tool falls back to ~/.claude/local/claude.
launch codex (not supported)
Codex is not supported by glm-launch. Current codex only speaks the OpenAI Responses API (it removed wire_api = "chat"), but Z.AI's GLM endpoints are Anthropic Messages and OpenAI Chat Completions only — there is no /responses endpoint, so codex requests return 404. The codex command is intentionally disabled and exits with this explanation.
Use launch claude instead — it uses Z.AI's Anthropic-compatible endpoint. If Z.AI later ships a Responses-compatible endpoint, codex support can be revisited.
shell
Print export lines that bootstrap your current shell with the GLM env vars — without launching anything. Eval the output and a plain claude (or any Anthropic SDK tool) will talk to Z.AI.
eval "$(uv run glm-launch shell)"
claude
Accepts the same model/auth options as launch claude (--model, --auth-token, --default-*-model, etc.). Secrets are shell-quoted; empty values are skipped. Sets ANTHROPIC_MODEL plus all the ANTHROPIC_* / CLAUDE_CODE_* vars listed under launch claude.
# Inspect what would be exported
uv run glm-launch shell
# Bootstrap with a specific model
eval "$(uv run glm-launch shell --model glm-4.7)"
models
List Z.AI GLM models. By default prints a built-in, annotated list; --remote fetches the live list from the Z.AI PaaS endpoint.
# Built-in list (no token needed)
uv run glm-launch models
# Live list from the API (needs GLM_AUTH_TOKEN)
uv run glm-launch models --remote
Options:
| Flag | Env var | Default | Description |
|---|---|---|---|
--remote / -r |
— | false |
Fetch the live list from the Z.AI API |
--models-url |
GLM_MODELS_URL |
https://api.z.ai/api/coding/paas/v4/models |
PaaS models endpoint (used with --remote) |
--auth-token |
GLM_AUTH_TOKEN |
— | Auth token (required with --remote) |
--timeout |
— | 30.0 |
Request timeout in seconds (must be greater than zero) |
The live endpoint is the OpenAI-compatible coding PaaS base (/api/coding/paas/v4/models) and uses Authorization: Bearer <token> — distinct from the Anthropic-style chat base (/api/anthropic) used by launch claude and bench. Coding Plan keys only work through the coding endpoints; if you have a general Z.AI API key instead, point --models-url at https://api.z.ai/api/paas/v4/models.
bench
Time a single /v1/messages round-trip against the configured GLM endpoint. Useful as a sanity check that your auth token, base URL, and chosen model are reachable.
uv run glm-launch bench
Options:
| Flag | Env var | Default | Description |
|---|---|---|---|
--model / -m |
— | glm-5.3 |
Model to benchmark |
--base-url |
GLM_BASE_URL |
https://api.z.ai/api/anthropic |
API endpoint |
--auth-token |
GLM_AUTH_TOKEN |
(required) | Auth token for the endpoint |
--timeout |
— | 30.0 |
Request timeout in seconds (must be greater than zero) |
--all |
— | false |
Probe every model in the registry instead of just --model |
Sends a minimal 32-token request and prints the round-trip time. Exits non-zero on HTTP error or timeout.
Example output:
glm-5.3 via https://api.z.ai/api/anthropic
OK (200) in 412ms
Verifying every model ID
glm-launch models prints the built-in registry, and models --remote lists
what the API advertises — but neither proves a given ID is callable with your
key. Z.ai rejects an unknown or unentitled model with HTTP 400
modelCode: does not exist, and the [1m] context IDs are a naming convention
that never appears in the API's model list at all. bench --all is the check
that actually calls each one:
uv run glm-launch bench --all
probing 14 models via https://api.z.ai/api/anthropic
ok glm-5.3 200 1980ms
FAIL glm-5.3-flash[1m] 400 427ms
ok glm-5.3-flash 200 1686ms
SKIP glm-5v-turbo 429 752ms
...
1 model(s) were rate limited and not verified: glm-5v-turbo
2 of 14 model(s) failed to resolve.
Rejected as `modelCode: does not exist`: glm-5.3-flash[1m], glm-5.2[1m]
A 429 is reported as SKIP, not a failure — it means the ID resolved but the
key is out of quota, which says nothing about whether the model exists. Only
IDs the API actually rejects count toward the non-zero exit.
The same probe is available as an opt-in test suite:
GLM_LIVE_TESTS=1 uv run pytest -m live -v
These are skipped by default (and in CI) since they need a real token and hit the network.
usage
Open the Z.AI usage/quota dashboard in your browser. Coding Plan quotas are tracked in 5-hour and weekly windows, and there is no API for quota data — the dashboard is the only place to see it.
uv run glm-launch usage
doctor
Check your environment for correct setup. Reports on environment variables and binary availability.
uv run glm-launch doctor
Checks performed:
- Authentication — Whether the required
GLM_AUTH_TOKENis set. The token is masked in output. - Environment variables — Whether the optional GLM, Anthropic default-model, and Claude Code env vars used by the launch commands are set.
- Binaries — Whether
claudeis found on PATH (with fallback to~/.claude/local/claude), including its version —[1m]-suffixed models need a recent Claude Code, so if claude reports the model doesn't exist, upgrade.
Exits with code 1 if GLM_AUTH_TOKEN is not set or the claude binary is missing, 0 otherwise.
Example output:
Environment variables:
GLM_BASE_URL: (not set)
GLM_API_KEY: (not set)
GLM_AUTH_TOKEN: zai_***
GLM_MODELS_URL: (not set)
API_TIMEOUT_MS: (not set)
ANTHROPIC_DEFAULT_HAIKU_MODEL: (not set)
ANTHROPIC_DEFAULT_SONNET_MODEL: (not set)
ANTHROPIC_DEFAULT_OPUS_MODEL: (not set)
ANTHROPIC_DEFAULT_FABLE_MODEL: (not set)
CLAUDE_CODE_SUBAGENT_MODEL: (not set)
CLAUDE_CODE_EFFORT_LEVEL: (not set)
CLAUDE_CODE_ATTRIBUTION_HEADER: (not set)
CLAUDE_CODE_AUTO_COMPACT_WINDOW: (not set)
CLAUDE_CODE_MAX_CONTEXT_TOKENS: (not set)
Binaries:
claude: /usr/local/bin/claude
All checks passed.
Environment variables
| Variable | Used by | Description |
|---|---|---|
GLM_BASE_URL |
launch claude, shell |
API base URL |
GLM_API_KEY |
launch claude, shell |
API key |
GLM_AUTH_TOKEN |
launch claude, shell, bench, models --remote |
Z.AI auth token (required) |
GLM_MODELS_URL |
models --remote |
PaaS models endpoint |
API_TIMEOUT_MS |
launch claude, shell |
Request timeout in milliseconds (positive integer) |
ANTHROPIC_DEFAULT_HAIKU_MODEL |
launch claude, shell |
Model for Haiku-tier requests |
ANTHROPIC_DEFAULT_SONNET_MODEL |
launch claude, shell |
Model for Sonnet-tier requests |
ANTHROPIC_DEFAULT_OPUS_MODEL |
launch claude, shell |
Model for Opus-tier requests |
ANTHROPIC_DEFAULT_FABLE_MODEL |
launch claude, shell |
Model for Fable-tier requests |
CLAUDE_CODE_SUBAGENT_MODEL |
launch claude, shell |
Model used for spawned subagents |
CLAUDE_CODE_EFFORT_LEVEL |
launch claude, shell |
Validated effort level for the agent loop |
CLAUDE_CODE_ATTRIBUTION_HEADER |
launch claude, shell |
Attribution header toggle (0 or 1) |
CLAUDE_CODE_AUTO_COMPACT_WINDOW |
launch claude, shell |
auto, empty, or a positive token count |
CLAUDE_CODE_MAX_CONTEXT_TOKENS |
launch claude, shell |
auto, empty, or a positive token count |
How it works
launch claude follows three steps:
- Resolve the
claudebinary on PATH (falling back to~/.claude/local/claude) - Set up the GLM environment variables
os.execvpe()the binary — fully replacing the glm process withclaudefor direct stdio passthrough
Z.AI exposes an Anthropic-compatible endpoint at https://api.z.ai/api/anthropic, so no local proxy is needed. The CLI sets the standard ANTHROPIC_* env vars and Claude Code talks directly to Z.AI.
Development
Common tasks are wrapped in a justfile. Run just with no arguments to list them.
Before committing, run the same core checks used by CI:
uv run pytest
uv tool run prek run --all-files
uv build
| Recipe | Description |
|---|---|
just bootstrap |
Upgrade pip/uv, then uv sync |
just sync |
uv sync the project dependencies |
just lock |
uv lock the dependency versions |
just build |
uv build the wheel and sdist |
just bump *ARGS |
Bump the CalVer version with bumpver (e.g. just bump) |
just bump-dry *ARGS |
Preview a version bump without writing changes |
just release *ARGS |
Bump, relock, and push the tag — CI then publishes to PyPI |
just lint *ARGS |
Run the prek hooks (defaults to --all-files) |
just fmt |
Format the justfile itself |
just demo |
Smoke-test the CLI by listing models |
Versioning follows CalVer (YYYY.MM.INC1), and lint hooks (ruff, pyupgrade, validate-pyproject) are configured in .pre-commit-config.yaml and run with prek. CI runs the tests, lint checks, and package build; the release workflow reruns the tests before publishing.
Releases are automated. Run just release to bump the CalVer version, relock, and push the tag in one step. Pushing a YYYY.MM.INC1 tag triggers the GitHub Actions release workflow, which builds and publishes to PyPI via trusted publishing (OIDC, no API token). A plain git push never publishes — only the tag does.
Metadata
Release files for glm-launch 2026.8.5
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| glm_launch-2026.8.5.tar.gz | 25.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| glm_launch-2026.8.5-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 42.2 kB
Release files / glm_launch-2026.8.5.tar.gz
| Download URL | glm_launch-2026.8.5.tar.gz |
|---|---|
| Size | 25.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
53beca31deb136542495c19d2b93e76dbb7e05021e03c3903c1e3229fc7c4b0a
|
|
BLAKE2b-256 checksum How to use checksums |
d7c2dc1f00ea58bef3871f5136ed5ebcb9836584901504ed210aaa1970677e64
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 31, 2026.
Transparency logRelease files / glm_launch-2026.8.5-py3-none-any.whl
| Download URL | glm_launch-2026.8.5-py3-none-any.whl |
|---|---|
| Size | 16.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
dc7825046a2d4bd044789efa0a0c86afa03a0a351c6398c1a082a11c23987110
|
|
BLAKE2b-256 checksum How to use checksums |
5940868a2155b9ebc44b5c6a9571c6531950a6f2841f683ba544a569114769e4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 31, 2026.
Transparency log