Skip to main content

Wrap Claude Code with Z.ai GLM settings

Project description

glm-launch

A Python CLI tool that wraps Claude Code with GLM settings. Instead of running a local proxy, it configures environment variables, then exec's the claude binary directly. (codex is not supported — Z.AI has no OpenAI Responses API endpoint.)

It works with Z.AI and their GLM series of models. You'll need a Z.AI API key — grab one with a Z.AI Coding Plan subscription. Using that referral link gives you 10% off and gets me 10% off too. Prefer not to? Here's a non-affiliate link.

Requires Python 3.13+.

Usage

# 1. Set your Z.AI auth token
export GLM_AUTH_TOKEN="your-zai-api-key"

# 2. Launch Claude Code routed through Z.AI (defaults to glm-5.2[1m])
uv run glm-launch              # bare command defaults to `claude`
uv run glm-launch claude       # same thing, explicit

# Pick a different model
uv run glm-launch --model glm-5.1               # bare options also go to `claude`
uv run glm-launch claude --model glm-5.1        # long-horizon flagship
uv run glm-launch claude --model glm-5-turbo    # fast
uv run glm-launch claude --model glm-4.5-air    # cheap

# Bootstrap your current shell so a plain `claude` uses Z.AI
eval "$(uv run glm-launch shell)"
claude

# See available models (built-in list, or --remote for the live API list)
uv run glm-launch models
uv run glm-launch models --remote

# Sanity-check connectivity / latency
uv run glm-launch bench

Examples use the installed glm-launch entrypoint. Before uv sync you can run the script directly with uv run src/main.py … — the two are interchangeable.

Installation

uv sync

This installs a glm-launch entrypoint. Run commands via uv run glm-launch <command>, or uv tool install . to get glm-launch on your PATH directly. You can also run the script without installing via uv run src/main.py <command>.

Run without cloning (uvx)

glm-launch is on PyPI, so you can run it directly with uvx (uv tool run) — no clone or manual install needed.

# From PyPI
uvx glm-launch launch claude

# Or straight from GitHub
uvx --from git+https://github.com/jefftriplett/glm-launch glm-launch launch claude

# Pin to a tag/branch/commit
uvx --from git+https://github.com/jefftriplett/glm-launch@main glm-launch models

Commands

launch claude

Launch Claude Code with GLM environment settings. Sets Anthropic env vars to route requests through Z.AI's Anthropic-compatible endpoint, then exec's the claude binary.

The launch prefix is optional: glm-launch claude is equivalent to glm-launch launch claude, and a bare glm-launch defaults to claude. Claude options can also be passed directly, so glm-launch --model glm-5.1 is equivalent to glm-launch claude --model glm-5.1.

uv run glm-launch launch claude

Options:

Flag Env var Default Description
--model / -m glm-5.2[1m] Model name passed to claude --model; the [1m] suffix enables the 1M context tier
--base-url GLM_BASE_URL https://api.z.ai/api/anthropic API endpoint
--api-key GLM_API_KEY "" API key
--auth-token GLM_AUTH_TOKEN (required) Z.AI auth token
--api-timeout-ms API_TIMEOUT_MS 3000000 Request timeout in milliseconds (positive integer)
--default-haiku-model ANTHROPIC_DEFAULT_HAIKU_MODEL glm-4.5-air Model for Haiku-tier requests
--default-sonnet-model ANTHROPIC_DEFAULT_SONNET_MODEL glm-5.2[1m] Model for Sonnet-tier requests
--default-opus-model ANTHROPIC_DEFAULT_OPUS_MODEL glm-5.2[1m] Model for Opus-tier requests
--default-fable-model ANTHROPIC_DEFAULT_FABLE_MODEL glm-5.2[1m] Model for Fable-tier requests
--subagent-model CLAUDE_CODE_SUBAGENT_MODEL glm-4.5-air Model used for spawned subagents
--effort-level CLAUDE_CODE_EFFORT_LEVEL max Effort level for the agent loop: low, medium, high, xhigh, max, or ultracode (see Effort levels)
--attribution-header CLAUDE_CODE_ATTRIBUTION_HEADER 0 Attribution header toggle (0 or 1; 0 disables it)
--auto-compact-window CLAUDE_CODE_AUTO_COMPACT_WINDOW auto Auto-compact context window (auto, empty, or a positive integer)
--max-context-tokens CLAUDE_CODE_MAX_CONTEXT_TOKENS auto Maximum context budget (auto, empty, or a positive integer)
--dry-run false Print the resolved command and masked GLM environment without launching

The following env vars are set before exec'ing claude:

  • ANTHROPIC_BASE_URL — from --base-url / GLM_BASE_URL
  • ANTHROPIC_API_KEY — from --api-key / GLM_API_KEY
  • ANTHROPIC_AUTH_TOKEN — from --auth-token / GLM_AUTH_TOKEN
  • API_TIMEOUT_MS — from --api-timeout-ms / API_TIMEOUT_MS
  • ANTHROPIC_DEFAULT_HAIKU_MODEL — from --default-haiku-model
  • ANTHROPIC_DEFAULT_SONNET_MODEL — from --default-sonnet-model
  • ANTHROPIC_DEFAULT_OPUS_MODEL — from --default-opus-model
  • ANTHROPIC_DEFAULT_FABLE_MODEL — from --default-fable-model
  • CLAUDE_CODE_SUBAGENT_MODEL — from --subagent-model
  • CLAUDE_CODE_EFFORT_LEVEL — from --effort-level
  • CLAUDE_CODE_ATTRIBUTION_HEADER — from --attribution-header
  • CLAUDE_CODE_AUTO_COMPACT_WINDOW — from --auto-compact-window (only when non-empty)
  • CLAUDE_CODE_MAX_CONTEXT_TOKENS — from --max-context-tokens (only when non-empty)

[!NOTE] With the default auto, the context settings are sized to the selected --model automatically: glm-5.2[1m] gets 1M tokens, most other models 200K, and glm-4.5/glm-4.5-air 128K (unknown models fall back to 200K). The [1m] suffix is what enables Z.AI's 1M context tier — plain glm-5.2 serves the standard 200K window. Pass an explicit number to override, or an empty string to leave the env vars unset. Run glm-launch models to see each model's window.

Effort levels

GLM-5.2 collapses Claude Code's effort ladder into two effective tiers (source):

Claude Code effort GLM-5.2 actual effort
low, medium, high high
xhigh, max, ultracode max

So --effort-level is effectively a two-position switch: high (faster) or max (deeper reasoning). Z.AI recommends max for coding, which is the default here. You can also switch mid-session with the /effort command in Claude Code.

[!TIP] The GLM coding models are text-only — pasting images into Claude Code won't work through Z.AI. Coding Plan subscribers get image understanding via Z.AI's Vision MCP server (backed by glm-4.6v) instead; see #3. Also note that Team Plan API keys are separate from regular Z.AI keys — only a Team key draws Team quota, so a mismatched key can look like an auth failure.

Examples:

# Use defaults (glm-5.2[1m] with 1M context, Z.AI endpoint)
uv run glm-launch launch claude

# Flagship with the 1M context tier (the default)
uv run glm-launch launch claude --model "glm-5.2[1m]"

# Flagship on the standard 200K window (cheaper)
uv run glm-launch launch claude --model glm-5.2

# Long-horizon agentic flagship
uv run glm-launch launch claude --model glm-5.1

# Fast, speed-optimized GLM-5 variant
uv run glm-launch launch claude --model glm-5-turbo

# Lightweight, low-cost model for cheaper runs
uv run glm-launch launch claude --model glm-4.5-air

# Tune the model tiers independently (e.g. cheap subagents, flagship main)
uv run glm-launch launch claude \
  --model "glm-5.2[1m]" \
  --subagent-model glm-4.5-air \
  --default-haiku-model glm-4.5-air

# Pass extra args through to claude
uv run glm-launch launch claude -- --verbose

# Inspect the command/env without launching claude
uv run glm-launch launch claude --dry-run

# Override via env vars
GLM_AUTH_TOKEN="my-token" uv run glm-launch launch claude

--dry-run does not require the claude binary to be installed.

Run uv run glm-launch models to see all valid model names (or --remote for the live list).

If claude is not on your PATH, the tool falls back to ~/.claude/local/claude.

launch codex (not supported)

Codex is not supported by glm-launch. Current codex only speaks the OpenAI Responses API (it removed wire_api = "chat"), but Z.AI's GLM endpoints are Anthropic Messages and OpenAI Chat Completions only — there is no /responses endpoint, so codex requests return 404. The codex command is intentionally disabled and exits with this explanation.

Use launch claude instead — it uses Z.AI's Anthropic-compatible endpoint. If Z.AI later ships a Responses-compatible endpoint, codex support can be revisited.

shell

Print export lines that bootstrap your current shell with the GLM env vars — without launching anything. Eval the output and a plain claude (or any Anthropic SDK tool) will talk to Z.AI.

eval "$(uv run glm-launch shell)"
claude

Accepts the same model/auth options as launch claude (--model, --auth-token, --default-*-model, etc.). Secrets are shell-quoted; empty values are skipped. Sets ANTHROPIC_MODEL plus all the ANTHROPIC_* / CLAUDE_CODE_* vars listed under launch claude.

# Inspect what would be exported
uv run glm-launch shell

# Bootstrap with a specific model
eval "$(uv run glm-launch shell --model glm-5.1)"

models

List Z.AI GLM models. By default prints a built-in, annotated list; --remote fetches the live list from the Z.AI PaaS endpoint.

# Built-in list (no token needed)
uv run glm-launch models

# Live list from the API (needs GLM_AUTH_TOKEN)
uv run glm-launch models --remote

Options:

Flag Env var Default Description
--remote / -r false Fetch the live list from the Z.AI API
--models-url GLM_MODELS_URL https://api.z.ai/api/coding/paas/v4/models PaaS models endpoint (used with --remote)
--auth-token GLM_AUTH_TOKEN Auth token (required with --remote)
--timeout 30.0 Request timeout in seconds (must be greater than zero)

The live endpoint is the OpenAI-compatible coding PaaS base (/api/coding/paas/v4/models) and uses Authorization: Bearer <token> — distinct from the Anthropic-style chat base (/api/anthropic) used by launch claude and bench. Coding Plan keys only work through the coding endpoints; if you have a general Z.AI API key instead, point --models-url at https://api.z.ai/api/paas/v4/models.

bench

Time a single /v1/messages round-trip against the configured GLM endpoint. Useful as a sanity check that your auth token, base URL, and chosen model are reachable.

uv run glm-launch bench

Options:

Flag Env var Default Description
--model / -m glm-5.2 Model to benchmark
--base-url GLM_BASE_URL https://api.z.ai/api/anthropic API endpoint
--auth-token GLM_AUTH_TOKEN (required) Auth token for the endpoint
--timeout 30.0 Request timeout in seconds (must be greater than zero)

Sends a minimal 32-token request and prints the round-trip time. Exits non-zero on HTTP error or timeout.

Example output:

  glm-5.2 via https://api.z.ai/api/anthropic
  OK (200) in 412ms

usage

Open the Z.AI usage/quota dashboard in your browser. Coding Plan quotas are tracked in 5-hour and weekly windows, and there is no API for quota data — the dashboard is the only place to see it.

uv run glm-launch usage

doctor

Check your environment for correct setup. Reports on environment variables and binary availability.

uv run glm-launch doctor

Checks performed:

  • Authentication — Whether the required GLM_AUTH_TOKEN is set. The token is masked in output.
  • Environment variables — Whether the optional GLM, Anthropic default-model, and Claude Code env vars used by the launch commands are set.
  • Binaries — Whether claude is found on PATH (with fallback to ~/.claude/local/claude), including its version — the default glm-5.2[1m] model needs a recent Claude Code, so if claude reports the [1m] model doesn't exist, upgrade.

Exits with code 1 if GLM_AUTH_TOKEN is not set or the claude binary is missing, 0 otherwise.

Example output:

Environment variables:
  GLM_BASE_URL: (not set)
  GLM_API_KEY: (not set)
  GLM_AUTH_TOKEN: zai_***
  GLM_MODELS_URL: (not set)
  API_TIMEOUT_MS: (not set)
  ANTHROPIC_DEFAULT_HAIKU_MODEL: (not set)
  ANTHROPIC_DEFAULT_SONNET_MODEL: (not set)
  ANTHROPIC_DEFAULT_OPUS_MODEL: (not set)
  ANTHROPIC_DEFAULT_FABLE_MODEL: (not set)
  CLAUDE_CODE_SUBAGENT_MODEL: (not set)
  CLAUDE_CODE_EFFORT_LEVEL: (not set)
  CLAUDE_CODE_ATTRIBUTION_HEADER: (not set)
  CLAUDE_CODE_AUTO_COMPACT_WINDOW: (not set)
  CLAUDE_CODE_MAX_CONTEXT_TOKENS: (not set)

Binaries:
  claude: /usr/local/bin/claude

All checks passed.

Environment variables

Variable Used by Description
GLM_BASE_URL launch claude, shell API base URL
GLM_API_KEY launch claude, shell API key
GLM_AUTH_TOKEN launch claude, shell, bench, models --remote Z.AI auth token (required)
GLM_MODELS_URL models --remote PaaS models endpoint
API_TIMEOUT_MS launch claude, shell Request timeout in milliseconds (positive integer)
ANTHROPIC_DEFAULT_HAIKU_MODEL launch claude, shell Model for Haiku-tier requests
ANTHROPIC_DEFAULT_SONNET_MODEL launch claude, shell Model for Sonnet-tier requests
ANTHROPIC_DEFAULT_OPUS_MODEL launch claude, shell Model for Opus-tier requests
ANTHROPIC_DEFAULT_FABLE_MODEL launch claude, shell Model for Fable-tier requests
CLAUDE_CODE_SUBAGENT_MODEL launch claude, shell Model used for spawned subagents
CLAUDE_CODE_EFFORT_LEVEL launch claude, shell Validated effort level for the agent loop
CLAUDE_CODE_ATTRIBUTION_HEADER launch claude, shell Attribution header toggle (0 or 1)
CLAUDE_CODE_AUTO_COMPACT_WINDOW launch claude, shell auto, empty, or a positive token count
CLAUDE_CODE_MAX_CONTEXT_TOKENS launch claude, shell auto, empty, or a positive token count

How it works

launch claude follows three steps:

  1. Resolve the claude binary on PATH (falling back to ~/.claude/local/claude)
  2. Set up the GLM environment variables
  3. os.execvpe() the binary — fully replacing the glm process with claude for direct stdio passthrough

Z.AI exposes an Anthropic-compatible endpoint at https://api.z.ai/api/anthropic, so no local proxy is needed. The CLI sets the standard ANTHROPIC_* env vars and Claude Code talks directly to Z.AI.

Development

Common tasks are wrapped in a justfile. Run just with no arguments to list them.

Before committing, run the same core checks used by CI:

uv run pytest
uv tool run prek run --all-files
uv build
Recipe Description
just bootstrap Upgrade pip/uv, then uv sync
just sync uv sync the project dependencies
just lock uv lock the dependency versions
just build uv build the wheel and sdist
just bump *ARGS Bump the CalVer version with bumpver (e.g. just bump)
just bump-dry *ARGS Preview a version bump without writing changes
just release *ARGS Bump, relock, and push the tag — CI then publishes to PyPI
just lint *ARGS Run the prek hooks (defaults to --all-files)
just fmt Format the justfile itself
just demo Smoke-test the CLI by listing models

Versioning follows CalVer (YYYY.MM.INC1), and lint hooks (ruff, pyupgrade, validate-pyproject) are configured in .pre-commit-config.yaml and run with prek. CI runs the tests, lint checks, and package build; the release workflow reruns the tests before publishing.

Releases are automated. Run just release to bump the CalVer version, relock, and push the tag in one step. Pushing a YYYY.MM.INC1 tag triggers the GitHub Actions release workflow, which builds and publishes to PyPI via trusted publishing (OIDC, no API token). A plain git push never publishes — only the tag does.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

glm_launch-2026.7.6.tar.gz (20.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

glm_launch-2026.7.6-py3-none-any.whl (14.3 kB view details)

Uploaded Python 3

File details

Details for the file glm_launch-2026.7.6.tar.gz.

File metadata

  • Download URL: glm_launch-2026.7.6.tar.gz
  • Upload date:
  • Size: 20.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for glm_launch-2026.7.6.tar.gz
Algorithm Hash digest
SHA256 ce7f99a1c56ca05e41f5bde5d91c222723395113ae8afd40d4b68960dbdec3cf
MD5 3a8fd3b119f99302dc058af9d3103cc4
BLAKE2b-256 3993eab1e25478c101d71fac50f78a0a58e2b0dc62bf13dc86f2a80aebf483a8

See more details on using hashes here.

Provenance

The following attestation bundles were made for glm_launch-2026.7.6.tar.gz:

Publisher: release.yml on jefftriplett/glm-launch

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file glm_launch-2026.7.6-py3-none-any.whl.

File metadata

  • Download URL: glm_launch-2026.7.6-py3-none-any.whl
  • Upload date:
  • Size: 14.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for glm_launch-2026.7.6-py3-none-any.whl
Algorithm Hash digest
SHA256 250d394d83c5a7cf3c604517bb466aa87f7b4c8d4cf55b94d113b392c78d469d
MD5 33a9eaeae85fc9b4f968dff11a8f40e3
BLAKE2b-256 f08e088f025f2fb175f2e8e53ed69cccd94dfde599ad92139b454a32143036e1

See more details on using hashes here.

Provenance

The following attestation bundles were made for glm_launch-2026.7.6-py3-none-any.whl:

Publisher: release.yml on jefftriplett/glm-launch

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page