Skip to main content

check-llm-quota

Check usage quota, rate limits, remaining sessions and supported models across LLM API providers from one command. Stdlib only, no dependencies.

Supported providers:

Provider Subcommand What it reports
Z.ai (Zhipu AI / 智谱AI) GLM Coding Plan zai Plan level, per-limit usage % (time, tokens, sessions, requests), reset times, model availability per plan
Kimi (Moonshot AI) Coding API kimi Plan, weekly quota, rolling-window rate limit, max parallel requests, model list

This package merges and replaces the standalone zai-quota and kimi-quota tools. Both old command names still work as aliases.

Quick start

# Run without installing (uv: https://docs.astral.sh/uv/)
uvx check-llm-quota all

# Or install it
uv tool install check-llm-quota     # or: pip install check-llm-quota
check-llm-quota all
=== kimi ===
🔑 Kimi Coding API — Quota Report
   Plan: Paid (BASIC)

📋 Weekly Quota:
   Used: 20/100 (20%)
   Remaining: 80
   Resets: 2026-06-21 08:14 SGT

📊 Rate Limit (300minute rolling window):
   Used: 1/100 (1%)
   Remaining: 99
   Resets: 2026-06-17 09:14 SGT

⚡ Max Parallel Requests: 10

=== zai ===
  Z.ai GLM Quota
  Plan: Lite
  -------------------------------------
  [green] Time Limit
     Used: 0% | Remaining: 100
       - search-prime: 0
       - web-reader: 0
       - zread: 0
     Resets: 2026-04-16 10:31 SGT

  [green] Tokens
     Used: 18%
     Resets: in 3h

  -------------------------------------

Usage

check-llm-quota <provider|all> [--key KEY] [--json] [provider flags]

check-llm-quota zai                   # Z.ai quota report
check-llm-quota zai --models          # GLM models and which plan unlocks them
check-llm-quota zai --endpoint cn     # force open.bigmodel.cn (default: api.z.ai, then cn)
check-llm-quota kimi                  # Kimi quota report
check-llm-quota kimi --models         # Kimi model list
check-llm-quota kimi --base-url URL   # custom Kimi-compatible endpoint
check-llm-quota all                   # every provider with a key configured
check-llm-quota all --json            # {"kimi": {...}, "zai": {...}}
check-llm-quota --json zai            # --key/--json work before or after the provider

python3 -m check_llm_quota all        # module form, same CLI
zai-quota --models                    # alias for: check-llm-quota zai --models
kimi-quota --json                     # alias for: check-llm-quota kimi --json

Exit codes: 0 success, 1 missing key or API failure (with all, 1 if any provider failed), 2 invalid arguments.

API keys per provider

Resolution order is always --key > environment variables > provider fallback.

zai

Source Notes
--key
ZAI_API_KEY from z.ai or open.bigmodel.cn
~/.hermes/auth.json Hermes Agent credential pool: first entry of credential_pool.zai[] with a non-empty access_token (a legacy top-level zai: [...] list is also read)

The quota endpoint is an unofficial monitoring API (/api/monitor/usage/quota/limit). It works today but Z.ai could change it without notice. api.z.ai is tried first and open.bigmodel.cn second; pass --endpoint intl|cn to pin one.

kimi

Source Notes
--key
MOONSHOT_API_KEY checked first
KIMI_API_KEY checked second
KIMI_BASE_URL / --base-url override https://api.kimi.com/coding/v1

all

all only uses environment variables and the Hermes file. Providers without a key are skipped with a note on stderr, and one provider failing does not stop the others. --key is rejected because it is ambiguous across providers.

Cron example

Log a JSON snapshot every hour and warn in the system log when any provider fails:

# m h dom mon dow   command
0 * * * *  ZAI_API_KEY=... MOONSHOT_API_KEY=... /usr/local/bin/uvx check-llm-quota all --json >> "$HOME/llm-quota.log" 2>&1 || logger -t check-llm-quota "quota check failed"

Set the keys in the crontab (as above), in a sourced env file, or rely on ~/.hermes/auth.json for Z.ai. Use the absolute path to uvx (which uvx) since cron has a minimal PATH. For a plain-text daily report, drop --json and pipe to mail or your notifier of choice.

Agent skill

The repository is also an Agent Skills bundle. Install into Claude Code, Codex, Cursor, OpenCode, Gemini CLI, Hermes and 45+ other agents:

npx skills add SeeYangZhi/check-llm-quota

Three skills are provided:

  • check-llm-quota: the full multi-provider skill. Its scripts/check_quota.py is a thin dispatcher that runs uvx check-llm-quota <args> and falls back to python3 -m check_llm_quota <args>.
  • zai-quota and kimi-quota: short pointer skills so prompts such as "check my GLM quota" or "how much Kimi usage is left" still trigger.

Agents run, for example:

python3 skills/check-llm-quota/scripts/check_quota.py all

Adding a provider

Each provider is one file in src/check_llm_quota/providers/. The registry imports every module there and picks up any module-level PROVIDER that is a check_llm_quota.core.Provider instance, so there is nothing else to wire up.

  1. Create src/check_llm_quota/providers/<name>.py:

    import argparse
    from typing import Any
    
    from .. import core
    from ..core import Provider, fmt_sgt, pct, pct_status
    
    
    class ExampleProvider(Provider):
        name = "example"                      # subcommand name
        description = "Example AI usage and limits"
        env_vars = ("EXAMPLE_API_KEY",)       # checked in order
        hermes_pool = None                    # or a ~/.hermes/auth.json pool name
        missing_key_message = "No Example key. Set EXAMPLE_API_KEY or pass --key."
    
        def add_arguments(self, parser: argparse.ArgumentParser) -> None:
            # --key and --json are added for you
            parser.add_argument("--models", action="store_true", help="List models")
    
        def fetch(self, key: str, args: argparse.Namespace) -> Any:
            # Raise core.QuotaError for user-facing failures; never sys.exit.
            return core.fetch_json("https://api.example.com/v1/usage", key, label="Example")
    
        def render(self, payload: Any, args: argparse.Namespace) -> None:
            used, total = payload["used"], payload["limit"]
            print(f"[{pct_status(pct(used, total))}] Used {used}/{total} ({pct(used, total)}%)")
            print(f"Resets: {fmt_sgt(payload['resetTime'])}")
    
        # Optional: shape what --json prints (defaults to the raw payload).
        # def to_json(self, payload, args): ...
    
    
    PROVIDER = ExampleProvider()
    
  2. Helpers in core you can reuse: fetch_json (Bearer auth, multi-URL fallback on 404/connection errors, 401 fails fast), resolve_key, hermes_pool_key, format_reset (ms epoch to relative/SGT), fmt_sgt (ISO UTC to SGT), pct, pct_status.

  3. Add fixture JSON under tests/fixtures/ and tests next to the existing tests/test_zai.py / tests/test_kimi.py. Tests are offline: conftest.py blocks urllib and provides fake_fetch to map URL substrings to fixtures.

  4. Document the env vars in this README and, if useful, add a pointer skill in skills/<name>-quota/SKILL.md.

Development

uv run pytest        # offline test suite  (make test)
make check           # --help smoke test for every entry point
uv build             # wheel + sdist       (make build)

Requires Python 3.9+. CI runs the tests on 3.9, 3.12 and 3.13.

License

MIT

Metadata

Release files for check-llm-quota 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for check-llm-quota 0.1.0
File Size Uploaded
check_llm_quota-0.1.0.tar.gz 23.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for check-llm-quota 0.1.0
File Interpreter ABI Platform
check_llm_quota-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 41.3 kB

Release files / check_llm_quota-0.1.0.tar.gz

Download URL check_llm_quota-0.1.0.tar.gz
Size 23.6 kB
Tags Source
SHA-256 checksum
How to use checksums
3d983daa830c1f1db6618d37f0f5c4e41a1a1c751235e00b655b7717ce4ccccd
BLAKE2b-256 checksum
How to use checksums
df892a169e63e7887bdc1b310d97a431cec4ab1613835804109d5f12044a11fc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release files / check_llm_quota-0.1.0-py3-none-any.whl

Download URL check_llm_quota-0.1.0-py3-none-any.whl
Size 17.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8e509541b4af2098e09f9bac1aa0439bbc603f3970e13b28767c5f12d9fcb690
BLAKE2b-256 checksum
How to use checksums
ed6476a5b6513f1e6309c7016207b4e40bee191edd506c1237b0d48083328f5a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page