Skip to main content

jobgpumonitor-server

The consumer side of jobgpumonitor: reads the JSONL events written by the emitter and the scheduler probe, keeps one document per run, sends notifications (ntfy, Telegram, webhook) and serves a small read-only API.

jobgpumonitor (in the job) ──┐
                             ├──> $JGM_DIR/runs/**.jsonl ──> jgmd serve ──> ntfy / Telegram / webhook
jgm scheduler (login node) ──┘                                   └──> SQLite ──> HTTP API

Quick start (login node)

pip install "jobgpumonitor-server[api]"
jgmd init                         # writes ~/.config/jgm-server/config.toml
$EDITOR ~/.config/jgm-server/config.toml   # put an ntfy topic or a Telegram bot in [notify.*]
jgmd notify-test                  # phone should buzz
jgmd serve --api                  # keep it in tmux / systemd --user

Then, still on the login node, jgm scheduler from the emitter package so that OOM, time-outs, preemptions and the queue are reported too.

What you get notified about

Alert When
🚀 started the job leaves the queue (with node, GPUs, time queued, time limit)
✅ finished clean end: duration, last metrics, GPU utilisation and idle share, peak memory
❌ failed / OOM / timed out / cancelled / preempted with the traceback or the scheduler's reason, exit code, last stderr lines, path of the .out file
⚠️ correction the scheduler's verdict contradicts what the process reported
⏱ will not finish in time tqdm ETA overshoots the job deadline by more than 2 minutes
🧠 memory above 90 % of the cgroup / requested memory
🥱 GPU idle every visible GPU under 5 % for 15 minutes
💀 stopped reporting no heartbeat for 3 intervals while the scheduler still says RUNNING (15 min when nothing at all arrives from the cluster: the agent or the proxy is then the likelier culprit)
🔄 reporting again heartbeats are back after a 💀; the next silence is reported again
⚠️ correction the scheduler's real verdict arrives after an UNKNOWN_ENDED (accounting was down)

Every alert fires once per run and is recorded, so restarting the server never re-sends. A channel that fails (ntfy down, no network) gets the alert again later, with backoff, for up to 6 hours; the other channels are not sent it twice. Raw events are kept keep_events_days after they are received (not after they were emitted).

API

jgmd serve --api (needs the [api] extra) exposes on 127.0.0.1:21834:

GET /health
GET /runs?phase=running
GET /runs/<cluster>/<job>/<restart>
GET /runs/<cluster>/<job>/<restart>/events?after=0&types=metric.log,progress.update
GET /runs/<cluster>/<job>/<restart>/stream          # server-sent events
GET /runs/<cluster>/<job>/<restart>/logs?stream=stdout   # the job's .out/.err, live
GET /alerts

Interactive docs at /docs. From your laptop: ssh -L 21834:localhost:21834 cluster.

CLI

jgmd serve [--once] [--api]    ingest, rules, notifications
jgmd runs [--phase running]    table of runs
jgmd show <run_id | job id>    full run document (--events N for raw events)
jgmd alerts                    what was sent, and through which channel
jgmd notify-test               send a test message
jgmd init                      write the example config

Configuration lives in ~/.config/jgm-server/config.toml (see jgmd init); JGMD_DIRS, JGMD_NTFY_TOPIC, JGMD_TELEGRAM_TOKEN / JGMD_TELEGRAM_CHAT_ID, JGMD_WEBHOOK_URL work without a file.

Development

uv venv && uv pip install -e ".[dev]" && uv run pytest

Metadata

Release files for jobgpumonitor-server 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for jobgpumonitor-server 0.4.0
File Size Uploaded
jobgpumonitor_server-0.4.0.tar.gz 62.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for jobgpumonitor-server 0.4.0
File Interpreter ABI Platform
jobgpumonitor_server-0.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 93.3 kB

Release files / jobgpumonitor_server-0.4.0.tar.gz

Download URL jobgpumonitor_server-0.4.0.tar.gz
Size 62.7 kB
Tags Source
SHA-256 checksum
How to use checksums
ca17c31a44dd71e72430869d58e2bab6cf650199e7c91ccff13284420ade051c
BLAKE2b-256 checksum
How to use checksums
fc6b23313c1c39a2da2b734bc3f13cbd91d8bd270568d5f5eb4e6bde2fe2aaf3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.

Transparency log

Release files / jobgpumonitor_server-0.4.0-py3-none-any.whl

Download URL jobgpumonitor_server-0.4.0-py3-none-any.whl
Size 30.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
8b74875533780a041ffb3f21ba0458e8d8900ba96cfc60a7d3ec9c6e8e908c9c
BLAKE2b-256 checksum
How to use checksums
22ca14b691f6a82b584a6b5f90703fef755110be53ed75b904bd206bf9360f70
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.4.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page