jobgpumonitor-server
The consumer side of jobgpumonitor: reads the JSONL events written by the emitter and the scheduler probe, keeps one document per run, sends notifications (ntfy, Telegram, webhook) and serves a small read-only API.
jobgpumonitor (in the job) ──┐
├──> $JGM_DIR/runs/**.jsonl ──> jgmd serve ──> ntfy / Telegram / webhook
jgm scheduler (login node) ──┘ └──> SQLite ──> HTTP API
Quick start (login node)
pip install "jobgpumonitor-server[api]"
jgmd init # writes ~/.config/jgm-server/config.toml
$EDITOR ~/.config/jgm-server/config.toml # put an ntfy topic or a Telegram bot in [notify.*]
jgmd notify-test # phone should buzz
jgmd serve --api # keep it in tmux / systemd --user
Then, still on the login node, jgm scheduler from the emitter package so that OOM,
time-outs, preemptions and the queue are reported too.
What you get notified about
| Alert | When |
|---|---|
| 🚀 started | the job leaves the queue (with node, GPUs, time queued, time limit) |
| ✅ finished | clean end: duration, last metrics, GPU utilisation and idle share, peak memory |
| ❌ failed / OOM / timed out / cancelled / preempted | with the traceback or the scheduler's reason, exit code, last stderr lines, path of the .out file |
| ⚠️ correction | the scheduler's verdict contradicts what the process reported |
| ⏱ will not finish in time | tqdm ETA overshoots the job deadline by more than 2 minutes |
| 🧠 memory | above 90 % of the cgroup / requested memory |
| 🥱 GPU idle | every visible GPU under 5 % for 15 minutes |
| 💀 stopped reporting | no heartbeat for 3 intervals while the scheduler still says RUNNING (15 min when nothing at all arrives from the cluster: the agent or the proxy is then the likelier culprit) |
| 🔄 reporting again | heartbeats are back after a 💀; the next silence is reported again |
| ⚠️ correction | the scheduler's real verdict arrives after an UNKNOWN_ENDED (accounting was down) |
Every alert fires once per run and is recorded, so restarting the server never re-sends.
A channel that fails (ntfy down, no network) gets the alert again later, with backoff, for
up to 6 hours; the other channels are not sent it twice. Raw events are kept
keep_events_days after they are received (not after they were emitted).
API
jgmd serve --api (needs the [api] extra) exposes on 127.0.0.1:21834:
GET /health
GET /runs?phase=running
GET /runs/<cluster>/<job>/<restart>
GET /runs/<cluster>/<job>/<restart>/events?after=0&types=metric.log,progress.update
GET /runs/<cluster>/<job>/<restart>/stream # server-sent events
GET /runs/<cluster>/<job>/<restart>/logs?stream=stdout # the job's .out/.err, live
GET /alerts
Interactive docs at /docs. From your laptop: ssh -L 21834:localhost:21834 cluster.
CLI
jgmd serve [--once] [--api] ingest, rules, notifications
jgmd runs [--phase running] table of runs
jgmd show <run_id | job id> full run document (--events N for raw events)
jgmd alerts what was sent, and through which channel
jgmd notify-test send a test message
jgmd init write the example config
Configuration lives in ~/.config/jgm-server/config.toml (see jgmd init); JGMD_DIRS,
JGMD_NTFY_TOPIC, JGMD_TELEGRAM_TOKEN / JGMD_TELEGRAM_CHAT_ID, JGMD_WEBHOOK_URL work
without a file.
Development
uv venv && uv pip install -e ".[dev]" && uv run pytest
Metadata
Release files for jobgpumonitor-server 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| jobgpumonitor_server-0.4.0.tar.gz | 62.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| jobgpumonitor_server-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 93.3 kB
Release files / jobgpumonitor_server-0.4.0.tar.gz
| Download URL | jobgpumonitor_server-0.4.0.tar.gz |
|---|---|
| Size | 62.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ca17c31a44dd71e72430869d58e2bab6cf650199e7c91ccff13284420ade051c
|
|
BLAKE2b-256 checksum How to use checksums |
fc6b23313c1c39a2da2b734bc3f13cbd91d8bd270568d5f5eb4e6bde2fe2aaf3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.
Transparency logRelease files / jobgpumonitor_server-0.4.0-py3-none-any.whl
| Download URL | jobgpumonitor_server-0.4.0-py3-none-any.whl |
|---|---|
| Size | 30.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8b74875533780a041ffb3f21ba0458e8d8900ba96cfc60a7d3ec9c6e8e908c9c
|
|
BLAKE2b-256 checksum How to use checksums |
22ca14b691f6a82b584a6b5f90703fef755110be53ed75b904bd206bf9360f70
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 7, 2026.
Transparency log