django-q-watchdog
Find the django-q tasks that crashed, timed out or froze, and say which ones.
When a django-q worker process dies (out-of-memory kill, segfault, container restart), the
task it was running disappears: no row in Failed tasks, no error, no traceback. django-q only
logs reincarnated worker Process-1:3 after death. On the original django-q (1.3.x), a task
killed by the timeout disappears the same way. And a task that hangs looks exactly like one
that is working until the timeout finally kills it.
django-q-watchdog records every task as it starts, keeps a heartbeat while it runs, and tells you what happened to the ones that never finished:
| Status | Meaning |
|---|---|
running |
Heartbeat is fresh. |
frozen |
Heartbeat is fresh, but memory hasn't moved for FROZEN_AFTER_SECONDS. Reported as waiting (almost no CPU: usually a network call or lock with no timeout) or busy (high CPU: possibly a loop). |
timed_out |
The worker was killed by django-q's timeout. |
lost |
The worker process died while running the task. |
Lost and timed-out tasks are saved into django-q's own Failed tasks (admin and
Failure model), with the worker, host and reason, so they show up where you already look.
What django-q itself records
| Situation | django-q 1.3.x | django-q2 | with django-q-watchdog |
|---|---|---|---|
| Task raises an exception | Failed task | Failed task | unchanged |
| Task exceeds the timeout | lost | Failed task (raised inside the task) | timed out, saved as failed (1.3.x) |
| Worker process dies mid-task | lost | lost | lost, named, saved as failed |
| Task hangs | silent until the timeout | silent until the timeout | frozen: waiting or busy |
Install
pip install django-q-watchdog
INSTALLED_APPS = [
# ...
"django_q",
"django_q_watchdog",
]
That's all the workers need: the heartbeat starts with each task. Then make sure problems get reported without anyone looking. Run the check every few minutes from cron or your scheduler:
python manage.py qwatchdog --fail-on-problems
It prints a table, logs one error per lost or frozen task to the django_q_watchdog logger,
and exits non-zero when something needs attention. Use --json for machine-readable output.
Run it from cron or your platform's scheduler rather than as a django-q schedule: if the cluster itself is stuck, a django-q schedule won't run either.
Alerts
Every problem is reported once, through the logger and a Django signal you can connect to Slack, Sentry, PagerDuty and so on:
from django.dispatch import receiver
from django_q_watchdog.signals import task_frozen, task_lost
@receiver(task_lost)
def notify_lost(sender, report, **kwargs):
slack.post(f"Task {report['name']} ({report['func']}) {report['status']}: {report['reason']}")
report contains the task ID, name, function, group, tenant, host, process, PID, start time,
running time, memory and CPU.
Status endpoint (optional)
urlpatterns = [
path("ops/q-watchdog/", include("django_q_watchdog.urls")),
]
Returns the in-flight tasks and a summary (running, frozen, timed out, lost, queued) as JSON.
It's staff-only by default because it names tasks, hosts and tenants. Change who can see it
with ACCESS_CHECK.
Example response, GET /ops/q-watchdog/:
{
"summary": {
"running": 1,
"frozen": 1,
"timed_out": 0,
"lost": 1,
"queued": 0
},
"tasks": [
{
"task_id": "1a2b3c4d5e6f47a8b9c0d1e2f3a4b5c6",
"name": "kilo-tango-river-seven",
"func": "integrations.tasks.sync_employees",
"group": null,
"tenant": "globex",
"host": "worker-1",
"process": "Process-1:4",
"pid": 4187,
"timeout": 600,
"started": "2026-10-09T03:43:59.735551+00:00",
"last_seen": "2026-10-09T05:13:29.735551+00:00",
"rss_mb": 150.2,
"rss_start_mb": 150.0,
"memory_changed_at": "2026-10-09T03:44:29.735551+00:00",
"cpu_seconds": 1.2,
"cpu_percent": 0,
"cpu_recent_percent": 0,
"status": "frozen",
"reason": "waiting: almost no CPU, likely blocked on I/O or a lock",
"running_for_seconds": 5400
},
{
"task_id": "7c6b5a4f3e2d41c0b9a8f7e6d5c4b3a2",
"name": "ruby-hotel-falcon-two",
"func": "documents.tasks.generate_pdf",
"group": null,
"tenant": "acme",
"host": "worker-1",
"process": "Process-1:2",
"pid": 4179,
"timeout": 600,
"started": "2026-10-09T05:03:19.735551+00:00",
"last_seen": "2026-10-09T05:07:09.735551+00:00",
"rss_mb": 1985.7,
"rss_start_mb": 190.3,
"memory_changed_at": "2026-10-09T05:07:09.735551+00:00",
"cpu_seconds": 120.5,
"cpu_percent": 52,
"cpu_recent_percent": 98,
"status": "lost",
"reason": "the worker process died while running it",
"running_for_seconds": 640
},
{
"task_id": "9f8b2c1e4d7a4b6c8e0f1a2b3c4d5e6f",
"name": "oscar-delta-nine-lemon",
"func": "reports.tasks.export_payroll",
"group": null,
"tenant": "acme",
"host": "worker-1",
"process": "Process-1:3",
"pid": 4182,
"timeout": 600,
"started": "2026-10-09T05:12:24.735551+00:00",
"last_seen": "2026-10-09T05:13:47.735551+00:00",
"rss_mb": 212.4,
"rss_start_mb": 180.1,
"memory_changed_at": "2026-10-09T05:13:47.735551+00:00",
"cpu_seconds": 41.0,
"cpu_percent": 43,
"cpu_recent_percent": 51,
"status": "running",
"reason": null,
"running_for_seconds": 95
}
]
}
Reading it:
sync_employeesis frozen and waiting. It has run for 90 minutes on almost no CPU, and its memory hasn't moved. That usually means a call to another system with no timeout.generate_pdfwas lost. Its memory grew from 190 MB to almost 2 GB before the worker died, which points to an out-of-memory kill. It is now also in django-q's Failed tasks.export_payrollis running normally.last_seenis the last heartbeat.rss_start_mbandrss_mbare the worker's memory when the task started and at the last heartbeat, andmemory_changed_atis when it last moved by 1 MB or more.cpu_recent_percentis CPU use since the previous heartbeat, which says what the task is doing now.cpu_percentis the average since it started.tenantisnullin projects without django-tenants.
The same data is available from python manage.py qwatchdog --json.
Multi-tenant projects
With django-tenants and
django-tenants-q, each record carries the task's schema_name, and lost tasks are saved as
failures in that tenant's schema. Without django-tenants, nothing changes.
Configuration
All optional:
Q_WATCHDOG = {
"CACHE": "default", # must be django-redis, Django's RedisCache, or LocMemCache
"HEARTBEAT_SECONDS": 60,
"STALE_AFTER_SECONDS": 180, # no heartbeat for this long: the worker is gone
"FROZEN_AFTER_SECONDS": 3600, # None turns frozen detection off
"RECORD_TTL_SECONDS": 604800,
"SAVE_LOST_AS_FAILED": True, # put lost and timed-out tasks into django-q's Failed tasks
"TENANT_KWARG": "schema_name",
"ACCESS_CHECK": "django_q_watchdog.conf.staff_only",
}
The cache must be shared by your web processes and your cluster, which is why a Redis cache is
expected in production. Memory figures come from /proc and are Linux-only; everything else
works on any platform.
How it works
- In each worker: django-q's
pre_executesignal writes a record for the task, and a small thread refreshes it everyHEARTBEAT_SECONDSwith memory and CPU usage. A refresh only updates a record that still exists, so a finished task is never brought back. - When django-q saves the task (success or failure), the record is deleted.
- The check looks at what's left: a record whose heartbeat stopped belongs to a worker that died or was killed; a record still refreshing but whose memory hasn't moved for an hour belongs to a frozen task. Whether it's waiting or busy is decided by its CPU use since the previous heartbeat, not its average, so a task that worked hard and then got stuck is reported as waiting.
Tested end to end against a real cluster, by killing worker processes and letting tasks hang past the timeout, on django-q 1.3.9 with Django 4.2 and on django-q2 with Django 5.2.
Development
pip install -e ".[dev]"
pytest
License
MIT
Metadata
Release files for django-q-watchdog 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| django_q_watchdog-0.1.1.tar.gz | 20.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| django_q_watchdog-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 36.6 kB
Release files / django_q_watchdog-0.1.1.tar.gz
| Download URL | django_q_watchdog-0.1.1.tar.gz |
|---|---|
| Size | 20.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5e73daf22836956d568b1dd3d6231765b3e808640a74a6af115e1a43812368b9
|
|
BLAKE2b-256 checksum How to use checksums |
3f15d2584f03b0ee983e575a8256fb7d308a8bfec9764994c3f187263cd71466
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Release files / django_q_watchdog-0.1.1-py3-none-any.whl
| Download URL | django_q_watchdog-0.1.1-py3-none-any.whl |
|---|---|
| Size | 15.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
51cdb087cabc912e8b573a1c0186b19dfaa6f20b6896f87dbf6139231e3e1ed7
|
|
BLAKE2b-256 checksum How to use checksums |
44bfd2bd03eb1f477454e0b0b10156fbc70b720a2f1c670124555ee0ca4fc358
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|