Skip to main content

django-q-watchdog

Find the django-q tasks that crashed, timed out or froze, and say which ones.

When a django-q worker process dies (out-of-memory kill, segfault, container restart), the task it was running disappears: no row in Failed tasks, no error, no traceback. django-q only logs reincarnated worker Process-1:3 after death. On the original django-q (1.3.x), a task killed by the timeout disappears the same way. And a task that hangs looks exactly like one that is working until the timeout finally kills it.

django-q-watchdog records every task as it starts, keeps a heartbeat while it runs, and tells you what happened to the ones that never finished:

Status Meaning
running Heartbeat is fresh.
frozen Heartbeat is fresh, but memory hasn't moved for FROZEN_AFTER_SECONDS. Reported as waiting (almost no CPU: usually a network call or lock with no timeout) or busy (high CPU: possibly a loop).
timed_out The worker was killed by django-q's timeout.
lost The worker process died while running the task.

Lost and timed-out tasks are saved into django-q's own Failed tasks (admin and Failure model), with the worker, host and reason, so they show up where you already look.

What django-q itself records

Situation django-q 1.3.x django-q2 with django-q-watchdog
Task raises an exception Failed task Failed task unchanged
Task exceeds the timeout lost Failed task (raised inside the task) timed out, saved as failed (1.3.x)
Worker process dies mid-task lost lost lost, named, saved as failed
Task hangs silent until the timeout silent until the timeout frozen: waiting or busy

Install

pip install django-q-watchdog
INSTALLED_APPS = [
    # ...
    "django_q",
    "django_q_watchdog",
]

That's all the workers need: the heartbeat starts with each task. Then make sure problems get reported without anyone looking. Run the check every few minutes from cron or your scheduler:

python manage.py qwatchdog --fail-on-problems

It prints a table, logs one error per lost or frozen task to the django_q_watchdog logger, and exits non-zero when something needs attention. Use --json for machine-readable output.

Run it from cron or your platform's scheduler rather than as a django-q schedule: if the cluster itself is stuck, a django-q schedule won't run either.

Alerts

Every problem is reported once, through the logger and a Django signal you can connect to Slack, Sentry, PagerDuty and so on:

from django.dispatch import receiver
from django_q_watchdog.signals import task_frozen, task_lost


@receiver(task_lost)
def notify_lost(sender, report, **kwargs):
    slack.post(f"Task {report['name']} ({report['func']}) {report['status']}: {report['reason']}")

report contains the task ID, name, function, group, tenant, host, process, PID, start time, running time, memory and CPU.

Status endpoint (optional)

urlpatterns = [
    path("ops/q-watchdog/", include("django_q_watchdog.urls")),
]

Returns the in-flight tasks and a summary (running, frozen, timed out, lost, queued) as JSON. It's staff-only by default because it names tasks, hosts and tenants. Change who can see it with ACCESS_CHECK.

Example response, GET /ops/q-watchdog/:

{
  "summary": {
    "running": 1,
    "frozen": 1,
    "timed_out": 0,
    "lost": 1,
    "queued": 0
  },
  "tasks": [
    {
      "task_id": "1a2b3c4d5e6f47a8b9c0d1e2f3a4b5c6",
      "name": "kilo-tango-river-seven",
      "func": "integrations.tasks.sync_employees",
      "group": null,
      "tenant": "globex",
      "host": "worker-1",
      "process": "Process-1:4",
      "pid": 4187,
      "timeout": 600,
      "started": "2026-10-09T03:43:59.735551+00:00",
      "last_seen": "2026-10-09T05:13:29.735551+00:00",
      "rss_mb": 150.2,
      "rss_start_mb": 150.0,
      "memory_changed_at": "2026-10-09T03:44:29.735551+00:00",
      "cpu_seconds": 1.2,
      "cpu_percent": 0,
      "cpu_recent_percent": 0,
      "status": "frozen",
      "reason": "waiting: almost no CPU, likely blocked on I/O or a lock",
      "running_for_seconds": 5400
    },
    {
      "task_id": "7c6b5a4f3e2d41c0b9a8f7e6d5c4b3a2",
      "name": "ruby-hotel-falcon-two",
      "func": "documents.tasks.generate_pdf",
      "group": null,
      "tenant": "acme",
      "host": "worker-1",
      "process": "Process-1:2",
      "pid": 4179,
      "timeout": 600,
      "started": "2026-10-09T05:03:19.735551+00:00",
      "last_seen": "2026-10-09T05:07:09.735551+00:00",
      "rss_mb": 1985.7,
      "rss_start_mb": 190.3,
      "memory_changed_at": "2026-10-09T05:07:09.735551+00:00",
      "cpu_seconds": 120.5,
      "cpu_percent": 52,
      "cpu_recent_percent": 98,
      "status": "lost",
      "reason": "the worker process died while running it",
      "running_for_seconds": 640
    },
    {
      "task_id": "9f8b2c1e4d7a4b6c8e0f1a2b3c4d5e6f",
      "name": "oscar-delta-nine-lemon",
      "func": "reports.tasks.export_payroll",
      "group": null,
      "tenant": "acme",
      "host": "worker-1",
      "process": "Process-1:3",
      "pid": 4182,
      "timeout": 600,
      "started": "2026-10-09T05:12:24.735551+00:00",
      "last_seen": "2026-10-09T05:13:47.735551+00:00",
      "rss_mb": 212.4,
      "rss_start_mb": 180.1,
      "memory_changed_at": "2026-10-09T05:13:47.735551+00:00",
      "cpu_seconds": 41.0,
      "cpu_percent": 43,
      "cpu_recent_percent": 51,
      "status": "running",
      "reason": null,
      "running_for_seconds": 95
    }
  ]
}

Reading it:

  • sync_employees is frozen and waiting. It has run for 90 minutes on almost no CPU, and its memory hasn't moved. That usually means a call to another system with no timeout.
  • generate_pdf was lost. Its memory grew from 190 MB to almost 2 GB before the worker died, which points to an out-of-memory kill. It is now also in django-q's Failed tasks.
  • export_payroll is running normally.
  • last_seen is the last heartbeat. rss_start_mb and rss_mb are the worker's memory when the task started and at the last heartbeat, and memory_changed_at is when it last moved by 1 MB or more.
  • cpu_recent_percent is CPU use since the previous heartbeat, which says what the task is doing now. cpu_percent is the average since it started.
  • tenant is null in projects without django-tenants.

The same data is available from python manage.py qwatchdog --json.

Multi-tenant projects

With django-tenants and django-tenants-q, each record carries the task's schema_name, and lost tasks are saved as failures in that tenant's schema. Without django-tenants, nothing changes.

Configuration

All optional:

Q_WATCHDOG = {
    "CACHE": "default",  # must be django-redis, Django's RedisCache, or LocMemCache
    "HEARTBEAT_SECONDS": 60,
    "STALE_AFTER_SECONDS": 180,  # no heartbeat for this long: the worker is gone
    "FROZEN_AFTER_SECONDS": 3600,  # None turns frozen detection off
    "RECORD_TTL_SECONDS": 604800,
    "SAVE_LOST_AS_FAILED": True,  # put lost and timed-out tasks into django-q's Failed tasks
    "TENANT_KWARG": "schema_name",
    "ACCESS_CHECK": "django_q_watchdog.conf.staff_only",
}

The cache must be shared by your web processes and your cluster, which is why a Redis cache is expected in production. Memory figures come from /proc and are Linux-only; everything else works on any platform.

How it works

  • In each worker: django-q's pre_execute signal writes a record for the task, and a small thread refreshes it every HEARTBEAT_SECONDS with memory and CPU usage. A refresh only updates a record that still exists, so a finished task is never brought back.
  • When django-q saves the task (success or failure), the record is deleted.
  • The check looks at what's left: a record whose heartbeat stopped belongs to a worker that died or was killed; a record still refreshing but whose memory hasn't moved for an hour belongs to a frozen task. Whether it's waiting or busy is decided by its CPU use since the previous heartbeat, not its average, so a task that worked hard and then got stuck is reported as waiting.

Tested end to end against a real cluster, by killing worker processes and letting tasks hang past the timeout, on django-q 1.3.9 with Django 4.2 and on django-q2 with Django 5.2.

Development

pip install -e ".[dev]"
pytest

License

MIT

Metadata

Release files for django-q-watchdog 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for django-q-watchdog 0.1.1
File Size Uploaded
django_q_watchdog-0.1.1.tar.gz 20.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for django-q-watchdog 0.1.1
File Interpreter ABI Platform
django_q_watchdog-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 36.6 kB

Release files / django_q_watchdog-0.1.1.tar.gz

Download URL django_q_watchdog-0.1.1.tar.gz
Size 20.8 kB
Tags Source
SHA-256 checksum
How to use checksums
5e73daf22836956d568b1dd3d6231765b3e808640a74a6af115e1a43812368b9
BLAKE2b-256 checksum
How to use checksums
3f15d2584f03b0ee983e575a8256fb7d308a8bfec9764994c3f187263cd71466
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / django_q_watchdog-0.1.1-py3-none-any.whl

Download URL django_q_watchdog-0.1.1-py3-none-any.whl
Size 15.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
51cdb087cabc912e8b573a1c0186b19dfaa6f20b6896f87dbf6139231e3e1ed7
BLAKE2b-256 checksum
How to use checksums
44bfd2bd03eb1f477454e0b0b10156fbc70b720a2f1c670124555ee0ca4fc358
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page