A Python library for managing background worker processes with persistent state, automatic recovery, and a CLI.
Project description
Crazy Workers
A Python library for managing background worker processes with persistent state, automatic crash recovery, and a built-in CLI.
Features
- Persistent State — SQLite database tracks worker status, PIDs, and parameters across restarts.
- Backend Integration — Co-locate crazy_workers' tables in your project's database (pass a SQLAlchemy engine or URL), inject a shared
DATABASE_URLinto every worker, and recover workers automatically when the backend boots. See Backend integration. - Process Management — Start, stop, and monitor background Python scripts as independent OS processes.
- Automatic Recovery — Detects crashed workers and restarts them on application boot.
- Reconcile Daemon — A single-owner loop (
python -m crazy_workers.daemon) drives actual state toward the desired state that clients (WorkerClient, CLI) write to the shared database: it starts and stops workers, restarts crashes with exponential backoff, and recycles a running worker whose parameters changed in the database — the process is never left silently running with a stale spec. - Automatic Boot-Restore — On Linux and Windows, starting a worker transparently installs a per-user OS hook (a systemd user unit / a logon Scheduled Task) that runs recovery after a machine reboot — no host application required. Opt out with
CRAZY_WORKERS_NO_BOOT. See Automatic boot-restore. - Child Process Control — On stop, terminates unmanaged subprocesses while preserving independently-managed nested workers.
- CLI Interface — Manage workers from the terminal with interactive prompts and auto-discovery (see CLI.md).
- Security — Worker types and keys are restricted to a safe identifier charset (
A-Z a-z 0-9 _ -), with a defence-in-depth check that the resolved script path stays inside the workers directory. This blocks path traversal on both Unix and Windows (including drive-relative names likec:evil). - Observability — Per-worker file logging; all service files (DB, lock, logs) live in a
.service/folder inside your workers directory. - Zombie Protection — Distinguishes active processes from zombies using
psutil. - PID-Reuse Safe — Each worker is tagged with an identity token on its command line; recovery and stop confirm a PID still belongs to the worker before acting, so a recycled PID is never mistaken for (or worse, killed as) a live worker. Works on both Unix and Windows.
- Gunicorn-safe — File-based lock prevents concurrent recovery runs across multiple workers.
- Testable — Drive your orchestration with a
FakeBackend(no real processes) and import polling helpers for the few genuine end-to-end tests. See Testing your app.
Installation
pip install crazy-workers
Or from source:
git clone https://github.com/Vanni-broUser/crazy-workers
cd crazy-workers
pip install .
Quick Start
1. Create a worker script
# workers/my_worker.py
import time
from crazy_workers import parse_params
params = parse_params()
duration = params.get('duration', 60)
for _ in range(duration):
time.sleep(1)
2. Manage it from Python
from crazy_workers import WorkerManager
manager = WorkerManager('workers')
# Start
success, result = manager.start_worker(
'my_worker',
worker_key='job_1',
parameters={'duration': 30},
)
print(result['pid']) # OS process ID
print(result['status']) # 'RUNNING'
# List
for w in manager.list_workers():
print(w['worker_key'], w['status'])
# Stop
manager.stop_worker('job_1')
# Recover crashed workers (call on app startup)
restarted = manager.recover_workers()
manager.dispose() # releases DB connection; does NOT kill workers
3. Or from the CLI
crazy-workers status
crazy-workers start my_worker --key job_1 --params '{"duration": 30}'
crazy-workers stop job_1
See CLI.md for full CLI documentation.
API Reference
WorkerManager(workers_dir, create_dir=True, ...)
| Parameter | Type | Default | Description |
|---|---|---|---|
workers_dir |
str |
'workers' |
Directory containing worker .py scripts |
create_dir |
bool |
True |
Create workers_dir and .service/ if they don't exist |
backend |
ProcessBackend |
None |
Process backend; the default spawns real subprocesses (tests inject a fake) |
auto_boot |
bool |
True |
Install the per-user OS boot-restore hook on first start — see Automatic boot-restore |
boot_provider |
BootProvider |
None |
Override the boot-restore mechanism (mainly a test seam) |
db_url |
str |
None |
SQLAlchemy URL for worker state; defaults to SQLite under .service/ |
engine |
Engine |
None |
Reuse an existing SQLAlchemy engine so the tables live in your database; not disposed by crazy_workers |
worker_env |
dict |
None |
Environment variables injected into every spawned worker (e.g. DATABASE_URL) |
auto_recover |
bool |
True |
Recover dead-but-RUNNING workers when the manager is constructed |
create_tables |
bool |
True |
Create crazy_workers' own tables on init; set False when the host owns the schema via its migrations |
See Backend integration for db_url / engine / worker_env / auto_recover / create_tables.
start_worker(worker_type, worker_key=None, parameters=None, env=None)
| Parameter | Type | Default | Description |
|---|---|---|---|
worker_type |
str |
— | Filename (without .py) of the worker script |
worker_key |
str |
worker_type |
Unique identifier; allows multiple instances of the same type |
parameters |
dict |
{} |
JSON-serializable dict passed as sys.argv[1] to the worker |
env |
dict |
None |
Extra environment variables injected into the worker process |
Returns (bool, dict | str) — (True, worker_dict) on success, (False, error_message) on failure.
Note on
RUNNING: success means the worker was spawned and survived a brief startup grace period that catches immediate launch failures (bad import, missing module). It does not guarantee the worker will run to completion — a worker that fails later is still reportedRUNNINGuntil the nextlist_workers()/recover_workers()reconciles its state.
stop_worker(worker_key)
Gracefully terminates the worker (SIGTERM → SIGKILL after timeout). Returns (bool, str).
list_workers()
Returns a list of worker dicts including RUNNING, STOPPED, CRASHED, and NEVER_STARTED (filesystem-discovered) workers.
recover_workers()
Restarts any worker whose DB status is RUNNING but whose process is dead. Uses a file lock to prevent concurrent recovery. Returns a list of restarted keys.
You rarely call this directly: it runs automatically when the manager is constructed (
auto_recover=True) and after a machine reboot via the boot hook. It remains available as an explicit, idempotent trigger.
dispose()
Closes the database connection and clears internal process references. Does not kill background workers — they continue running independently.
Worker Script Contract
A worker receives its parameters as a JSON string in sys.argv[1]. Use
parse_params instead of decoding it by hand:
from crazy_workers import parse_params
params = parse_params()
# ... do work ...
Without arguments it is lenient: a worker launched with no parameters gets
{}. Pass required= to abort before any real work when a parameter the
worker cannot run without is missing or empty — the worker exits with code 1
and a message on stderr:
params = parse_params(required=('device_id', 'output_dir'))
parse_params(argv=...) accepts an explicit argv list, which makes the
parsing trivially testable without patching sys.argv.
A worker is a separate process, so it cannot be handed a live object (e.g. a DB
connection). Pass configuration instead: the manager's worker_env (and any
per-call env) is injected as environment variables, so a worker reads, say,
os.environ['DATABASE_URL'] and opens its own connection. See
example_app/workers/db_writer.py.
Testing your app
Code that uses WorkerManager has two kinds of logic worth testing:
orchestration (which workers start, pairing, rollback, recovery) and the
work itself (does the worker actually do its job). crazy_workers.testing
makes both fast and non-flaky.
Orchestration — without launching a single process
WorkerManager.for_testing() wires a FakeBackend: spawning and termination
are faked, but the state machine (SQLite, recovery, validation) stays real
and runs in-process. The backend is exposed as manager.test for assertions.
from crazy_workers import WorkerManager
def test_starts_recorder_and_renamer():
manager = WorkerManager.for_testing('workers') # FakeBackend, no processes
manager.start_worker('recorder', worker_key='cam1', parameters={'device': 'cam1'})
manager.start_worker('renamer', worker_key='renamer_cam1', parameters={'output_dir': '/data/cam1'})
assert manager.test.started_types == ['recorder', 'renamer']
assert manager.test.is_running('cam1')
assert manager.test.parameters_for('renamer_cam1') == {'output_dir': '/data/cam1'}
manager.dispose()
def test_recovery_restarts_a_crash():
manager = WorkerManager.for_testing('workers')
manager.start_worker('recorder', worker_key='cam1')
manager.test.crash('cam1') # simulate an unexpected death
manager.recover_workers() # the real recovery path runs in-process
assert manager.test.start_count('cam1') == 2
assert manager.test.is_running('cam1')
manager.dispose()
workers_dir must still contain the <type>.py files (start_worker checks
the script exists), but the fake backend never executes them — empty files are
enough.
manager.test exposes:
| Member | Returns |
|---|---|
started_types / started_keys |
every spawn, in order (a restart appears twice) |
running_keys |
keys whose latest process is still "alive" |
is_running(key) |
bool |
start_count(key) |
how many times the key was (re)started |
parameters_for(key) |
parameters of the most recent spawn |
crash(key) |
simulate an unexpected death (without a stop) |
The real thing — polling helpers, not sleep
The few tests that must launch real processes should wait on conditions,
never fixed sleeps. crazy_workers.testing exposes the helpers used by the
library's own suite:
from crazy_workers.testing import wait_for_worker_status, wait_for_log, wait_for_pid_dead
manager = WorkerManager('workers')
ok, res = manager.start_worker('recorder', worker_key='cam1')
wait_for_worker_status(manager, 'cam1', 'RUNNING')
wait_for_log('workers/.service/logs/cam1.log', 'recording started')
# ... assert the worker actually did its job ...
manager.stop_worker('cam1')
wait_for_pid_dead(res['pid'])
Available: wait_for, wait_for_file, wait_for_log, wait_for_worker_status,
wait_for_worker_in_db, wait_for_worker_pid, wait_for_pid_dead. Each raises
AssertionError with a useful message on timeout.
What the fake covers — and what it doesn't.
for_testing/FakeBackend tests orchestration, not "the worker actually records / sends / converts". Keep a small number of real end-to-end tests for that, made stable with the polling helpers above, and move everything else — the bulk — into the fast, deterministic fake world.
Project Structure
crazy_workers/ # Library package
core/ # WorkerManager, process engine, recovery lock
boot/ # Automatic per-user boot-restore (systemd user unit / scheduled task)
cli/ # CLI entry point, commands, discovery
database/ # SQLAlchemy schema and pluggable storage (SQLite, or a shared engine/URL)
testing/ # FakeBackend + polling helpers for consumer test suites
example_app/ # Flask demo application
app.py
workers/ # Example worker scripts
tests/
core/ # Unit tests for core modules
cli/ # Unit tests for CLI modules
database/ # Unit tests for storage layer
testing/ # Tests for the FakeBackend and polling helpers
integration/ # Full-stack integration tests (real processes)
app/ # Tests for the example Flask app
Flask Integration
⚠️ Security:
start_worker()runs the worker script named by the caller. Exposing it over HTTP makes it a privileged operation — anyone who can reach the route can launch any script in your workers directory. Put such routes behind authentication, and prefer validatingworker_typeagainst a known allow-list of expected workers. The example below is a minimal demo with no auth.
from crazy_workers import WorkerManager
def create_app(db_engine, db_url):
app = Flask(__name__)
manager = WorkerManager(
'workers',
engine=db_engine, # crazy_workers' tables live in YOUR database
worker_env={'DATABASE_URL': db_url}, # injected into every worker
# auto_recover=True (default): when the app boots, workers left RUNNING
# with a dead PID are restored automatically — no explicit call needed.
)
@app.route('/workers/start', methods=['POST'])
def start():
data = request.json
success, result = manager.start_worker(
data['worker_type'],
worker_key=data.get('worker_key'),
parameters=data.get('parameters', {}),
)
return (jsonify(result), 200) if success else (jsonify({'error': result}), 400)
return app
See example_app/app.py for a complete example.
Backend integration
When crazy_workers runs inside a backend, three options let it cooperate with the project instead of living off to the side:
- Co-locate its tables in your database. Pass an existing SQLAlchemy
engine(or adb_url) toWorkerManager. crazy_workers creates its ownworkerstable inside your database and inherits its persistence and backups — so state survives even if the process/container is recreated. A shared engine is never disposed by crazy_workers; its owner manages it. - Let your migrations own the schema. If your project tracks its schema with
a migration tool (Alembic, etc.), pass
create_tables=Falseso crazy_workers issues no DDL: theworkerstable becomes a normal migration in your history, with a single source of truth and no create-on-import side effect. You own the ordering — the table must exist before the manager queries it, so build the manager after your migrations run (and keepauto_recover=Falseuntil then, since recovery reads that table). See theworkersschema incrazy_workers/database/schema.pyfor the columns your migration must create. - Give workers the connection they need. A worker is a separate process, so
it can't receive a live DB connection — pass the configuration instead.
worker_env={'DATABASE_URL': ...}is injected into every spawned worker (overridable per call viastart_worker(..., env=...)); the worker opens its own connection from it. - Recovery on construction.
auto_recover=True(default) restores dead-but-RUNNINGworkers when the manager is built, so a restarting backend brings its workers back with no explicit call. The CLI and boot entrypoint set it toFalse(management/one-shot, not supervision).
Gunicorn / Multi-Process Servers
When using a pre-fork server like Gunicorn:
- Recovery is atomic — a file lock (
.service/workers.db.recovery.lock) ensuresrecover_workers()runs once even when multiple workers boot simultaneously. - Workers outlive their parent — if a Gunicorn worker is recycled, background processes keep running. The next recovery cycle re-attaches or restarts them.
Automatic boot-restore
Starting a worker transparently installs a per-user OS hook for its workers
directory, so workers come back after a reboot without any host application
running. The hook calls the internal entrypoint python -m crazy_workers.boot --workers-dir <dir>, which runs recover_workers(). The install is
best-effort and happens at most once per directory (a marker lives in
.service/boot.json); a failure never blocks the worker from starting and is
reported by crazy-workers status.
| Platform | Mechanism | When it runs |
|---|---|---|
| Linux | systemd user unit (~/.config/systemd/user) |
at user login — or at true boot if lingering is enabled |
| Windows | logon Scheduled Task (schtasks /SC ONLOGON) |
at user logon |
Unattended boot — the honest caveat. The recovery itself is durable across
consecutive reboots (a worker left RUNNING with a dead PID is restarted, and
this is PID-reuse safe). But running it without any login depends on the OS:
- Linux: enable lingering once, as root —
loginctl enable-linger <user>— so the user's systemd starts at boot (the same model as the Docker daemon restarting containers). Without it, restore runs at the next login. - Windows: a user task fires at logon; true pre-logon start needs autologon or an administrator-installed task.
crazy-workers status always reports whether the hook runs at boot or
at login, so this is never silently wrong.
Opting out. Set CRAZY_WORKERS_NO_BOOT (any non-empty value) to disable the
automatic install entirely — useful in containers, CI, or when you manage boot
yourself. You can also pass WorkerManager(..., auto_boot=False).
Logs
Each worker's stdout/stderr is appended to .service/logs/<worker_key>.log. These files are written directly by the worker process, so the library does not rotate them — they grow until you act. For long-lived deployments, rotate them externally (e.g. logrotate with copytruncate) or have your worker script configure its own logging.handlers.RotatingFileHandler instead of writing to stdout/stderr.
Development
Setup
git clone https://github.com/Vanni-broUser/crazy-workers
cd crazy-workers
pip install -e .[dev]
Commands
# Lint and format
ruff check . --fix && ruff format .
# Run tests
pytest
# Run tests with coverage
coverage run -m pytest && coverage report
Standards
See AI.md for the full coding and testing standards used in this project.
Support the project ❤️
crazy-workers is free and open-source (MIT). If it saves you time or powers your work, consider supporting its development:
- GitHub Sponsors — recurring or one-time, 0% platform fee: https://github.com/sponsors/Vanni-broUser
- Stripe — one-time card donation via secure checkout: Stripe Payment Link (coming soon)
- ⭐ Star the repo — free, and it really helps visibility.
The Stripe link is configured via a Payment Link. Replace the placeholder above (and in
.github/FUNDING.yml) with your realhttps://buy.stripe.com/...URL once created.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file crazy_workers-1.6.1.tar.gz.
File metadata
- Download URL: crazy_workers-1.6.1.tar.gz
- Upload date:
- Size: 51.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.2.0 CPython/3.11.15
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8d73d3de81635bba9b1f00a27292cd777ef9093f01e8f679cc368e756054a56d
|
|
| MD5 |
7a9c7cb58b6d20d2a74ed82920385b3f
|
|
| BLAKE2b-256 |
54f36a64d6d8039a50c6f8f32a800d732be831622861b835efa4e40e83871f17
|
File details
Details for the file crazy_workers-1.6.1-py3-none-any.whl.
File metadata
- Download URL: crazy_workers-1.6.1-py3-none-any.whl
- Upload date:
- Size: 56.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.2.0 CPython/3.11.15
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
917872a5cf74bad98a1ae4aee742a1fd7e1afc66b13637c891936981d5dffdbf
|
|
| MD5 |
056678a12e0a859fd96d665c1a1c6679
|
|
| BLAKE2b-256 |
e1eb8045bda7adbbba025be8e98a603965b98763bae434ba9c1e28a2521de813
|