Landing Zones
Automated data transfer system using rsync with cron job generation.
Quick Start
# Install
pip install -e .
# Generate cron files, transfer scripts, and validation wrappers
landingzones --help
landingzones --config config/config.yaml build
landingzones build
# Check deployment readiness
landingzones validate deployment
# Run a hop-local validation
landingzones validate hop <flow_group> preflight
landingzones validate hop <flow_group>
# Run toy data through the configured flows
landingzones validate integration
Project Structure
landingzones/
├── src/landingzones/ # Main package
│ ├── cli.py # Top-level operator CLI
│ ├── generate_cron_files.py # Cron generation tool
│ ├── check_deployment_readiness.py
│ ├── plot_transfer_status.py
│ └── config/transfers.tsv # Default config
├── input/ # Default input directory
├── output/ # Default output directory
│ ├── crontab.d/ # Generated cron files
│ ├── scripts/ # Generated transfer scripts
│ └── validation_scripts/ # Generated validation wrappers
├── log/ # Default log directory
├── tests/ # Test suite
├── pyproject.toml # Package config
└── README.md
Configuration
The system is configured via a tab-separated transfers.tsv file:
| Column | Description | Example |
|---|---|---|
identifiers |
Unique transfer ID used for generated shell script names | transfer_001, server1_to_server2 |
runtime_id |
Required deploy/artifact identity used for cron grouping and filtering | server1_prod.user1 |
system |
Configured system key used for managed paths and flock settings | server1, localhost |
users |
Optional user/account context for review and generated headers | user1, local |
source |
Source directory path | /srv/data/src/ |
source_port |
SSH port for remote sources (optional) | 2222 |
destination |
Destination (local or remote) | user@host:/dest/ |
destination_port |
SSH port (optional) | 225 |
rsync_options |
Additional rsync flags | --chown=:group |
io_nice |
Optional ionice settings for rsync |
-c2 -n7 |
log_file |
Log file name resolved under the system log folder | transfers.log |
flock_file |
Lock file name resolved under the system flock folder | transfer.lock |
flow_group |
Optional logical flow label shared by multi-hop transfers | labnet_to_seqdata |
is_entry_point |
Optional TRUE marker for the first hop of a logical flow |
TRUE |
is_end_point |
Optional TRUE marker for the final hop of a logical flow |
TRUE |
readiness_policy |
Entry-point readiness policy; defaults to current direct behavior | direct, stable_snapshot |
readiness_stable_observations |
Number of identical source inventories required before stable_snapshot is eligible |
2 |
readiness_quiet_seconds |
Minimum seconds since the last observed source change before stable_snapshot is eligible |
300 |
readiness_fingerprint_mode |
Source inventory fingerprint mode | path_size_mtime |
Future todo: add an optional second, per-remote-host lock for cross-server
transfers. The existing flock_file prevents one transfer from overlapping
with itself; a host-level lock would limit concurrent SSH/rsync handshakes
against the same remote server when many transfer rows run on the same cron
schedule.
Example
identifiers runtime_id system users source source_port destination destination_port rsync_options io_nice log_file flock_file
local_copy localhost_test.testuser localhost testuser input/* output/ transfers.log landingzones.lock
CLI Commands
# Generate cron files with defaults
landingzones build
# Generate only selected runtime IDs from a shared transfers.tsv
landingzones build --runtime-id server1_prod.user1 --runtime-id server2_prod.user2
# Check deployment readiness
landingzones validate deployment
# Run a hop-local validation wrapper through the CLI
landingzones validate hop <flow_group>
# Seed toy data and run the real scripts/logs/locks
landingzones validate integration
# Generate an HTML health dashboard from a shared transfer TSV log
landingzones report transfers output/log/Landing_Zone_server1_prod.user1.transfers.tsv
# Synchronize expected routes, ingest one or more Event Spools, and serve live monitoring
landingzones monitor sync-definitions
landingzones monitor ingest output/log/Landing_Zone_server1.transfers.tsv
landingzones monitor serve
Database-Backed Transfer Event Monitoring
Generated runtimes write schema-version-1 Transfer Events to the existing
per-system common-status path. The file is now an append-only Event Spool with
immutable event_id values and separate run_id and attempt_id identities.
The same event identity is retained in applicable portable per-run history.
Runtime scripts never connect to the monitoring database, so database or
ingestor availability cannot block transfer execution.
Configure the separately invoked monitoring processes with a synchronous SQLite SQLAlchemy URL:
monitoring_database_url: sqlite:///output/landingzones-monitoring.sqlite
The same value can be supplied with --database-url or
LZ_MONITORING_DATABASE_URL. Keep the SQLite file on a host-local filesystem;
schema version 1 does not claim support for other database backends.
Synchronize current Transfer Definitions separately from operational events, then ingest one or more independently produced spools:
landingzones monitor sync-definitions \
--config config/config.yaml \
--runtime-id server1_prod.user1
landingzones monitor ingest \
output/log/Landing_Zone_server1.transfers.tsv \
--spool-id server1-prod
Ingestion consumes only complete newline-terminated rows, checkpoints each
spool, and safely ignores replayed event_id values. It rejects unversioned or
unsupported headers rather than guessing their layout. During cutover, stop
old writers, archive the old common-status TSV if desired, and create a fresh
file (or let the first schema-v1 runtime create it); never append v1 rows below
the old unversioned header. A malformed row is reported as a warning, skipped,
and included in the checkpoint so one bad row cannot hold back later valid
events; header and checkpoint errors remain fatal.
Start the live service with:
landingzones monitor serve --host 127.0.0.1 --port 8080
To run ingestion and reporting in one service, configure explicit schema-v1 spool paths and listener settings:
monitoring_spools:
- output/log/events.tsv
monitoring_ingest_interval: 60
monitoring_host: 127.0.0.1
monitoring_port: 8080
Then use:
landingzones --config config/config.yaml monitor serve
The service ingests at startup and polls independently of browser requests.
Missing spools are retried because a writer may create them later. An
incompatible existing spool fails startup; later ingestion errors are logged
and retried while the report remains available. Do not run a separate ingestion
loop for the same sources when the combined service owns them.
Without monitoring_spools, serving retains its database-reader behavior.
Listener CLI options override the config. LZ_MONITORING_SPOOLS accepts a
comma-separated list; LZ_MONITORING_INGEST_INTERVAL, LZ_MONITORING_HOST and
LZ_MONITORING_PORT override their corresponding YAML settings.
Service startup loads the configured transfer file and synchronizes Transfer
Definitions before accepting requests; use --transfers and repeatable
--runtime-id options to override the configured inventory or selection.
Every HTML page load and JSON request queries the current database, and HTML
monitoring pages show their last queried UTC time and auto-refresh every 60
seconds while preserving the current URL and query filters. The run list
defaults to transfers with a discovered directory; route-only observations,
including configured-but-never-observed routes and pre-discovery failures, are
available through the Show them toggle. The JSON API is available at /api/runs and
/api/runs/<run_id>. Repeatable query parameters include runtime_id,
system, execution_user, transfer_identifier, tag, state, and
reason_code. The existing
landingzones report transfers command remains a legacy static schema-0
reporting surface; it is not the operational reader for schema-version-1
Event Spools.
Remote source discovery records source_missing only when the remote probe
successfully reports that the directory is absent. SSH failures retain the
SSH exit code and diagnostic text and use one of ssh_timeout,
ssh_authentication_failed, ssh_host_unreachable, or ssh_failed as the
stable reason_code for downstream use cases.
Generated Cron Format
*/15 * * * * /bin/sh output/scripts/local_copy.sh
Transfer Catalog Loading Modes
The transfer catalog is the owner of transfer loading invariants. Future transfer-column, filtering, endpoint-expansion, validation, artifact-naming, boolean, tag, and warning behavior should land at the catalog seam before command code consumes the rows.
load_runtime_transfer_catalog is the build/runtime loading mode. It keeps
runtime validation enabled, including required runtime file columns such as
log_file and flock_file, disabled-row filtering, exact runtime_id
filtering, path-variable endpoint expansion, generated artifact names, boolean
normalization, tag normalization, and shared lock/log warning metadata.
load_reporting_transfer_catalog is the reporting/analysis loading mode. It
uses the same normalized transfer facts, filtering, endpoint expansion,
booleans, tags, and artifact naming, but reporting analysis can omit
runtime-only log_file and flock_file columns because dashboards inspect
transfer metadata and status logs instead of generating runnable scripts.
Command boundaries:
landingzones builduses the runtime catalog.landingzones validate deploymentuses the runtime catalog.landingzones validate integrationuses the runtime catalog.landingzones validate separationuses the reporting catalog.landingzones report transfersuses the reporting catalog.
Entry-Point Readiness Policies
Entry-point readiness checks protect flows where the upstream producer writes directly into a visible source tree. They are a mitigation, not a guarantee: only producer-controlled atomic publish or a producer completion marker can prove that a run is complete.
readiness_policy=direct keeps the existing behavior. When a row has
is_entry_point=TRUE, the generated script mints portable metadata, archives
the visible run directory, and removes the unpacked source contents from that
hop before transfer. Use this for producers that publish atomically, or where
the operator accepts the current direct-archive contract.
readiness_policy=stable_snapshot is the consumer-side mitigation for local
entry-point sources. The generated script observes each top-level run without
modifying producer-visible contents, records durable state under the
entry-root .landing_zones_readiness sidecar directory, and waits until the
same path/size/mtime fingerprint has been seen for
readiness_stable_observations observations and the configured
readiness_quiet_seconds has elapsed. Once eligible, it copies the run into a
Landing Zones-owned private snapshot, confirms that the source still matches
the snapshot, and archives from that private copy. The original producer-visible
run is left in place for separate retention or explicit operator cleanup.
Generated stable-snapshot scripts use the Python interpreter that ran
landingzones build for the readiness helper; set LANDINGZONES_PYTHON in the
runtime environment only when that interpreter must be overridden.
Shared Main Transfer Locks
flock_file is the top-level transfer lock for one generated script. Remote
transfers run startup jitter before acquiring the main flock_file, then use
non-blocking flock -n for the main lock. This spreads same-minute cron starts
before the lock race, but a transfer can still skip a run if another process
keeps the same main lock busy after the jitter. This behavior does not change
the cron cadence in transfers.tsv; for example, a */2 * * * * row remains a
two-minute schedule.
landingzones build and landingzones validate deployment warn with Shared main transfer locks detected when multiple transfer rows in the same runtime
resolve to the same main flock_file. Treat that warning as an operator review
item: either the shared lock is intentional serialization, or one of the rows
should get a distinct flock_file.
The shared main-lock warning is different from the generated
transfer-status and notification-status locks. Files such as
Landing_Zone_<system>.transfers.lock and
Landing_Zone_<system>.notifications.lock protect short TSV append sections and
are expected to be shared by all scripts for a runtime.
To verify a generated script manually, hold its resolved flock_file, run the
script with debug output, and check the message order:
script=/path/to/generated_transfer.sh
lock=$(sed -n 's/^flock_file="\([^"]*\)"/\1/p' "$script" | sed "s|\$HOME|$HOME|")
mkdir -p "$(dirname "$lock")"
(
exec 200>"$lock"
flock -n 200
LZ_DEBUG_CLI=1 /bin/sh "$script"
) 2>&1
For remote transfers, debug output should show startup delay ... before
using lock file ...; if the held lock wins, the script then reports
lock busy, exiting.
Installation
# Development mode
pip install -e ".[report]"
# With test dependencies
pip install -e ".[test]"
# Production
pip install .
Lab Sequencer Bundle
For lab machines where a managed Python environment is awkward, build a
relocatable bundle using a python-build-standalone runtime. Build it on a
machine that matches the lab sequencer OS, architecture, and libc family.
Download or provide a python-build-standalone install_only archive, then run:
cd app
python scripts/build_python_standalone_bundle.py --python-archive /path/to/cpython-*-install_only.tar.*
If you already extracted the runtime, point at its Python executable instead:
python scripts/build_python_standalone_bundle.py --python-bin /path/to/python/install/bin/python3
With Pixi, the app includes a packaging task that downloads a matching
python-build-standalone runtime using getpybs:
cd app
pixi run build-standalone
Build the lab Linux artifact on Linux. A bundle built on macOS contains a macOS
Python runtime and will fail on the sequencer with cannot execute binary file.
Before copying a tarball to the lab host, verify the bundled runtime:
file packaging/dist/landingzones-standalone/python/bin/python3
packaging/dist/landingzones-standalone/python/bin/python3 -c "import platform; print(platform.system(), platform.machine())"
Expected for the current lab machines is Linux/x86_64.
The standalone bundle installs the core operator CLI without pandas, so it is
intended for build, validate, and deploy on locked-down lab machines.
landingzones report transfers remains a reporting extra and should run from
an environment with landingzones[report] installed.
The same bundle can be produced by the GitHub Actions workflow
Build Standalone Bundle. Run it manually from Actions, or push a v* tag.
It uploads landingzones-standalone-linux-x86_64 containing:
landingzones-standalone-linux-x86_64.tar.gz
For v* tags, the workflow also creates or updates the matching GitHub Release
and uploads landingzones-standalone-linux-x86_64.tar.gz as a release asset.
To release a version, update src/landingzones/__init__.py and pixi.toml
to the same version and merge the changes into main. Tag that commit with
v<version> and push the tag to GitHub. The standalone workflow builds from
that tag, verifies the bundle CLI, and publishes the matching release asset.
A version bump on main alone does not create a tag or release.
For recovery or rebuilds, manually dispatch Build Standalone Bundle against
the version tag. This replaces the existing standalone release asset. A manual
run against a branch uploads an Actions artifact only.
The bundle is written to:
app/packaging/dist/landingzones-standalone/
app/packaging/dist/landingzones-standalone.tar.gz
Copy the tarball to the lab machine, extract it, and run it like the normal CLI:
./landingzones --config config/config.yaml build
./landingzones --config config/config.yaml validate deployment
For offline builds, pass --wheelhouse /path/to/wheels so dependencies are
installed from local wheels. The legacy shell wrapper still works:
./scripts/build_python_standalone_bundle.sh --python-archive /path/to/cpython-*-install_only.tar.*
The bundle carries Python and Python packages only; the target machine still
needs system tools such as rsync, ssh, flock, curl, and cron.
Testing
For a local lab-machine → Cluster A → Cluster B scenario using real SSH/rsync and synthetic Linux users/groups, see the container transfer lab. It includes project isolation, preprocessing, and outage/retry checks without access to production servers.
# Run all tests
pytest
# Verbose
pytest -v
# With coverage
pytest --cov=landingzones --cov-report=html
# Specific test
pytest tests/test_generate_cron_files.py::TestClassName::test_method
Validation Modes
The operator-facing validation surface has three modes:
landingzones validate deploymentlandingzones validate hop <flow_group> [preflight|run]landingzones validate integration
landingzones validate integration is the heavier integration-style test mode. It copies toy data into the configured starting locations, generates the real shell scripts, and runs the transfers using the normal log and flock paths.
Use landingzones validate integration --slow when you want the harness to print the result of each completed step and wait for Enter before running the next one.
Generated transfer scripts create portable .landing_zones sidecars for every enabled transfer. flow_group is optional sidecar metadata: when a transfer mints a new sidecar the value may be blank, and downstream transfers preserve the value already stored in the sidecar.
When a row has is_entry_point=TRUE, the generated script archives top-level
run directories into .landing_zones/landingzone-run-archive.tar before
transfer. The default readiness_policy=direct archives from the visible source
and removes unpacked contents from that hop. The stable_snapshot policy
archives from a Landing Zones-owned private snapshot and leaves the original
producer-visible run in place. Intermediate hops then move the archive plus the
.landing_zones metadata instead of thousands of individual payload files. When a later row has
is_end_point=TRUE, the generated script extracts the archive after staging
promotion and removes the archive from the final destination. Archive extraction
validates that tar entries are relative paths before unpacking.
Generated Validation Wrappers
Each flow_group with exactly one is_entry_point=TRUE row gets a generated wrapper in the configured validation-scripts directory:
output/validation_scripts/lz_run_validation_<flow_group>.sh
Use landingzones validate hop <flow_group> as the main interface. The generated wrapper remains available directly and bakes in:
- the entry directory for that flow
- the immediate next hop for preflight checks
- the default fixture directory under
test_data - the
flow_groupand producer labels used in theLZTEST_...folder name
Typical usage:
# Regenerate scripts after changing config/transfers
landingzones --config config/config.yaml build
# Check only the current hop structure and immediate next-hop access
landingzones validate hop local_labnet_to_server1_data preflight
# Inject a validation run with the baked-in defaults
landingzones validate hop local_labnet_to_server1_data
# Inject a validation run with an explicit token suffix
landingzones validate hop local_labnet_to_server1_data --token ABCD
# Direct wrapper execution still works if needed
./output/validation_scripts/lz_run_validation_local_labnet_to_server1_data.sh
Wrapper/CLI behavior:
- no action defaults to
run preflightchecks only the current hop plus its immediate next hop- options-only invocation such as
--token ABCDalso defaults torun
Use landingzones validate hop for lightweight producer-side validation. Use landingzones validate integration when you want the heavier integration test that seeds toy data and executes the full generated transfer chain.
Required config in your deployment config.yaml:
transfers_file: input/transfers.tsv
test_data: tests/toy_data/
validation_scripts_dir: output/validation_scripts/
rit_managed_locations:
test_local: tests/test_local
flock_paths:
test_local: /opt/homebrew/bin/flock
rit_managed_folder_structure:
log: output/log/
flock: output/flock/
sh_output: output/scripts/
crontabs: output/crontab.d/
Typical local fixture layout:
deploy/local/
├── config/config.yaml
├── input/transfers.tsv
├── tests/toy_data/
└── tests/test_local/
How to run it:
# Run from the deployment root that owns config/, input/, and tests/
cd deploy/local
# Generate deployment artifacts
landingzones build --config config/config.yaml
# Run the heavier integration test
landingzones --config config/config.yaml validate integration
What it does:
- Filters
transfers.tsvto the currentsystemanduser - Seeds each initial source root from
test_data - Generates scripts into the configured
sh_outputdirectory - Uses the configured
logandflockdirectories - Executes the scripts in transfer order
- Validates that the seeded top-level directories reached the terminal destinations
Before seeding, integration validation summarizes pre-existing endpoint entries.
Expected toy-data directories and visible source entries are blockers because a
new run would mix old and new test data. Unrelated destination leftovers are
reported as extras. An empty .staging directory is reported as managed
persistent staging state and does not block reruns, because successful Landing
Zone Runtime transfers preserve that group-writable staging root. A .staging
path still blocks when it is non-empty, not a usable directory, or cannot be
inspected.
After a successful run it asks whether you want cleanup. Answer y to remove the propagated test directories plus generated log and lock artifacts so the next run starts from the initial state. Answer n to inspect the final tree and logs.
Deployment
- Configure
transfers.tsvwith your routes - Generate cron files:
landingzones build - Deploy:
cp output/crontab.d/*.cron ~/crontab.d/ cat ~/crontab.d/*.cron | crontab -
Or use automated deployment:
landingzones deploy cron
landingzones deploy cron defaults to --cron-scope execution-context. This
scope replaces the active crontab with a previewed Cron Activation Plan for
the current system/user Execution Context: selected runtime cron fragments
are activated, same-context staged runtime cron fragments are preserved,
unidentified staged .cron files are preserved, and foreign or unresolved
runtime fragments are excluded unless you choose a broader scope.
Cron scopes:
execution-context: safe default for shared managed hosts.expected: uses Generated Runtime Metadata, while staying bounded by the current Execution Context.replace-selected: activates only the selected Runtime Selection plus unidentified staged.cronfiles; this is the explicit replacement path for older selected-runtime behavior.staged: activates every staged*.cronfile after preview.
Exact staged filenames can be omitted with repeated CLI flags or config:
landingzones deploy cron \
--exclude-cron-fragment old-runtime.Landing_Zone.cron
cron_fragment_exclusions:
- old-runtime.Landing_Zone.cron
Missing exclusion filenames are shown in the preview and do not block
activation. Non-interactive cron activation still requires
--confirm-cron-activation.
Upgrade note: the older selected-runtime default is now the explicit
replace-selected scope. The compatibility name selected still works, but it
prints a warning and maps to replace-selected.
Development
# Make changes in src/landingzones/
# Run tests
pytest
# Test CLI
landingzones --help
Requirements
- Python >= 3.8
- PyYAML >= 5.0.0
- pandas >= 1.0.0 only for
landingzones report transfers/landingzones[report] - SQLAlchemy >= 2.0,<3 for schema-version-1 SQLite monitoring
- System: rsync, ssh, flock
- System for archived entry/end-point flows: tar
Metadata
Release files for landingzones 1.1.17
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| landingzones-1.1.17.tar.gz | 162.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| landingzones-1.1.17-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 267.4 kB
Release files / landingzones-1.1.17.tar.gz
| Download URL | landingzones-1.1.17.tar.gz |
|---|---|
| Size | 162.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
bd5956e6f5e0ad1a42730ed3f0199e3d8f4ba7ec0a151d1aedfd0119f0ba41ee
|
|
BLAKE2b-256 checksum How to use checksums |
54f3a5e2a0e1f9febecac5722b8684be021d8a52dd4d1e1ea3df6c3b8dfbce5d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.
Transparency logRelease files / landingzones-1.1.17-py3-none-any.whl
| Download URL | landingzones-1.1.17-py3-none-any.whl |
|---|---|
| Size | 104.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f805e4f3f41276c582a1748cd4e6b850363a6ceb6ac17ddb6321a1450055b2b9
|
|
BLAKE2b-256 checksum How to use checksums |
869005fdf32d8d65be2dabfaaeb698ce815548d1b110ef55655dba614b8f7a2a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.
Transparency log