Local, folder-driven Dagster application for CSV transformations with hot-folder sensors and Tkinter configuration UI.
Project description
Air-Gapped Desktop Automation Engine
A local Dagster application for folder-driven CSV transformations. Source files, saved column choices, generated outputs, and Dagster run storage remain on the machine.
How it works
Each business pipeline has an inbox and a Dagster sensor. The sensor waits for
the required number of stable files, moves them to staging, and launches a job.
The job prepares atom parameters, executes the transformation, then moves source
files to processed/ or rejected/.
An atom is a reusable transformation module with an
execute(params_json) -> result_json interface. Atoms are domain-agnostic —
they have no opinion on which column or business context. They do read and write
files, so they are not side-effect-free.
A business pipeline applies an atom to specific domain data with its own folder lifecycle, sensor, and saved configuration.
Registered jobs
| Job | Files needed | Description |
|---|---|---|
vlookup_rollnumber |
2 CSV | Left-join on user-chosen columns (first run: Tkinter picker, then saved) |
groupby_religion |
1 CSV | Group by user-chosen column with counts, sorted descending |
mock_generator |
1+ CSV | Generate synthetic data from file headers using Faker |
vlookup_then_groupby |
2 CSV | Composite: VLOOKUP → GroupBy in a single job |
Each job has a same-named sensor. Sensors must be enabled in Dagster UI before they process inbox files.
Installation
Python 3.10+ required.
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
python -m pip install -r requirements.txt
Tkinter must be available in the Python installation for first-run config.
Run
DAGSTER_HOME="$(pwd)" dagster dev -m definitions
Open http://localhost:3000, enable the desired sensor, and place files in
pipelines/<job-name>/inbox/.
Manual execution without UI:
DAGSTER_HOME="$(pwd)" dagster job execute -m definitions -j vlookup_rollnumber
Source modules
Entry points
| File | Role |
|---|---|
definitions.py |
Registers all jobs and sensors with Dagster |
pipeline.py |
Builds single-atom business jobs via the factory |
runners/composite.py |
Defines the vlookup_then_groupby composite job |
logic/atoms/
| Module | Purpose |
|---|---|
vlookup.py |
Left join two CSVs on specified columns with dtype validation |
groupby.py |
Group by one column, count, sort descending |
mock_generate.py |
Infer Faker generators from headers, write synthetic CSVs |
runners/
| Module | Purpose |
|---|---|
pipeline_factory.py |
Assembles load → preprocess → execute Dagster jobs |
atom_runner.py |
Reusable load/execute ops with processed/rejected/notification lifecycle |
ops.py |
First-run config ops for VLOOKUP, GroupBy, and mock generator |
composite.py |
VLOOKUP → GroupBy chain with shared config |
sensors/
| Module | Purpose |
|---|---|
hot_folder_sensor.py |
Per-pipeline sensor factory: stability check, staging, sweeper |
helpers/
| Module | Purpose |
|---|---|
folders.py |
PipelineFolders class — per-pipeline directory lifecycle |
config_store.py |
JSON config read/write per pipeline |
column_picker.py |
Tkinter UI for VLOOKUP join/output column selection |
notifier.py |
macOS/Windows toast notifications, opens output folder on success |
file_picker.py |
Standalone Tkinter file dialog (not used by registered jobs) |
mock_generator.py |
Standalone CLI mock generator (separate from the atom/job) |
Runtime data
pipelines/<job>/— inbox, staging, processed, rejected, output, config.jsondagster.yaml— SQLite storage configurationdagster_storage/— Dagster run history (persists across restarts)
File and configuration lifecycle
inbox/ → sensor stability check → staging/ → job runs
├→ processed/ + output/ (success)
└→ rejected/ + *.error.txt (failure)
Non-matching files (not *.csv/*.xlsx) are swept to rejected/ after 3
minutes with an error description.
Config init pattern
- First run: pre-process op detects no
config.json→ shows Tkinter UI → saves - Future runs: loads config silently, fully automated
- To reconfigure: delete
pipelines/<job>/config.json
Saved config fields
| Job | Fields |
|---|---|
vlookup_rollnumber |
left_column, right_column, left_output_columns, right_output_columns |
groupby_religion |
group_column |
vlookup_then_groupby |
All VLOOKUP fields + group_column |
mock_generator |
None |
Important behavior
- Two-file jobs: lexical filename order determines left/right assignment.
- Atom failures and config cancellation move staged files to
rejected/. - Unexpected crashes may leave files in
staging/— recover manually. - Notifications are best-effort. Linux has no implementation.
- On success, macOS/Windows opens the output folder immediately (not a click action).
Adding a new business pipeline
- Create or reuse an atom in
logic/atoms/withexecute(params_json) -> str - Create a pre-process op in
runners/ops.py(with config init pattern) - Register in
pipeline.py:my_pipeline = build_business_pipeline( pipeline_name="my_pipeline", atom_module=my_atom, atom_label="My Pipeline", pre_process_op=my_pre_process_op, )
- Add sensor in
definitions.py:my_sensor = build_hot_folder_sensor("my_pipeline", "my_pipeline", min_files=1)
- Folders are created automatically on import.
Development status
No automated tests, lint, or pinned dependency versions. Validate by loading definitions and exercising jobs with disposable CSV inputs.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file etlai-0.1.1.tar.gz.
File metadata
- Download URL: etlai-0.1.1.tar.gz
- Upload date:
- Size: 18.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
fffe07431cc299d5b348b64369234f7a6b12160de705577e26920fcd1474eb6a
|
|
| MD5 |
112d8db4b71d79c7731833a69720d9ff
|
|
| BLAKE2b-256 |
0edf06eb0d605e665d63a3a2f22333347a120b1b6ffaf05e4411ffd8fe432e6f
|
File details
Details for the file etlai-0.1.1-py3-none-any.whl.
File metadata
- Download URL: etlai-0.1.1-py3-none-any.whl
- Upload date:
- Size: 25.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f4e2e6c6666cb505bf5965d08f657a1f32df9c545bbfc9c9c18d0b6e4337149b
|
|
| MD5 |
3e93d6dc3bdb860537969f03258c4538
|
|
| BLAKE2b-256 |
239b4722707515e75237dd88e6a05a6bf4f13325d76186cdc1b22466006e0f58
|